Clinical AI benchmarks poorly reflect how clinicians actually use language models

Analysis of 127,833 clinical queries found that routine use centers on documentation and information retrieval, tasks largely absent from standard evaluations.

Independent
Krithik Vishwanath · Haitong Lin · Anton Alyakin · Jin Vivian Lee · D. Brock Hewitt · Jie J. Yao · +16 more
Research Digest··3 min read
Vishwanath and colleagues compared eight months of real-world use of an institutional language model assistant with 58 public clinical AI benchmarks.

The authors analyzed 127,833 queries submitted by 6,342 physicians, advanced practice providers and nurses across 35 specialties at a large academic health system.

Why this paper

Independent

In one line

Clinical LLM benchmarks poorly represent deployed use, which centers on documentation and knowledge retrieval rather than diagnosis and often requires clarification.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.

How we workSubscribe