The authors analyzed 127,833 queries submitted by 6,342 physicians, advanced practice providers and nurses across 35 specialties at a large academic health system.
Clinical AI benchmarks poorly reflect how clinicians actually use language models
Analysis of 127,833 clinical queries found that routine use centers on documentation and information retrieval, tasks largely absent from standard evaluations.
Independent
Krithik Vishwanath · Haitong Lin · Anton Alyakin · Jin Vivian Lee · D. Brock Hewitt · Jie J. Yao · +16 more
Research Digest··3 min read
Vishwanath and colleagues compared eight months of real-world use of an institutional language model assistant with 58 public clinical AI benchmarks.
Why this paper
Independent
In one line
Clinical LLM benchmarks poorly represent deployed use, which centers on documentation and knowledge retrieval rather than diagnosis and often requires clarification.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§