KlinikeBench contains 333 tasks created with contributions from more than 35 clinicians.
Clinical language models diagnose well but often mishandle full encounters
KlinikeBench tests whether models gather necessary history, use clinical tools appropriately and diagnose within an interactive consultation.
Big Tech
Xueting Fang · Zehui Li · Yang Yang · Camilla Giovino · Shubh K. Patel · Shailly Prajapati · +2 more
Zhejiang University · Imperial College London · Nanchang University · University of Toronto · Microsoft
Research Digest··3 min read
Thread:Process Agent Benchmarks
Fang et al.
Why this paper
From Microsoft and 4 others · Part of Process Agent Benchmarks, now 9 papers
In one line
KlinikeBench reveals that top language models succeed on less than 30% of interactive clinical tasks despite 90.7% diagnostic accuracy.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§