Clinical language models diagnose well but often mishandle full encounters

KlinikeBench tests whether models gather necessary history, use clinical tools appropriately and diagnose within an interactive consultation.

Big Tech
Xueting Fang · Zehui Li · Yang Yang · Camilla Giovino · Shubh K. Patel · Shailly Prajapati · +2 more

Zhejiang University · Imperial College London · Nanchang University · University of Toronto · Microsoft

Research Digest··3 min read
Fang et al.

KlinikeBench contains 333 tasks created with contributions from more than 35 clinicians.

Why this paper

From Microsoft and 4 others · Part of Process Agent Benchmarks, now 9 papers

In one line

KlinikeBench reveals that top language models succeed on less than 30% of interactive clinical tasks despite 90.7% diagnostic accuracy.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.