Clinical agents falter when health records cannot support the question

EHR-RobustGym tests whether agents detect missing or contradictory evidence instead of producing plausible but unsupported clinical answers.

Chinese Tech
Yitong Qiao · Yancheng Jin · Lei Liu · Yue Shen · Jian Wang · Jinjie Gu · +1 more

Zhejiang University · Ant Healthcare · Ant Group

Research Digest··3 min read
Qiao et al.

The authors created 5,486 Clean-Noise question pairs covering six clinical intents and both individual-patient and population-level queries.

Why this paper

From Ant Group and 2 others

In one line

A benchmark of 5,486 noisy clinical queries shows LLM success drops from 62.2% to 37.9% when questions lack supporting EHR evidence.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.

Clinical agents falter when health records cannot support the question | Zotpaper