The authors created 5,486 Clean-Noise question pairs covering six clinical intents and both individual-patient and population-level queries.
Clinical agents falter when health records cannot support the question
EHR-RobustGym tests whether agents detect missing or contradictory evidence instead of producing plausible but unsupported clinical answers.
Chinese Tech
Yitong Qiao · Yancheng Jin · Lei Liu · Yue Shen · Jian Wang · Jinjie Gu · +1 more
Zhejiang University · Ant Healthcare · Ant Group
Research Digest··3 min read
Qiao et al.
Why this paper
From Ant Group and 2 others
In one line
A benchmark of 5,486 noisy clinical queries shows LLM success drops from 62.2% to 37.9% when questions lack supporting EHR evidence.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§