The authors evaluated retrieval-based factuality checking on the open-ended MedExpert dataset and three closed-ended biomedical datasets.
Retrieval-based checks miss many factual errors in open-ended medical answers
Across multiple retrievers, evidence sources and verifier models, the authors find persistent failures that standard closed-ended benchmarks obscure.
Top University
Center for Language and Speech Processing · Johns Hopkins University
Research Digest··2 min read
Huang et al.
Why this paper
From Johns Hopkins University and Center for Language and Speech Processing
In one line
Retrieval-based factuality evaluation of long-form medical answers fails because scaling retrievers, verifiers, reasoning, and medical tuning does not fix fundamental error modes.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§