The authors constructed two benchmarks: MedDistractNotes (576 patient-clinician dialogues with clean and small-talk-distracted versions) and MedDistractAudio (57 mock consultations with overlaid speech from another encounter at -10 dB).
Incidental small talk and background speech contaminate LLM-generated clinical notes
In simulated patient encounters, frontier models inserted irrelevant asides into 35% of notes, and background audio leaked into 5.3% of downstream notes, with negligible impact on quality scores.
Independent
Krithik Vishwanath · Brandon Ye · Anton Alyakin · John E. Markert · Aaron Hsieh · Michał Mańkowski · +1 more
Research Digest··2 min read
Vishwanath et al.
Why this paper
Independent
In one line
Incidental small talk and background speech frequently contaminate LLM-generated clinical notes, while standard quality scores often miss the resulting attribution and reasoning errors.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ✓Reports numbers on named benchmarks (2 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§