Real tumor board discussions expose gaps in medical language models

Across 611 cancer cases, leading models only partially reproduced specialists’ answers and multidisciplinary treatment conclusions.

Big Tech
Anqi Li · Zhixuan Ge · Yixuan Duan · Jiarong Qian · Chi-Yu Chen · MingYu Lu · +6 more

Rice University · University of Illinois at Urbana-Champaign · University of Washington · National Yang Ming Chiao Tung University · Microsoft

Research Digest··2 min read
Li and colleagues built OpenTumorBoard from recordings of real multidisciplinary cancer discussions, preserving the sequence of questions, specialist responses and consensus decisions.

The authors transcribed 12,534 minutes of publicly available YouTube tumor board recordings, producing 611 patient cases and 19,157 discussion turns.

Why this paper

From Microsoft and 5 others

In one line

OpenTumorBoard benchmarks LLMs on tumor board discussions; best models score 3.43/5 in clinical equivalence.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks (2 benchmarks)

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.