The authors constructed supervised fine-tuning (SFT) and preference datasets from MedCalc-Bench, a benchmark of 45 medical calculators covering 2,547 clinical scenarios.
Training LLMs to write code boosts medical calculation accuracy by 20-30 points
MedCode framework fine-tunes 7B-8B models to generate executable scripts that offload arithmetic to a deterministic interpreter, achieving 82-94% accuracy on medical benchmarks.
Chinese Tech
Tianshi Ming · Yingying Zhang · Xian Wu
Carnegie Mellon University · Tencent YouTu Lab
Research Digest··3 min read
Ming et al.
Why this paper
From Tencent YouTu Lab and Carnegie Mellon University
In one line
Generating executable code for deterministic calculation improves medical-calculation accuracy in 7B and 8B language models by 20 to 30 percentage points.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§