The authors begin with on-policy self-distillation, in which a student generates reasoning from a problem while a second copy of the model sees both the problem and a verified solution.
Co-evolving teacher and student models improves mathematical reasoning
A recursive self-distillation method substantially outperformed a frozen-teacher baseline across four competition-level mathematics benchmarks.
Big Tech
Meta AI · University of California, Riverside
Research Digest··2 min read
Yin et al.
Why this paper
From Meta AI and University of California, Riverside
In one line
A recursive distillation method with a co-evolving privileged teacher improves reasoning accuracy more than using a frozen teacher.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§