The authors identify the 'Gap Curse' in reasoning distillation: as teacher models become more sophisticated, their output distributions increasingly diverge from what student models can approximate, leading to performance degradation.
Teacher alignment method closes capacity gap in reasoning distillation without discarding data
The authors introduce TeacherGRPO, a reinforcement learning approach that adapts teacher models to student distributions, preserving both data coverage and reasoning quality.
Big Tech
Zhenyu Lei · Zihan Chen · Yaochen Zhu · Shangbin Feng · Zaiyi Zheng · Ruocheng Guo · +2 more
University of Virginia · Netflix · University of Washington · Florida State University · Microsoft
Research Digest··3 min read
The authors propose Teacher Alignment to address the Gap Curse in reasoning distillation, where larger teacher models produce distributions too complex for smaller students to approximate.
Why this paper
From Microsoft and 4 others
In one line
TeacherGRPO adapts teacher models to student capacity via reinforcement learning, closing the reasoning distillation gap curse.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§