The authors target structured pruning of reasoning LLMs followed by a two-stage recovery: offline SFT and on-policy distillation (OPD) where the pruned model generates trajectories and is supervised by the unpruned teacher.
Recovery-aware calibration improves distillation recovery for pruned reasoning models
ReCal reweights pruning criteria using teacher-probe disagreement, boosting mathematical reasoning after on-policy distillation by up to 16.7 points on AIME.
Chinese Tech
Houcheng Jiang · Mao Zheng · Mingyang Song · Qiyong Zhong · Jie Sun · Tianyu Zhang · +1 more
Zhongguancun Academy · Tencent · University of Science and Technology of China · National University of Singapore
Research Digest··3 min read
The authors propose ReCal, a plug-and-play calibration method that adjusts pruning criteria based on forward KL divergence between an unpruned teacher and a pruned probe.
Why this paper
From Tencent and 3 others
In one line
ReCal improves post-pruning on-policy distillation recovery by reweighting calibration statistics based on teacher-probe token-level forward KL.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§