The authors first analyze the implicit token-level advantages of forward KL (FKL) and reverse KL (RKL) in self-distillation.
Outcome-guided self-distillation improves LLM reasoning by adapting supervision
The authors propose OG-OPSD, which selects forward or reverse KL divergence based on answer correctness and shortens distillation on incorrect trajectories using teacher entropy, outperforming vanilla on-policy self-distillation.
Chinese Tech
ZheXu Wang · Mao-Lin Luo · Yankun Hong · Zi-Hao Zhou · Bo Ye · Jian Zhao · +3 more
Southeast University · Huawei Noah’s Ark Lab · Zhongguancun Academy · Zhongguancun Institute of Artificial Intelligence
Research Digest··3 min read
On-policy self-distillation (OPSD) provides dense token-level supervision but suffers from noise and instability.
Why this paper
From Huawei Noah’s Ark Lab and 3 others
In one line
Outcome-guided divergence selection and prefix truncation improve on-policy self-distillation for LLM reasoning, consistently beating vanilla OPSD across model scales.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§