The authors study on-policy distillation (OPD) in a setting where a teacher with access to a privileged think trace (explicit intermediate reasoning) provides dense token-level targets for a no-think student that generates its own rollouts.
On-policy distillation can transfer think-mode advantage to no-think responses via trace-response routing
Authors propose ThinkOPD, which ranks student responses by outcome benefit and trace-response compatibility, outperforming uniform distillation and baselines on math and code tasks.
Independent
Wanqi Ren · Jianxiang Wang · Danxuan Liu · Linyi Ding · Yuan Zhang
Research Digest··2 min read
The authors identify trace-response divergence (TRD) in a natural distillation setting where a think-mode teacher provides privileged supervision for a no-think student.
Why this paper
Independent · Part of Reasoning Distillation Alignment, now 7 papers
In one line
ThinkOPD outperforms uniform on-policy distillation by combining reward gain and trace-response divergence for routing.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§