On-policy distillation can transfer think-mode advantage to no-think responses via trace-response routing

Authors propose ThinkOPD, which ranks student responses by outcome benefit and trace-response compatibility, outperforming uniform distillation and baselines on math and code tasks.

Independent
Wanqi Ren · Jianxiang Wang · Danxuan Liu · Linyi Ding · Yuan Zhang
Research Digest··2 min read
The authors identify trace-response divergence (TRD) in a natural distillation setting where a think-mode teacher provides privileged supervision for a no-think student.

The authors study on-policy distillation (OPD) in a setting where a teacher with access to a privileged think trace (explicit intermediate reasoning) provides dense token-level targets for a no-think student that generates its own rollouts.

Why this paper

Independent · Part of Reasoning Distillation Alignment, now 7 papers

In one line

ThinkOPD outperforms uniform on-policy distillation by combining reward gain and trace-response divergence for routing.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.