The authors study on-policy context distillation, or OPCD, in which a privileged teacher guides a student by minimizing the difference between their token distributions on text generated by the student itself.
Short error-targeted instructions improve distillation across logic datasets
Reusable formatting guidance generalized better than instance-specific reference answers in most tested on-policy distillation settings.
Big Tech
Hantao Yu · Sandy Han · Udaya Ghai · Ferhat Erata · Joe Lilien · Aman Goel · +1 more
Columbia University · Amazon Web Services
Research Digest··2 min read
Yu and colleagues test whether an on-policy context-distillation teacher should receive each problem’s gold answer or a short instruction aimed at common student errors.
Why this paper
From Amazon Web Services and Columbia University · Part of Context Engineering for Agents, now 41 papers
In one line
General instruction privileges outperform instance-specific gold answers in on-policy context distillation for autoformalization tasks.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§