The authors applied On-Policy Context Distillation (OPCD) to train student models (Qwen3-Thinking and Olmo3-Thinking families) on autoformalization tasks, translating natural language statements into first-order logic (FOL).
General instructions outperform gold answers in on-policy context distillation
Using matched formatting instructions as teacher privileges improves out-of-distribution accuracy by 4–17 points while maintaining in-domain performance.
Big Tech
Hantao Yu · Xiaoxue Han · Udaya Ghai · Ferhat Erata · Joseph Lilien · Aman Goel · +1 more
Columbia University · Amazon Web Services
Research Digest··2 min read
Yu et al.
Why this paper
From Amazon Web Services and Columbia University · Part of Context Engineering for Agents, now 43 papers
In one line
Using general instruction privileges instead of instance-specific gold improves out-of-distribution performance in On-Policy Context Distillation.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§