The authors first tested OPD across nine Qwen3 student-teacher pairs, covering parameter-count ratios from 2x to 53x.
Post-training order determines how well language models learn next
Controlled Qwen3 experiments show that supervised warm-up, teacher adaptation and distillation order materially affect later reasoning gains.
Top University
Emre Can Acikgoz · Yang Li · Zeyu Leo Liu · Srijan Bansal · Dilek Hakkani-Tür · Shafiq Joty · +1 more
Salesforce AI Research · UIUC
Research Digest··3 min read
Acikgoz et al.
Why this paper
From UIUC and Salesforce AI Research
In one line
In LLM post-training, a stage that improves the current model can make the next stage less effective.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§