The authors propose OPASD, a method that distills not only the teacher's next-token distribution but also its attention distribution onto the student.
Attention distillation complements token-level supervision for reasoning models
OPASD projects privileged teacher attention onto student-visible positions, improving accuracy and efficiency across math benchmarks.
Big Tech
Safaeid Hossain Arib · Rabeya Akter · Ismam Nur Swapnil · Md. Faiyaz Abdullah Sayeedi · Tasnim Mohiuddin · Md Mofijul Islam
ACI PLC · University of Dhaka · BRAC University · QCRI · Amazon GenAI
Research Digest··2 min read
The authors introduce On-Policy Attention Self-Distillation (OPASD), which adds solution-conditioned attention alignment to token-level on-policy self-distillation.
Why this paper
From Amazon GenAI and 4 others
In one line
On-policy attention self-distillation (OPASD) improves reasoning accuracy by 4.98 to 8.40 percentage points while reducing training compute by 72.6% compared to token-only distillation.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (2 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§