Attention distillation complements token-level supervision for reasoning models

OPASD projects privileged teacher attention onto student-visible positions, improving accuracy and efficiency across math benchmarks.

Big Tech
Safaeid Hossain Arib · Rabeya Akter · Ismam Nur Swapnil · Md. Faiyaz Abdullah Sayeedi · Tasnim Mohiuddin · Md Mofijul Islam

ACI PLC · University of Dhaka · BRAC University · QCRI · Amazon GenAI

Research Digest··2 min read
The authors introduce On-Policy Attention Self-Distillation (OPASD), which adds solution-conditioned attention alignment to token-level on-policy self-distillation.

The authors propose OPASD, a method that distills not only the teacher's next-token distribution but also its attention distribution onto the student.

Why this paper

From Amazon GenAI and 4 others

In one line

On-policy attention self-distillation (OPASD) improves reasoning accuracy by 4.98 to 8.40 percentage points while reducing training compute by 72.6% compared to token-only distillation.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks (2 benchmarks)

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.