The authors identify a key mismatch in on-policy distillation (OPD) when converting autoregressive (AR) language models into block diffusion language models (dLLMs).
Future-aware distillation corrects teacher-student information mismatch in block diffusion language models
d-OPD modifies on-policy distillation to account for visible future context within each block, improving generation quality and training efficiency.
Big Tech
Ruitao Liu · Qinghao Hu · Song Han
Tsinghua University · MIT · NVIDIA
Research Digest··3 min read
The authors introduce d-OPD, a future-aware on-policy distillation method for converting autoregressive LLMs into block diffusion language models.
Why this paper
From NVIDIA and 2 others · Released code
In one line
Correcting an autoregressive teacher's next-token distribution with visible future context improves distillation of block diffusion language models from pretrained AR models.
What it released
Code
What we could check
- ✓Code link in the paper (github.com)
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§