The authors study distributional distillation for continuous diffusion language models (CDLMs), where a pretrained teacher supervises a student to generate samples in fewer steps by matching the noised student and data distributions under a reverse-KL objective.
Distributional distillation cuts diffusion language model sampling steps by half or more
Two new methods, Simplex-DMD and Reinforce-DMD, achieve competitive perplexity with just 4 to 256 network evaluations on 1,024-token sequences.
AI Startup
Paul Le Van Kiem · Dario Shariatian · Umut Simsekli · Alain Durmus
Inria · PSL Research University · Cohere · Ecole Polytechnique
Research Digest··3 min read
The authors propose two distributional distillation methods for continuous diffusion language models that significantly reduce the number of network evaluations (NFEs) needed for high-quality generation.
Why this paper
From Cohere and 3 others
In one line
Two distribution matching distillation methods reduce network evaluations for continuous diffusion language models by up to 49% and 20% at similar generation quality.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§