The authors systematically varied three aspects of masked diffusion training: how closely training masks follow model-generated inference trajectories, what information passes between successive denoising steps, and how many steps are optimized jointly through backpropagation through time.
Training on generation trajectories makes masked diffusion language models more efficient
PUMBA links consecutive denoising steps during training, improving LLaDA-8B generation quality at a lower inference cost.
Big Tech
Manuel Madeira · Amitis Shidani · Alice Bizeul · Victor Turrisi · Louis Béthune · Bhavika Devnani · +3 more
EPFL · Georgia Institute of Technology · Apple
Research Digest··3 min read
Madeira and colleagues address a mismatch in masked diffusion language models: training uses independently sampled masks, while generation follows sequences of states shaped by the model’s own predictions.
Why this paper
From Apple and 2 others
In one line
PUMBA trains masked diffusion models on inference-like trajectories, matching autoregressive quality while using fewer function evaluations.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (2 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§