Training on generation trajectories makes masked diffusion language models more efficient

PUMBA links consecutive denoising steps during training, improving LLaDA-8B generation quality at a lower inference cost.

Big Tech
Manuel Madeira · Amitis Shidani · Alice Bizeul · Victor Turrisi · Louis Béthune · Bhavika Devnani · +3 more

EPFL · Georgia Institute of Technology · Apple

Research Digest··3 min read
Madeira and colleagues address a mismatch in masked diffusion language models: training uses independently sampled masks, while generation follows sequences of states shaped by the model’s own predictions.

The authors systematically varied three aspects of masked diffusion training: how closely training masks follow model-generated inference trajectories, what information passes between successive denoising steps, and how many steps are optimized jointly through backpropagation through time.

Why this paper

From Apple and 2 others

In one line

PUMBA trains masked diffusion models on inference-like trajectories, matching autoregressive quality while using fewer function evaluations.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks (2 benchmarks)

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.