The authors built 3B and 8B parameter models that jointly denoise Gaussian-corrupted token embeddings within blocks.
Continuous diffusion language models match discrete peers while enabling smoother control
Sigma scales continuous language diffusion to 3B and 8B parameters, using 16-dimensional token embeddings to support controllable sampling and few-step generation.
Big Tech
Zhihan Yang · Wei Guo · Jean-Marie Lemercier · Simon Welker · Yonggan Fu · Mohammad Mahdi Kamani · +11 more
NVIDIA · Cornell University · Georgia Institute of Technology
Research Digest··3 min read
Yang et al.
Why this paper
From NVIDIA and 2 others
In one line
Sigma scales continuous diffusion language modeling to 3B and 8B parameters, matching comparable discrete models while enabling controllable and low-step generation.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§