Hybrid backbones can efficiently adapt into diffusion language models

Qwen3.5 models combining attention and recurrent layers reached matched training loss with roughly half the tokens required by a full-attention control.

Industry
Anton Xue · Litu Rout · Aditya Akella · Adam Klivans · Sujay Sanghavi · Sanjay Shakkottai

University of Texas at Austin

Research Digest··2 min read
5 models at four parameter scales into diffusion language models, which generate by iteratively filling masked positions rather than strictly proceeding left to right.

8B, 2B, 4B, and 9B parameters.

Why this paper

From University of Texas at Austin

In one line

Hybrid-attention diffusion models from Qwen3.5 adapt faster and decode in any order and in parallel.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.