Autoregressive video diffusion generates frames by denoising current content while writing key-value history for future predictions.
Separating context-writing and denoising parameters improves autoregressive video generation
SGF+ trains the two roles with independent parameters, improving VBench quality and extending generation from 5-second rollouts to 24 hours.
Top University
Zihan Su · Junhao Zhuang · Yaowei Li · Siwen Lu · Haoran Li · Lingen Li · +7 more
Tsinghua University · JD · The Chinese University of Hong Kong
Research Digest··2 min read
The authors introduce SGF+, which assigns separate parameters to context writing and denoising in autoregressive video diffusion.
Why this paper
From Tsinghua University and 2 others
In one line
Separating parameters for context writing and denoising in autoregressive video generation improves visual quality and long-horizon consistency without extra data or fine-tuning.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§