The authors constructed controlled oracle comparisons at matched attention densities.
Cached token selection narrows diffusion transformers’ sparse attention quality gap
MC-Sparse reuses exact token-level attention choices and residual corrections across denoising steps, accelerating video and 3D generation while closely matching dense attention.
Chinese Tech
Jiarui Chen · Zeqiang Lai · Jiangshan Wang · Ziheng Ouyang · Ye Huang · Xiangyu Yue · +2 more
Fudan University · Tencent HY · Shanghai Innovation Institute · MMLab, CUHK · Nankai University
Research Digest··3 min read
Chen et al.
Why this paper
From Tencent HY and 6 others
In one line
MC-Sparse achieves up to 2.32× denoising speedup over dense attention in diffusion transformers with negligible quality loss.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (2 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§