VSA2 uses a fine-grained router to estimate which key-value tokens matter for each query.
Fine-grained routing makes sparse video attention faster without quality loss
VSA2 adaptively allocates attention across video tokens and can replace full attention during ongoing diffusion-transformer pretraining.
Chinese Tech
Peiyuan Zhang · Guoqiang Wei · Yilong Zhao · Zixiang Zhang · Wei Zhou · Will Lin · +4 more
University of California, San Diego · ByteDance Seed · University of California, Berkeley · Georgia Institute of Technology
Research Digest··2 min read
Zhang et al.
Why this paper
From ByteDance Seed and 3 others
In one line
VSA2 makes video DiT attention 8.9x faster with comparable quality by using a fine-grained router and adaptive sparsity.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§