Fine-grained routing makes sparse video attention faster without quality loss

VSA2 adaptively allocates attention across video tokens and can replace full attention during ongoing diffusion-transformer pretraining.

Chinese Tech
Peiyuan Zhang · Guoqiang Wei · Yilong Zhao · Zixiang Zhang · Wei Zhou · Will Lin · +4 more

University of California, San Diego · ByteDance Seed · University of California, Berkeley · Georgia Institute of Technology

Research Digest··2 min read
Zhang et al.

VSA2 uses a fine-grained router to estimate which key-value tokens matter for each query.

Why this paper

From ByteDance Seed and 3 others

In one line

VSA2 makes video DiT attention 8.9x faster with comparable quality by using a fine-grained router and adaptive sparsity.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.