sparse attention
- Papers
- 4
- Released code
- 2
- First seen
- Sept 2026
- Latest
- Oct 2026
4 papers in the last two months, against 0 in the two before.
Who is working on it
ELLIS Institute Tübingen 2Max Planck Institute for Intelligent Systems 2
The papers
Most central to this idea first, not most recent.
- Chinese Techcs.CV
Fine-grained routing makes sparse video attention faster without quality loss
University of California, San Diego, ByteDance Seed · Oct 2026
- Research Labcs.LGcode
Training-free inference framework trims redundant computation and memory in looped Transformers
ELLIS Institute Tübingen, Max Planck Institute for Intelligent Systems · Oct 2026
- Research Labcode
Looped Transformer inference sped up by exploiting sparse cross-loop redundancy
ELLIS Institute Tübingen, Max Planck Institute for Intelligent Systems · Oct 2026
- Industrycs.LG
A Vestigial Attention Branch Signals Which Cache Entries to Keep
Yotta Labs · Sept 2026
Concepts are extracted from each paper and reused across the corpus, so this page grows on its own as the desk reads.