Subquadratic Attention
Designing attention mask families and architectures that achieve subquadratic complexity for long-context modeling while preserving expressivity.
7 papers
Where this stands
The written synthesis of this thread is for subscribers. Subscribe.
How this thread developed
Oct 2026 · The Hong Kong University of Science and Technology (Guangzhou), Tencent
Two-sided associative memory correction boosts linear attention performance
Presents Gated Slot Attention-2 (GSA2) which adds two-sided associative memory correction to linear attention, outperforming subquadratic baselines at 1.3B scale.
4 further papers
Oct 2026 · Stanford University
Switching attention layers close gap between linear and softmax
Blends linear and softmax attention states within a single layer to achieve the memory efficiency of subquadratic mechanisms with expressivity closer to full softmax attention.
Oct 2026 · Amazon AGI
Token graph communities speed long-context decoding without model retraining
Introduces a sparse attention mask family that achieves subquadratic complexity via token graph community pruning for long-context generation.
3 of 7 papers shown