← All threads

Subquadratic Attention

Designing attention mask families and architectures that achieve subquadratic complexity for long-context modeling while preserving expressivity.

7 papers

Where this stands

The written synthesis of this thread is for subscribers. Subscribe.

How this thread developed

  1. Oct 2026 · The Hong Kong University of Science and Technology (Guangzhou), Tencent

    Two-sided associative memory correction boosts linear attention performance

    Presents Gated Slot Attention-2 (GSA2) which adds two-sided associative memory correction to linear attention, outperforming subquadratic baselines at 1.3B scale.

  2. 4 further papers

    Oct 2026 · Stanford University

    Switching attention layers close gap between linear and softmax

    Blends linear and softmax attention states within a single layer to achieve the memory efficiency of subquadratic mechanisms with expressivity closer to full softmax attention.

  3. Oct 2026 · Amazon AGI

    Token graph communities speed long-context decoding without model retraining

    Introduces a sparse attention mask family that achieves subquadratic complexity via token graph community pruning for long-context generation.

3 of 7 papers shown