Looped Transformer inference sped up by exploiting sparse cross-loop redundancy

FlashLoop reduces computation and memory by lazily updating only changing tokens, attention, and KV residuals.

Research Lab

ELLIS Institute Tübingen · Max Planck Institute for Intelligent Systems · Tübingen AI Center

Research Digest··3 min read
The authors demonstrate that much of the computation in looped Transformers is redundant: as recurrence progresses, changes become concentrated on a small subset of tokens, attention-output differences are dominated by sparse key columns, and key-value (KV) residuals between adjacent loops become amenable to low-bit quantization.

6B) and identified three forms of redundancy: token-update redundancy (only a small fraction of token rows account for most hidden-state changes), attention-computation redundancy (cross-loop attention output differences are dominated by a sparse, predictable subset of key columns), and KV-storage redundancy (KV residuals between adjacent loops are much more friendly to quantization than full KV states).

Why this paper

From Max Planck Institute for Intelligent Systems and 2 others · Released code

In one line

FlashLoop speeds looped transformer inference by up to 1.64x and cuts KV cache memory by up to 6x without accuracy loss.

What it released

Code

What we could check

  • ✓Code link in the paper (github.com)
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.