Training-free inference framework trims redundant computation and memory in looped Transformers

FlashLoop exploits lazy-update dynamics across loops to deliver up to 1.64x end-to-end speedup and 6x KV-cache reduction without retraining or accuracy loss

Research Lab
Wanqi Yang · Shiwei Liu

ELLIS Institute Tübingen · Max Planck Institute for Intelligent Systems · Tübingen AI Center

Research Digest··1 min read
FlashLoop is a training-free inference framework that reduces cross-loop redundancy in looped Transformers through token-sparse updates, sparse attention, and KV-residual quantization.

The authors analyzed inference dynamics of looped Transformers, where a shared block is applied repeatedly to increase computational depth.

Why this paper

From Max Planck Institute for Intelligent Systems and 2 others · Released code · Part of Agent Harness Optimization, now 90 papers

In one line

FlashLoop achieves up to 1.64x speedup and 6x KV-cache reduction in looped transformers without accuracy loss.

What it released

Code

What we could check

  • ✓Code link in the paper (github.com)
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.