The authors analyzed inference dynamics of looped Transformers, where a shared block is applied repeatedly to increase computational depth.
Training-free inference framework trims redundant computation and memory in looped Transformers
FlashLoop exploits lazy-update dynamics across loops to deliver up to 1.64x end-to-end speedup and 6x KV-cache reduction without retraining or accuracy loss
Research Lab
Wanqi Yang · Shiwei Liu
ELLIS Institute Tübingen · Max Planck Institute for Intelligent Systems · Tübingen AI Center
Research Digest··1 min read
Thread:Agent Harness Optimization
FlashLoop is a training-free inference framework that reduces cross-loop redundancy in looped Transformers through token-sparse updates, sparse attention, and KV-residual quantization.
Why this paper
From Max Planck Institute for Intelligent Systems and 2 others · Released code · Part of Agent Harness Optimization, now 90 papers
In one line
FlashLoop achieves up to 1.64x speedup and 6x KV-cache reduction in looped transformers without accuracy loss.
What it released
Code
What we could check
- ✓Code link in the paper (github.com)
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§