The authors identify two sources of accuracy degradation when naively quantizing recurrent states: error accumulation from repeated quantization at every token, and large outliers in state rows and columns that widen the quantization range.
Training-free quantization halves memory traffic for linear attention states
LeapQuant uses per-window updates and outlier-compensating tokens to achieve 8-bit state compression without accuracy loss.
Big Tech
Yi Pan · Haocheng Xi · Kan Zhu · Xingyang Li · Yibo Wu · Mayank Mishra · +7 more
UC Berkeley · University of Washington · MIT · Perplexity AI · NVIDIA
Research Digest··3 min read
The authors propose LeapQuant, a training-free method for 8-bit quantization of recurrent states in linear attention layers.
Why this paper
From NVIDIA and 4 others
In one line
LeapQuant quantizes recurrent states to 8-bit with near-lossless accuracy, reducing memory and speeding inference in linear attention models.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§