The author evaluated decode-time KV-cache compression, where the model’s stored keys and values must be compressed as it generates a long chain of thought.
Low-precision caches preserve more reasoning history within fixed memory budgets
BreadthKV combines quantization and token eviction, using brief end-to-end calibration to choose how many bits each cached token receives.
Academic
Runguo Li
University of Illinois Urbana-Champaign
Research Digest··3 min read
Li studies how long-reasoning models should divide a fixed key-value cache budget between retaining more tokens and storing each token more precisely.
Why this paper
From University of Illinois Urbana-Champaign
In one line
Under fixed KV-cache bytes, retaining more low-precision tokens improves long-chain reasoning more than keeping fewer high-precision tokens through eviction alone.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors (2 noted)
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§