Long-running language-model agents accumulate key-value, or KV, caches that consume increasing GPU memory.
Streaming KV caches makes agent context compaction faster to train
Across text games and software engineering tasks, KV-streams cut training time while preserving the performance of conventional re-prefill compaction.
Big Tech
Emiliano Penaloza · Dane Malenfant · Dheeraj Vattikonda · Roger Creus Castanyer · Siddarth Venkatraman · Abhay Puri · +12 more
Mila · Microsoft · McGill University · Université de Montréal · Polytechnique Montréal
Research Digest··2 min read
Penaloza et al.
Why this paper
From Microsoft and 11 others
In one line
KV-streams streams KV cache across compactions instead of re-prefilling, achieving 2.6x to 5x faster agent training without performance loss.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (3 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§