The authors developed KV-streams, an inference and training mechanism that directly evicts selected entries from the attention key-value cache while preserving the remaining cache as one continuous stream.
Streaming KV caches makes compacted agent training up to five times faster
KV-streams avoids repeatedly processing retained context while preserving bounded memory and comparable task performance.
Big Tech
Emiliano Penaloza · Dane Malenfant · Dheeraj Vattikonda · Roger Creus Castanyer · Siddarth Venkatraman · Abhay Puri · +12 more
Mila · Microsoft · McGill University · Université de Montréal · Polytechnique Montréal
Research Digest··2 min read
The authors modify agent training so that key-value caches, the model’s internal attention state, continue across context-compaction events instead of being discarded and rebuilt.
Why this paper
From Microsoft and 12 others · Released code
In one line
KV-streams streams KV caches forward during compaction, achieving 2.6 to 5x wall-clock training speedup with no performance degradation.
What it released
Code
What we could check
- ✓Code link in the paper (github.com)
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§