Streaming KV caches makes compacted agent training up to five times faster

KV-streams avoids repeatedly processing retained context while preserving bounded memory and comparable task performance.

Big Tech
Emiliano Penaloza · Dane Malenfant · Dheeraj Vattikonda · Roger Creus Castanyer · Siddarth Venkatraman · Abhay Puri · +12 more

Mila · Microsoft · McGill University · Université de Montréal · Polytechnique Montréal

Research Digest··2 min read
The authors modify agent training so that key-value caches, the model’s internal attention state, continue across context-compaction events instead of being discarded and rebuilt.

The authors developed KV-streams, an inference and training mechanism that directly evicts selected entries from the attention key-value cache while preserving the remaining cache as one continuous stream.

Why this paper

From Microsoft and 12 others · Released code

In one line

KV-streams streams KV caches forward during compaction, achieving 2.6 to 5x wall-clock training speedup with no performance degradation.

What it released

Code

What we could check

  • ✓Code link in the paper (github.com)
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.