The authors study compressed key-value caches, compact representations of a document's transformer states that can be computed once and reused for many queries.
Compressed LLM memories can preserve skills beyond their source documents
The authors show that learned KV-cache compression can impair unrelated queries, then mitigate the problem through data mixing or relevance-based routing.
Big Tech
Sonia Laguna · Joao Monteiro · Marco Cuturi · Pierre Ablin · Eleonora Gualdoni
Apple
Research Digest··2 min read
Laguna and colleagues test whether compressed key-value caches remain safe to reuse when a user asks questions unrelated to the cached document.
Why this paper
From Apple
In one line
Cartridges++ retains off-context abilities in compressed KV caches with negligible cost.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§