The authors introduce a framework that separates memory writing from query-adaptive readout for vision-language models (VLMs) processing long videos or continuous streams.
Training-free visual memory framework compresses video tokens while preserving event coherence
KeyRec separates recent fine-grained visual cache from structured event bank, with adaptive readout budget allocation for streaming and long-video understanding.
Big Tech
Zihan Chen · Xuejian Rong · Xiaojuan Wang · Boqing Gong · Adi Zicher · Yael Pritch · +1 more
University of Virginia · Google
Research Digest··2 min read
Thread:Memory Management for Agents
The authors propose KeyRec, a training-free method that builds bounded visual memory by maintaining a dense recent cache and a structured event bank.
Why this paper
From Google and University of Virginia · Part of Memory Management for Agents, now 44 papers
In one line
KeyRec uses a bounded visual memory with a recent cache and event bank to compress video for VLMs, using 10% of the token budget while improving accuracy.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§