Training-free visual memory framework compresses video tokens while preserving event coherence

KeyRec separates recent fine-grained visual cache from structured event bank, with adaptive readout budget allocation for streaming and long-video understanding.

Big Tech
Zihan Chen · Xuejian Rong · Xiaojuan Wang · Boqing Gong · Adi Zicher · Yael Pritch · +1 more

University of Virginia · Google

Research Digest··2 min read
The authors propose KeyRec, a training-free method that builds bounded visual memory by maintaining a dense recent cache and a structured event bank.

The authors introduce a framework that separates memory writing from query-adaptive readout for vision-language models (VLMs) processing long videos or continuous streams.

Why this paper

From Google and University of Virginia · Part of Memory Management for Agents, now 44 papers

In one line

KeyRec uses a bounded visual memory with a recent cache and event bank to compress video for VLMs, using 10% of the token budget while improving accuracy.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.