Training-free visual memory framework compresses video tokens while preserving event coherence
The authors propose KeyRec, a training-free method that builds bounded visual memory by maintaining a dense recent cache and a structured event bank. It achieves the best compressed performance in 13 of 15 settings across four benchmarks using only 10% of the decoder-facing visual-token budget, outperforming compressed baselines by 2.21 to 18.37 points on real-time questions.