← All threads

KV Cache Compression

Methods for compressing the key-value cache in transformer inference to reduce memory while preserving reasoning quality.

4 papers

Where this stands

The written synthesis of this thread is for subscribers. Subscribe.

How this thread developed

  1. Oct 2026 · The Hong Kong University of Science and Technology (Guangzhou), Guangdong OPPO Mobile Telecommunications Corp., Ltd.

    Low-rank KV cache compression preserves all tokens during long reasoning

    Proposes iS-KV, an online low-rank KV-cache compression method that preserves recent tokens at full precision and folds older tokens into bounded-rank representations.

  2. Oct 2026 · Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences

    Unified retention and compensation for KV cache eviction via coverage calibration

    Derives an exact factorization of KV cache eviction error into evicted attention mass and directional gap, and proposes CORE for unified retention and compensation.

  3. 1 further paper

    Oct 2026 · Seoul National University, Korea University

    Test-time method prevents drift in long video generation by keeping caching in-distribution

    Identifies KV-provenance drift and proposes in-distribution caching to maintain quality in long video generation.

3 of 4 papers shown