The authors developed a KV cache compression framework specifically for LLM agents.
Action-aware cache compression makes long-running LLM agents more efficient
ActKV retains cache entries important for agent actions while adapting memory budgets and integrating compression with paged inference.
Industry
Zihan Wang · Cheng Tang · Lei Gong · Chao Wang · Wenqi Lou · Teng Wang · +1 more
University of Science and Technology of China · Suzhou Institute for Advanced Research, University of Science and Technology of China
Research Digest··2 min read
Wang et al.
Why this paper
From University of Science and Technology of China and Suzhou Institute for Advanced Research, University of Science and Technology of China · Part of Agent Harness Optimization, now 67 papers
In one line
ActKV compresses KV cache for agentic LLMs by retaining entries critical to actions, achieving 98.53% accuracy with 25.98% memory.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§