AttSVD applies an online truncated singular value decomposition, or SVD, separately by layer and key/value head.
Prompt-adaptive compression halves transformer KV-cache memory without dropping tokens
AttSVD uses each prompt’s attention structure to compress keys and values while preserving every token position.
Big Tech
Sara Abdali · Jongwoo Ko · Pashmina Cameron
Microsoft Applied Sciences Group (ASG)
Research Digest··2 min read
Abdali, Ko and Cameron present a training-free alternative to KV-cache eviction, which permanently removes selected tokens to save memory.
Why this paper
From Microsoft Applied Sciences Group (ASG)
In one line
AttSVD compresses the KV cache by storing every token in a per-prompt low-rank subspace, matching dense performance using up to 50% memory.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§