TaSQ modifies the VQ target space through three techniques: query-guided channel weighting to emphasize channels with higher sensitivity to attention logits, cross-head normalization to handle token-level magnitude outliers, and covariance-aware channel grouping to assign dependent channels to the same local VQ codebook.
Tailored vector quantization space enables accurate 1-bit KV cache compression
TaSQ uses query-guided weighting, cross-head normalization, and covariance-aware grouping to improve VQ quality at extreme compression rates.
Top University
Minsoo Cheong · Donghyun Son · Sungjoo Yoo
Seoul National University · Stanford University
Research Digest··3 min read
Thread:KV Cache Compression
The authors introduce TaSQ, a method that tailors the vector quantization (VQ) target space for 1-bit KV cache compression in large language models.
Why this paper
From Seoul National University and Stanford University · Part of KV Cache Compression, now 4 papers
In one line
TaSQ tailors vector quantization for 1-bit KV cache compression, improving accuracy and throughput.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§