The authors built an inverted-file product-quantization index over cached keys.
Approximate scores recover context lost by sparse long-context attention
PQ-HSA reuses quantized retrieval scores to represent omitted tokens while reading only a small fraction of the full key-value cache.
Chinese Tech
Kunming Shao · Jierun Chen · Yanli Wang · Ruoyu Wang · Haoli Bai · Kwang-Ting Cheng · +1 more
The Hong Kong University of Science and Technology · Huawei Technologies Ltd. · Sun Yat-sen University · Nanyang Technological University
Research Digest··3 min read
Shao et al.
Why this paper
From Huawei Technologies Ltd. and 3 others · Released code
In one line
PQ-HSA preserves accuracy at low retrieval budgets by reusing approximate ranking scores to include unselected KV-cache tokens while accelerating long-context decoding.
What it released
Code
What we could check
- ✓Code link in the paper (github.com)
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§