Approximate scores recover context lost by sparse long-context attention

PQ-HSA reuses quantized retrieval scores to represent omitted tokens while reading only a small fraction of the full key-value cache.

Chinese Tech
Kunming Shao · Jierun Chen · Yanli Wang · Ruoyu Wang · Haoli Bai · Kwang-Ting Cheng · +1 more

The Hong Kong University of Science and Technology · Huawei Technologies Ltd. · Sun Yat-sen University · Nanyang Technological University

Research Digest··3 min read
Shao et al.

The authors built an inverted-file product-quantization index over cached keys.

Why this paper

From Huawei Technologies Ltd. and 3 others · Released code

In one line

PQ-HSA preserves accuracy at low retrieval budgets by reusing approximate ranking scores to include unselected KV-cache tokens while accelerating long-context decoding.

What it released

Code

What we could check

  • ✓Code link in the paper (github.com)
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.