The authors profiled sparse-attention prefill and found that adjacent query tokens have strongly correlated access patterns.
Sharing key-value data across queries speeds sparse-attention prefill
QUILT groups neighboring queries so overlapping key-value entries are loaded and dequantized once, reducing long-context inference latency.
Chinese Tech
Zhenduo Zhao · Qihui Zhou · Mingcong Song · Zhiyi Chen · Chuangguan Ye · Fengfan Hou · +4 more
Huawei Technologies Co., Ltd.
Research Digest··3 min read
Zhao et al.
Why this paper
From Huawei Technologies Co., Ltd.
In one line
Neighboring queries in sparse-attention prefill share KV entries; jointly processing them and reusing shared entries reduces kernel latency by up to 55.1%.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§