The authors first observed that sparse attention mechanisms in LLMs dynamically select which tokens to attend to based on semantic relevance, creating secret-dependent memory accesses in the GPU memory hierarchy.
Sparse attention in LLMs leaks private data via GPU side channels
SparLeak exploits memory access patterns from sparse attention to infer query attributes and reconstruct responses with high accuracy.
Academic
Fahao Chen · Linkang Du · Jinhao Zhou · Peng Li · Zhou Su
Shandong University · Xi'an Jiaotong University · Waseda University
Research Digest··3 min read
Chen et al.
Why this paper
From Shandong University and 2 others
In one line
Sparse attention leaks prompt attributes and generated tokens through GPU cache and TLB contention, enabling attacks with average success rates of 90.9% and 87.3%.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (2 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§