Sparse attention in LLMs leaks private data via GPU side channels

SparLeak exploits memory access patterns from sparse attention to infer query attributes and reconstruct responses with high accuracy.

Academic
Fahao Chen · Linkang Du · Jinhao Zhou · Peng Li · Zhou Su

Shandong University · Xi'an Jiaotong University · Waseda University

Research Digest··3 min read
Chen et al.

The authors first observed that sparse attention mechanisms in LLMs dynamically select which tokens to attend to based on semantic relevance, creating secret-dependent memory accesses in the GPU memory hierarchy.

Why this paper

From Shandong University and 2 others

In one line

Sparse attention leaks prompt attributes and generated tokens through GPU cache and TLB contention, enabling attacks with average success rates of 90.9% and 87.3%.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks (2 benchmarks)

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.