Token graph communities speed long-context decoding without model retraining

CommunityKV groups tokens using existing attention scores, then retrieves relevant groups during generation while assigning new tokens at fixed update cost.

Big Tech
Joe McKenna · Anastasios Alexandridis · Nathan Susanj · Jing Liu

Amazon AGI

Research Digest··3 min read
McKenna and colleagues present CommunityKV, a training-free sparse-attention method for reducing the cost of long-context generation.

During prompt processing, CommunityKV converts the model’s already-computed query-key attention scores into a token graph.

Why this paper

From Amazon AGI · Part of Token Pruning Efficiency, now 10 papers

In one line

CommunityKV uses graph partitioning on attention scores to achieve up to 1.71x throughput for long-context decoding.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors (3 noted)
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.