The authors propose NAMOH (Native Sparse Attention from Mixture-of-Head), where a learned router selects the top-K heads per token from H total heads, with load balancing to distribute tokens evenly.
Routing attention heads by token enables joint scaling of parameters and context
NAMOH activates K of H heads per token, each attending only to its assigned subsequence, reducing KV access while improving quality
Top University
Zizhuo Fu · Runsheng Wang · Meng Li
Peking University
Research Digest··3 min read
Fu et al.
Why this paper
From Peking University
In one line
NAMOH activates K of H attention heads per token, so scaling heads shortens routed contexts and reduces KV access without increasing storage.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§