The authors compared four ways to parallelize attention.
SPLASH switches attention layouts live as language-model workloads change
The system moves running requests among four parallel attention layouts with low handoff overhead, improving throughput as concurrency and context lengths evolve.
Research Lab
Chuan Liu · Shuoming Zhang · Zhicheng Li · Qianqi Sun · Ruiyuan Xu · Qiuchu Yu · +3 more
State Key Laboratory of Processors · Institute of Computing Technology Chinese Academy of Sciences · University of Chinese Academy of Sciences
Research Digest··2 min read
Liu and colleagues present SPLASH, an LLM serving system that changes how attention is distributed across accelerators without draining requests or restarting workers.
Why this paper
From Institute of Computing Technology Chinese Academy of Sciences and 2 others · Released code
In one line
SPLASH switches attention parallel layouts on the fly with under 0.51% overhead, improving LLM serving throughput 1.3-1.73x on B200 GPUs.
What it released
Code
What we could check
- ✓Code link in the paper (github.com)
- ·No weights link found
- ·No dataset link found
- ✓Compute or model size stated (hardware B200 GPUs)
- ✓Limitations stated by the authors
- ✓Reports numbers on named benchmarks (3 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§