SPLASH switches attention layouts live as language-model workloads change

The system moves running requests among four parallel attention layouts with low handoff overhead, improving throughput as concurrency and context lengths evolve.

Research Lab
Chuan Liu · Shuoming Zhang · Zhicheng Li · Qianqi Sun · Ruiyuan Xu · Qiuchu Yu · +3 more

State Key Laboratory of Processors · Institute of Computing Technology Chinese Academy of Sciences · University of Chinese Academy of Sciences

Research Digest··2 min read
Liu and colleagues present SPLASH, an LLM serving system that changes how attention is distributed across accelerators without draining requests or restarting workers.

The authors compared four ways to parallelize attention.

Why this paper

From Institute of Computing Technology Chinese Academy of Sciences and 2 others · Released code

In one line

SPLASH switches attention parallel layouts on the fly with under 0.51% overhead, improving LLM serving throughput 1.3-1.73x on B200 GPUs.

What it released

Code

What we could check

  • ✓Code link in the paper (github.com)
  • ·No weights link found
  • ·No dataset link found
  • ✓Compute or model size stated (hardware B200 GPUs)
  • ✓Limitations stated by the authors
  • ✓Reports numbers on named benchmarks (3 benchmarks)

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.