The authors built Spexis on top of vLLM.
Speculative parallelism boosts multi-GPU LLM inference efficiency
Spexis introduces speculative-parallel scheduling that overlaps speculative decoding with normal execution to reduce memory pressure and improve throughput.
Top University
Hyungyu Jung · Jaehyeok Yu · Hoonseo Choi · Sungkyun Kim · Jinho Lee · Jiwon Seo
Seoul National University · Hanyang University
Research Digest··2 min read
Jung et al.
Why this paper
From Seoul National University and Hanyang University
In one line
Spexis accelerates multi-GPU LLM serving by overlapping speculative and normal execution and scheduling around predicted speculation quality and memory pressure, reaching 34% speedups.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ✓Compute or model size stated (gpu NVIDIA A40, RTX PRO 6000, H100, A100, L40S)
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§