DScale targets wasted computation during speculative decoding, where a small model drafts a block of tokens and a larger model verifies acceptable prefixes.
Adaptive verification speeds block-diffusion decoding under heavy concurrent demand
DScale reallocates limited verification capacity across requests while retaining useful draft tokens and fixed-shape GPU graphs.
Research Lab
Rongjian Chen · Minxian Xu · Zhengxin Fang · Kejiang Ye · Chengzhong Xu
Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences · University of Chinese Academy of Sciences · Victoria University of Wellington · Institute of AI and Brain Sciences · University of Macau
Research Digest··3 min read
Chen and colleagues introduce DScale, a runtime system for speculative decoding that keeps DFlash’s existing block-diffusion drafter and verifies selected candidate prefixes using half the native token capacity.
Why this paper
From Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences and 4 others
In one line
DScale improves throughput of block-diffusion speculative decoding by up to 48.8% over DFlash with only a 112K-parameter predictor.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ✓Compute or model size stated (params Qwen3-8B)
- ✓Limitations stated by the authors
- ✓Reports numbers on named benchmarks (2 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§