Speculative decoding uses a small model to propose several tokens, then asks the larger target model to verify them in parallel.
Iterative draft editing speeds up speculative language model decoding
DEdit repairs parallel token proposals using later draft tokens as context, increasing acceptance before target-model verification.
Big Tech
Longxuan Yu · Bingsen Chen · Peng Shi · Dongkyu Lee · Yi Xiang · Hideo Kobayashi · +6 more
University of California, Riverside · Amazon Web Services · New York University
Research Digest··2 min read
Yu, Chen and colleagues introduce DEdit, a diffusion-based drafter that revises an entire speculative draft before an autoregressive language model verifies it.
Why this paper
From Amazon Web Services and 2 others
In one line
DEdit improves speculative decoding speed by iteratively editing drafts with bidirectional context, achieving over 5.7x speedup on Qwen3 models.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§