The authors studied speculative decoding, in which a lightweight model drafts several tokens and a larger target model verifies them together.
Looped diffusion drafting improves decoding without extra model components
D-Loop revises parallel token proposals using a shared drafter, reducing repetition and extending the drafts accepted by larger language models.
Chinese Tech
Kecheng Chen · Yuyang He · Cheng Gong · Hui Liu · Guoping Long · Jiajun Li · +5 more
City University of Hong Kong · The Chinese University of Hong Kong · Huawei Research
Research Digest··3 min read
Chen et al.
Why this paper
From Huawei Research and 2 others
In one line
D-Loop adds intra-block causal conditioning to diffusion drafting, improving acceptance length and decoding speed without extra model components.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§