The authors proposed APEX, a two-level controller for speculative decoding.
Learned controller adapts speculative decoding to maximize speed and reduce waste
APEX selects proposal mechanism per request and adjusts draft depth per block, achieving up to 5.24x speedup and 41% fewer wasted tokens.
Big Tech
Manvi Jha · Zach Zhang · Zhichao Xu · Linbo Liu · Sai Muralidhar Jayanthi · Vinayak Arannil
University of Illinois Urbana-Champaign · AWS AI
Research Digest··3 min read
The authors introduce APEX, a learned controller that dynamically selects the speculative decoding mechanism (EAGLE-3, n-gram, or a draft model) and adapts draft depth at each verification block to balance throughput and wasted computation.
Why this paper
From AWS AI and University of Illinois Urbana-Champaign
In one line
APEX adaptively selects speculation mechanism and draft depth to balance speed and token waste.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§