The authors formulate token generation as a constrained optimization problem: maximize the probability of accepting the drafter’s output while bounding the divergence between the final and target distributions.
Mentored decoding can trade small distribution shifts for better predictions
The authors connect lossy speculative decoding to boosting theory, showing when a faster composite model can also improve prediction quality.
Big Tech
Vivien Tran-Thien · Richard Nock
Research Digest··3 min read
Thread:Agent Harness Optimization
Tran-Thien and Nock formalize mentored decoding, in which a fast draft model’s tokens are accepted while the resulting distribution remains within a chosen distance of a larger target model.
Why this paper
From Google · Part of Agent Harness Optimization, now 90 papers
In one line
Mentored decoding proves that lossy speculative decoding can beat the target model in quality while speeding inference.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§