5 billion tokens of backbone-generated chain-of-thought data.
Post-training multi-token prediction heads match joint pretraining with far less data
Lightweight post-training on backbone-generated chain-of-thought data enables frozen reasoning models to achieve comparable throughput speedups.
Big Tech
Prachi Badarayani · Aidan Jay · Chenghui Zhou · Dayquan Julienne · Yuan Gao · Tianwei Chen · +5 more
Microsoft
Research Digest··2 min read
5 billion tokens of target-generated chain-of-thought data matches or exceeds the expected speedup of jointly pretrained MTP heads on math, coding, and knowledge benchmarks, using 10^3 to 10^4 times fewer training tokens.
Why this paper
From Microsoft
In one line
Post-trained MTP heads on a frozen model match jointly-pretrained quality with 10^3-10^4x fewer tokens; relaxed verification and adaptive head count add speedups.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (3 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§