Post-training multi-token prediction heads match joint pretraining with far less data

Lightweight post-training on backbone-generated chain-of-thought data enables frozen reasoning models to achieve comparable throughput speedups.

Big Tech
Prachi Badarayani · Aidan Jay · Chenghui Zhou · Dayquan Julienne · Yuan Gao · Tianwei Chen · +5 more

Microsoft

Research Digest··2 min read
5 billion tokens of target-generated chain-of-thought data matches or exceeds the expected speedup of jointly pretrained MTP heads on math, coding, and knowledge benchmarks, using 10^3 to 10^4 times fewer training tokens.

5 billion tokens of backbone-generated chain-of-thought data.

Why this paper

From Microsoft

In one line

Post-trained MTP heads on a frozen model match jointly-pretrained quality with 10^3-10^4x fewer tokens; relaxed verification and adaptive head count add speedups.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks (3 benchmarks)

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.