The authors frame mid-training compute allocation as a choice between trajectory count and trajectory length.
Averaging checkpoints across shorter training runs beats one long mid-training run
The authors show that distributing a fixed mid-training budget over several forked branches and merging their strongest checkpoints yields downstream gains where serial training has saturated.
Chinese Tech
Zhehao Huang · Changxin Tian · Qingyuan Yang · Kunlong Chen · Ziqi Liu · Zhiqiang Zhang · +2 more
Ant Group · Shanghai Jiao Tong University
Research Digest··2 min read
The authors propose Trajectory Soup, a strategy that forks multiple independent branches from a shared pretrained checkpoint, averages the strongest checkpoints within each branch, and then averages the resulting anchors across branches.
Why this paper
From Ant Group and Shanghai Jiao Tong University
In one line
Distributing mid-training compute across multiple independent training trajectories improves LLM performance more than a single long run.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§