The authors trained models across a grid of parameter counts, epochs and training tokens per parameter.
Repeated training tokens lose value according to a simple scaling rule
Experiments across model sizes and training budgets show when repeated data remains useful, when its value declines, and why dataset ordering also matters.
Independent
Yekun Chai · Haoyi Xiong
Research Digest··3 min read
Chai and Xiong study language-model pretraining when a limited corpus must be reused for multiple epochs.
Why this paper
Independent
In one line
Repeated-token value follows a simple scaling variable, but ordering, allocation, source entropy, and tokenization also determine finite-data pretraining loss.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§