The authors propose a theoretical scaling law for reward optimization that jointly accounts for the amount of preference data (M comparisons) and the KL-divergence budget (K).
Joint scaling with data and budget governs reward optimization performance
Aouad et al. derive and empirically confirm that post-training gains scale as the square root of the smaller of logged preference data and divergence budget
Top University
Ali Aouad · Aymane El Gadarri · Vivek F. Farias
MIT
Research Digest··3 min read
The authors develop an information-theoretic model showing that post-training performance scales as Θ(√min(log M, K)), where M is the number of preference comparisons and K is the divergence budget.
Why this paper
From MIT
In one line
Post-training performance scales as Θ(√(min(log M, K))) with data M and divergence budget K.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§