The authors formulate policy-space scaling, expanding the set of deployable policies from a fixed RL history without extending training or increasing inference compute.
RL training runs can yield stronger policies than their best checkpoints
The authors demonstrate that combining two checkpoints from the same run via spectral fusion improves accuracy on math and coding benchmarks without extra training or inference compute.
Independent
Bangji Yang · Jiajun Fan · Hongba Ma · Ruihan Guo · Ge Liu
Research Digest··3 min read
The authors show that a completed reinforcement learning (RL) training history contains structure that can be exploited to construct policies stronger than any checkpoint visited during training.
Why this paper
Independent
In one line
An RL training history can yield policies stronger than its best checkpoint without extra training or inference compute.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (2 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§