Starting from random initialization, the authors trained two models jointly.
Self-play on generated programs produces transferable language-model capabilities
Two models trained from scratch—without natural training data—developed an adaptive synthetic curriculum whose benefits scaled predictably with compute.
Top University
Aditya Cowsik · Kfir Dolev · Michael Y. Li · G. Bruno De Luca · Nourya Cohen · Noah D. Goodman · +1 more
Independent Researcher · Tel Aviv University · Stanford University · LAPTh · USMB
Research Digest··2 min read
The authors pair a generator that writes programs with a learner that predicts the resulting byte sequences.
Why this paper
From Tel Aviv University and 5 others
In one line
A generator and learner, trained only on self-generated program outputs, achieve predictable zero-shot scaling on natural data without ever training on natural data.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§