Self-play on generated programs produces transferable language-model capabilities

Two models trained from scratch—without natural training data—developed an adaptive synthetic curriculum whose benefits scaled predictably with compute.

Top University
Aditya Cowsik · Kfir Dolev · Michael Y. Li · G. Bruno De Luca · Nourya Cohen · Noah D. Goodman · +1 more

Independent Researcher · Tel Aviv University · Stanford University · LAPTh · USMB

Research Digest··2 min read
The authors pair a generator that writes programs with a learner that predicts the resulting byte sequences.

Starting from random initialization, the authors trained two models jointly.

Why this paper

From Tel Aviv University and 5 others

In one line

A generator and learner, trained only on self-generated program outputs, achieve predictable zero-shot scaling on natural data without ever training on natural data.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.