Self-play on generated programs produces transferable language-model capabilities
The authors pair a generator that writes programs with a learner that predicts the resulting byte sequences. Although neither model sees natural data during training, the learner’s zero-shot loss on several natural datasets improves with self-play compute and it develops in-context learning behavior.