5 9B Thinking model on difficult tool-calling and coding datasets filtered so that the initial policy solved none of the selected tasks across 128 attempts.
Self-generated failure summaries help models solve previously impossible tasks
A 9-billion-parameter model learned from compact lessons derived from failed attempts, reaching double-digit success where standard reinforcement learning remained near zero.
Big Tech
Michael Kirchhof · Eleonora Gualdoni · Andrew Szot · Khashayar Gatmiry · Aryo Lotfi · Abbas Kazerouni · +3 more
Apple
Research Digest··2 min read
Kirchhof et al.
Why this paper
From Apple
In one line
RLTL;DR breaks the learning barrier by conditioning on self-generated insights and internalizing task-insight mappings.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§