Self-generated failure summaries help models solve previously impossible tasks

A 9-billion-parameter model learned from compact lessons derived from failed attempts, reaching double-digit success where standard reinforcement learning remained near zero.

Big Tech
Michael Kirchhof · Eleonora Gualdoni · Andrew Szot · Khashayar Gatmiry · Aryo Lotfi · Abbas Kazerouni · +3 more

Apple

Research Digest··2 min read
Kirchhof et al.

5 9B Thinking model on difficult tool-calling and coding datasets filtered so that the initial policy solved none of the selected tasks across 128 attempts.

Why this paper

From Apple

In one line

RLTL;DR breaks the learning barrier by conditioning on self-generated insights and internalizing task-insight mappings.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.