5-4B.
Training on self-explanations improves coding agents without reinforcement learning
Retrospection-only fine-tuning transferred lessons from an agent’s explanations into better actions on held-out software-engineering tasks.
Big Tech
Jonathan Light · Christopher Zhang Cui · Jeonghye Kim · Roger Creus Castanyer · Emiliano Penaloza · Zhengyan Shi · +4 more
RPI · UC San Diego · KAIST · Mila · Microsoft Research
Research Digest··2 min read
Thread:Agent Self-Improvement
The authors introduce Retrospection-Only Fine-Tuning, or ROFT, in which an agent attempts a task, explains its experience, then updates its weights using only the explanation tokens.
Why this paper
From Microsoft Research and 4 others · Part of Agent Self-Improvement, now 21 papers
In one line
Fine-tuning an agent on its own self-generated retrospections improves task performance without external rewards or RL.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (2 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§