Training on self-explanations improves coding agents without reinforcement learning

Retrospection-only fine-tuning transferred lessons from an agent’s explanations into better actions on held-out software-engineering tasks.

Big Tech
Jonathan Light · Christopher Zhang Cui · Jeonghye Kim · Roger Creus Castanyer · Emiliano Penaloza · Zhengyan Shi · +4 more

RPI · UC San Diego · KAIST · Mila · Microsoft Research

Research Digest··2 min read
The authors introduce Retrospection-Only Fine-Tuning, or ROFT, in which an agent attempts a task, explains its experience, then updates its weights using only the explanation tokens.

5-4B.

Why this paper

From Microsoft Research and 4 others · Part of Agent Self-Improvement, now 21 papers

In one line

Fine-tuning an agent on its own self-generated retrospections improves task performance without external rewards or RL.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks (2 benchmarks)

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.