Agent Self-Improvement
Benchmarks and methods for evaluating and enabling language-model agents to learn from their own past experiences and capability goals without external supervision.
10 papers
Where this stands
The written synthesis of this thread is for subscribers. Subscribe.
How this thread developed
Sept 2026
Models struggle to improve themselves from vague capability goals
ASPIRE benchmark reveals that agents struggle to interpret vague capability goals and improve themselves accordingly.
4 further papers
Sept 2026 · Waseda University, Adelaide University
Co-trained agent roles improve tool-based reasoning and verification
The cooperative framework has agents generate tasks and verifiers, enabling continual self-improvement of tool-based reasoning without external supervision.
Sept 2026 · Fujitsu Limited, The Institute of Statistical Mathematics
Synthetic demonstrations let robot policies escape sparse-reward failures
Uses synthetic demonstrations to escape sparse-reward failures, enabling policy improvement without real-world supervision.
Sept 2026 · Fudan University, Meituan Longcat Team
Trajectory shortcut trees improve agents without outcome labels or annotations
Extracts learning signals from trajectories without outcome labels, enabling self-improvement from unlabeled experience.
Sept 2026 · Zhejiang University, Alibaba Group
Checkpoint testing reveals hidden weaknesses in self-evolving agent memories
Tests self-evolving memories on held-out episodes, enabling agents to learn from past failures.
Sept 2026 · Weco AI
Research agent improves itself through seven successive code rewrites
Demonstrates an autonomous loop where an AI research agent improves its own code through recursive optimization, advancing self-improvement capabilities.
6 of 10 papers shown