← All threads

Agent Self-Improvement

Benchmarks and methods for evaluating and enabling language-model agents to learn from their own past experiences and capability goals without external supervision.

10 papers

Where this stands

The written synthesis of this thread is for subscribers. Subscribe.

How this thread developed

  1. Sept 2026

    Models struggle to improve themselves from vague capability goals

    ASPIRE benchmark reveals that agents struggle to interpret vague capability goals and improve themselves accordingly.

  2. 4 further papers

    Sept 2026 · Waseda University, Adelaide University

    Co-trained agent roles improve tool-based reasoning and verification

    The cooperative framework has agents generate tasks and verifiers, enabling continual self-improvement of tool-based reasoning without external supervision.

  3. Sept 2026 · Fujitsu Limited, The Institute of Statistical Mathematics

    Synthetic demonstrations let robot policies escape sparse-reward failures

    Uses synthetic demonstrations to escape sparse-reward failures, enabling policy improvement without real-world supervision.

  4. Sept 2026 · Fudan University, Meituan Longcat Team

    Trajectory shortcut trees improve agents without outcome labels or annotations

    Extracts learning signals from trajectories without outcome labels, enabling self-improvement from unlabeled experience.

  5. Sept 2026 · Zhejiang University, Alibaba Group

    Checkpoint testing reveals hidden weaknesses in self-evolving agent memories

    Tests self-evolving memories on held-out episodes, enabling agents to learn from past failures.

  6. Sept 2026 · Weco AI

    Research agent improves itself through seven successive code rewrites

    Demonstrates an autonomous loop where an AI research agent improves its own code through recursive optimization, advancing self-improvement capabilities.

6 of 10 papers shown