harness evolution
Iteratively adapting an agent's environment, prompts, or evaluation setup to improve task performance and generalization.
- Papers
- 18
- Released code
- 2
- First seen
- June 2026
- Latest
- Sept 2026
14 papers in the last two months, against 4 in the two before.
The papers
Most central to this idea first, not most recent.
- Industrycs.AI
Research agent improves itself through seven successive code rewrites
Weco AI · Sept 2026
- Independentcs.AI
Closed-loop development improves mobile agents across planning and tool use
Sept 2026
- Independentcs.SE
LLMs can build agent harnesses, but improvements transfer poorly
Sept 2026
- Research Labcs.SE
Cross-task failure diagnosis makes LLM agent harness training faster
Chengdu Institute of Computer Applications, Chinese Academy of Sciences, University of Chinese Academy of Sciences · Sept 2026
- Chinese Techcs.SE
Self-evolving harness turns general AI agents into stronger RCA specialists
The Chinese University of Hong Kong, Individual Researcher · Aug 2026
- Top Universitycs.CL
Generating Agent Harnesses on Demand Improves Models Across Benchmarks
LV-NUS Lab · Aug 2026
- Top Universitycs.AI
Evolved agent harnesses lift enterprise performance without retraining models
ServiceNow, Mila · Aug 2026
- Top Universitycs.AI
Behavior-aware checks make agent harness evolution cheaper and more reliable
Fudan University, Shanghai Key Laboratory of Data Science · Aug 2026
- Top Universitycs.CL
Adaptive task selection cuts the cost of optimizing LLM agent harnesses
The University of Tokyo · Aug 2026
- Top Universitycs.AIcode
Re-evaluation shows harness evolution for agents may not outperform simple test-time scaling
Allen Institute for AI, University of Washington · July 2026
- Chinese Techcs.AIcode
Evolutionary training harness co-evolves with LLM policies for RL
Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Tongyi Lab , Alibaba Group · June 2026
- Big Techcs.AI
Self-supervised method improves agent harnesses using past trajectories
City University of Hong Kong, Microsoft Research Asia · June 2026
- Top Universitycs.RO
Visual action rehearsal helps language models control robot manipulation
HKUST(GZ), CUHK · Sept 2026
- Chinese Techcs.AI
Self-improving context programs help frozen models understand long videos
Tencent · Aug 2026
- Chinese Techcs.AI
Restructuring agent environments improves performance on noisy, evolving tasks
Shanghai Jiao Tong University, Theseus Labs · Sept 2026
- Independentcs.CL
Models struggle to improve themselves from vague capability goals
Sept 2026
- Industrycs.AI
Layered agent harnesses improve software over repeated autonomous development cycles
Shanghai Artificial Intelligence Laboratory · Sept 2026
- Big Techcs.AI
Skill evolution improves agent performance in image generation workflows
University of Pennsylvania, Nvidia · July 2026
Concepts are extracted from each paper and reused across the corpus, so this page grows on its own as the desk reads.