The authors first characterize what they call in-context evolutionary stagnation, where a fixed language model repeatedly generates and refines algorithms but stops making meaningful progress.
Training on self-generated algorithms helps design agents escape performance plateaus
Population-Curated Policy Optimization retains diverse high-quality solutions and updates model weights so algorithm-design agents can reuse domain knowledge.
Chinese Tech
Chen Lu · Ke Xue · Siyuan Xu · Mingxuan Yuan · Chao Qian
Nanjing University · Huawei Noah’s Ark Lab
Research Digest··2 min read
Lu et al.
Why this paper
From Huawei Noah’s Ark Lab and Nanjing University
In one line
Population-Curated Policy Optimization enables self-evolving LLM agents to escape in-context evolutionary stagnation by learning from self-generated algorithms.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§