The authors divided an agent harness into functional modules, then assembled research pools relevant to each module.
Research-guided software evolution improves fixed-model agents on two benchmarks
ScholarEvolve retrieves ideas from research papers, turns them into modular harness changes, and tests combinations without retraining the underlying language model.
Big Tech
Jingbo Yang · Kwei-Herng Lai · Xiaowen Wang · Yaar Harari · Evgeniy Gabrilovich · Shiyu Chang
University of California, Santa Barbara · Microsoft
Research Digest··2 min read
Yang et al.
Why this paper
From Microsoft and University of California, Santa Barbara · Released code
In one line
ScholarEvolve automatically uses state-of-the-art research and topic modeling to guide agent harness evolution, improving task completion on AppWorld and Tau2-Bench.
What it released
Code
What we could check
- ✓Code link in the paper (github.com)
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (2 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§