Research-guided software evolution improves fixed-model agents on two benchmarks

ScholarEvolve retrieves ideas from research papers, turns them into modular harness changes, and tests combinations without retraining the underlying language model.

Big Tech
Jingbo Yang · Kwei-Herng Lai · Xiaowen Wang · Yaar Harari · Evgeniy Gabrilovich · Shiyu Chang

University of California, Santa Barbara · Microsoft

Research Digest··2 min read
Yang et al.

The authors divided an agent harness into functional modules, then assembled research pools relevant to each module.

Why this paper

From Microsoft and University of California, Santa Barbara · Released code

In one line

ScholarEvolve automatically uses state-of-the-art research and topic modeling to guide agent harness evolution, improving task completion on AppWorld and Tau2-Bench.

What it released

Code

What we could check

  • ✓Code link in the paper (github.com)
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks (2 benchmarks)

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.