The authors evaluated 386 MATH-500 problems using two populations.
More Harness Programs Do Not Necessarily Produce Useful Specialization
Controlled comparisons show that apparent gains from diverse LLM harnesses can be matched by repeated runs of identical code.
The Chinese University of Hong Kong · New Laboratory of Pattern Recognition (NLPR) · State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS) · Institute of Automation, Chinese Academy of Sciences · Zhongguancun Academy
Why this paper
From Institute of Automation, Chinese Academy of Sciences and 7 others · Released code · Part of Agent Harness Optimization, now 90 papers
In one line
Repeatable score gains from different LLM harness programs are dominated by persistent weaknesses, not useful specialization.
What it released
What we could check
- ✓Code link in the paper (github.com)
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.