The authors ran five models through three pinned production harnesses on 45 tasks from the two hardest SWE-bench Verified difficulty bands.
Coding harnesses change token costs more reliably than success rates
Across three production harnesses, pass-rate differences largely resembled rerun noise, while per-task costs varied by up to threefold.
Academic
Yangze Liu · Zhongyi Han
Shandong University
Research Digest··2 min read
Liu and Han held language models fixed while comparing Claude Code, mini-SWE-agent and OpenCode on SWE-bench Verified, a benchmark of real repository repair tasks.
Why this paper
From Shandong University
In one line
Swapping a coding agent's harness moves pass rate no more than rerunning the same harness.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors (3 noted)
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§