The study evaluated 35 public coding-agent submissions, organized into 20 prespecified version pairs, on 250 SWE-bench Verified issues.
Fixed LLM judges can mismeasure upgrades as agents change
Across coding and customer-service benchmarks, judge errors depended on the agent version being evaluated, sometimes producing upgrade conclusions unsupported by execution-based rewards.
Independent
Jiapeng Li
Research Digest··3 min read
Li tests whether a language-model judge calibrated on one agent version remains reliable when comparing newer versions.
Why this paper
Independent
In one line
Fixed LLM judges make version-dependent errors when evaluating agent upgrades, making paired audits preferable.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors (2 noted)
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§