Fixed LLM judges can mismeasure upgrades as agents change

Across coding and customer-service benchmarks, judge errors depended on the agent version being evaluated, sometimes producing upgrade conclusions unsupported by execution-based rewards.

Independent
Jiapeng Li
Research Digest··3 min read
Li tests whether a language-model judge calibrated on one agent version remains reliable when comparing newer versions.

The study evaluated 35 public coding-agent submissions, organized into 20 prespecified version pairs, on 250 SWE-bench Verified issues.

Why this paper

Independent

In one line

Fixed LLM judges make version-dependent errors when evaluating agent upgrades, making paired audits preferable.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors (2 noted)
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.