New benchmark tests AI's ability to spot invalid code reviews

CRJudgeBench evaluates whether agents can correctly judge technical trustworthiness of plausible but incorrect review comments.

Big Tech
Yue Pan · Jiawei Li · Ziyuan Zhang · Xiangxin Zhao · He Ye

University College London · Amazon · Zhejiang University

Research Digest··2 min read
The authors introduce CRJudgeBench, a benchmark of 1199 instances built from real pull requests and expert-verified perturbations, assessing whether AI agents can distinguish technically valid code review comments from plausible but invalid ones.

The authors constructed CRJudgeBench from real pull requests (PRs) across diverse repositories, collecting both trustworthy review comments and manually crafted untrustworthy comments that are plausible but contain technically incorrect claims.

Why this paper

From Amazon and 2 others · Released data · Part of Process Agent Benchmarks, now 8 papers

In one line

AI agents can judge if a code review comment is technically correct by gathering repository evidence, outperforming general LLMs.

What it released

Data

What we could check

  • ·No code link found
  • ·No weights link found
  • ✓Dataset link in the paper (huggingface.co)
  • ·No compute details found
  • ✓Limitations stated by the authors (2 noted)
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.