The authors constructed CRJudgeBench from real pull requests (PRs) across diverse repositories, collecting both trustworthy review comments and manually crafted untrustworthy comments that are plausible but contain technically incorrect claims.
New benchmark tests AI's ability to spot invalid code reviews
CRJudgeBench evaluates whether agents can correctly judge technical trustworthiness of plausible but incorrect review comments.
Big Tech
Yue Pan · Jiawei Li · Ziyuan Zhang · Xiangxin Zhao · He Ye
University College London · Amazon · Zhejiang University
Research Digest··2 min read
Thread:Process Agent Benchmarks
The authors introduce CRJudgeBench, a benchmark of 1199 instances built from real pull requests and expert-verified perturbations, assessing whether AI agents can distinguish technically valid code review comments from plausible but invalid ones.
Why this paper
From Amazon and 2 others · Released data · Part of Process Agent Benchmarks, now 8 papers
In one line
AI agents can judge if a code review comment is technically correct by gathering repository evidence, outperforming general LLMs.
What it released
Data
What we could check
- ·No code link found
- ·No weights link found
- ✓Dataset link in the paper (huggingface.co)
- ·No compute details found
- ✓Limitations stated by the authors (2 noted)
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§