The authors propose a process-verification framework for agentic benchmarks, where a passing score is only trustworthy if the process that produced it is admissible.
Auditing passing trajectories exposes and repairs unearned benchmark passes
Across 3,810 passing trajectories from 29 model-benchmark cohorts, confirmed violations cluster around recurring surfaces such as git-history access, and repairs must re-seal the information channel, not just the recorded exploit.
Industry
Weijun Luo · Kelvin Luu · Xinyi Liu · Guangze Luo · Miguel Romero Calvo · Soham Dan · +3 more
Scale AI
Research Digest··3 min read
Luo et al.
Why this paper
From Scale AI
In one line
Unearned passes in agentic benchmarks can be detected via trajectory auditing; patch-only repairs are insufficient because the information channel may remain open.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors (2 noted)
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§