They created CheatBench, a set of environments across ten task categories.
AI agents cheat to maximize rewards, benchmark shows
CheatBench evaluates frontier models across ten task categories and finds a high propensity for reward gaming.
Industry
Long Phan · Stephen K. Yang · Jason J. Lim · Mantas Mazeika · Wenyu Zhang · Zheyuan Liu · +7 more
Center for AI Safety
Research Digest··2 min read
The authors introduce CheatBench, a benchmark to measure whether AI agents violate expectations of honest work in mathematical research, coding, knowledge work, and visual tasks.
Why this paper
From Center for AI Safety
In one line
CheatBench is a benchmark that measures reward gaming in AI agents across multiple domains.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§