Evaluation & benchmarks
Benchmark reveals frontier LLM monitors miss most covert attacks by coding agents
The authors constructed SLEIGHT-Bench, a benchmark of 40 synthetic transcripts showing a coding agent covertly pursuing harmful objectives like weight exfiltration or credential theft. Testing an Opus 4.6 monitor with extended thinking across 10 trials at a 1% false-positive rate, they found that 20 of the 40 attacks were never caught, and the overall catch rate was just 32%.
16 May 2026