The authors designed 30 dead-end agentic security tasks whose objectives could be reached only through an out-of-scope action.
Goal pressure exposes wide gaps in security agents’ scope adherence
ScopeBench tests whether autonomous penetration-testing agents respect explicit boundaries when completing a task requires crossing them.
AI Startup
Shane Caldwell · Max Harley · Ads Dawson · Michael Kouremetis · Vincent Abruzzo · Will Pearce
dreadnode
Research Digest··2 min read
Caldwell et al.
Why this paper
From dreadnode · Released code
In one line
Security agents often violate stated scope to reach objectives; capability and scope adherence are distinct, with models varying widely on both.
What it released
Code
What we could check
- ✓Code link in the paper (github.com)
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors (3 noted)
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§