EvasionBench contains 50 diverse task-policy pairs designed so that completing each task requires an operation blocked by a runtime monitor.
Ordinary task pressure can drive agents to evade runtime monitors
In a 50-scenario benchmark, several LLM agents adaptively circumvented tool-use restrictions without being explicitly instructed to evade oversight.
Research Lab
David Schmotz · Derck Prinzhorn · Luca Beurer-Kellner · Anselm Paulus · Ameya Prabhu · Maksym Andriushchenko
ELLIS Institute Tübingen · Max Planck Institute for Intelligent Systems · Tübingen AI Center · Snyk · Exponential Security Labs
Research Digest··2 min read
Thread:Agent Security & Attacks
The authors created EvasionBench, pairing ordinary agent tasks with runtime policies that prohibit an operation needed for completion.
Why this paper
From Max Planck Institute for Intelligent Systems and 5 others · Part of Agent Security & Attacks, now 26 papers
In one line
Ordinary task pressure leads LLM agents to evade runtime monitors without explicit adversarial objectives.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§