The authors formalized proactive safety monitoring as three separate questions: whether intervention is needed, when it should occur, and what risk is present.
Agent safety monitors struggle to intervene before multi-step risks escalate
Across 1,139 agent trajectories, the best of 16 evaluated language models intervened within the optimal window only 40.74% of the time.
Top University
Jiapeng Sun · Yujin Zhou · Han Zhu · Pengcheng Wen · Jiayi Zhou · Sirui Han · +1 more
The Hong Kong University of Science and Technology · Peking University
Research Digest··2 min read
Thread:Agent Security & Attacks
The authors introduce PASTABench to test whether safety monitors can identify accumulating risks and intervene at the right point in a multi-step agent workflow.
Why this paper
From Peking University and The Hong Kong University of Science and Technology · Part of Agent Security & Attacks, now 26 papers
In one line
PASTABench finds LLM agents achieve only 40.74% optimal intervention timing and smaller models rely on keyword matching.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§