Agent safety monitors struggle to intervene before multi-step risks escalate

Across 1,139 agent trajectories, the best of 16 evaluated language models intervened within the optimal window only 40.74% of the time.

Top University
Jiapeng Sun · Yujin Zhou · Han Zhu · Pengcheng Wen · Jiayi Zhou · Sirui Han · +1 more

The Hong Kong University of Science and Technology · Peking University

Research Digest··2 min read
The authors introduce PASTABench to test whether safety monitors can identify accumulating risks and intervene at the right point in a multi-step agent workflow.

The authors formalized proactive safety monitoring as three separate questions: whether intervention is needed, when it should occur, and what risk is present.

Why this paper

From Peking University and The Hong Kong University of Science and Technology · Part of Agent Security & Attacks, now 26 papers

In one line

PASTABench finds LLM agents achieve only 40.74% optimal intervention timing and smaller models rely on keyword matching.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.