Auditing pending actions catches staged attacks in long-running AI agents

A benchmark of 479 paired executions shows that tracing how external content shapes consequential actions can outperform existing runtime defenses.

Research Lab
Jingkai Liu · Yufei Han · Xiaoting Lyu · Wei Wang · Ting Yu

Beijing Jiaotong University · Mohamed bin Zayed University of Artificial Intelligence · INRIA Rennes-Bretagne-Atlantique · Xi’an Jiaotong University

Research Digest··3 min read
Liu and colleagues demonstrate staged prompt injections that enter an agent through external content, propagate across multiple steps, and later cause harmful actions without preventing task completion.

The authors built a feedback-guided pipeline that designs attacks using a task’s complete benign execution, then reruns the task unscripted with a fixed injection as the only deliberate environmental change.

Why this paper

From INRIA Rennes-Bretagne-Atlantique and 3 others

In one line

Long-horizon agents are vulnerable to staged prompt injection; boundary action auditing with Path-Aligned Attribution blocks attacks at high recall and low false-block rate.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.