Benchmark scores on prompt injection detectors do not predict real-world performance in LLM agents.

Fifteen detectors and two LLM judges were evaluated on agent tool outputs, showing that rankings depend heavily on benchmark choice and that false-positive rates are more consistent than detection rates.

Industry
Zhuowen Liu

Japan Advanced Institute of Science and Technology (JAIST)

Research Digest··1 min read
Zhuowen Liu et al.

The authors replayed the ground-truth tool calls of two agent benchmarks, AgentDojo and tau-bench, without an LLM to obtain benign tool outputs.

Why this paper

From Japan Advanced Institute of Science and Technology (JAIST) · Released code

In one line

Detector rankings on injection benchmarks do not predict performance inside LLM agents; training data form drives real-world detection.

What it released

Code

What we could check

  • ✓Code link in the paper (github.com)
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks (3 benchmarks)

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.