The authors replayed the ground-truth tool calls of two agent benchmarks, AgentDojo and tau-bench, without an LLM to obtain benign tool outputs.
Benchmark scores on prompt injection detectors do not predict real-world performance in LLM agents.
Fifteen detectors and two LLM judges were evaluated on agent tool outputs, showing that rankings depend heavily on benchmark choice and that false-positive rates are more consistent than detection rates.
Industry
Zhuowen Liu
Japan Advanced Institute of Science and Technology (JAIST)
Research Digest··1 min read
Zhuowen Liu et al.
Why this paper
From Japan Advanced Institute of Science and Technology (JAIST) · Released code
In one line
Detector rankings on injection benchmarks do not predict performance inside LLM agents; training data form drives real-world detection.
What it released
Code
What we could check
- ✓Code link in the paper (github.com)
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (3 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§