The authors derived stage-level labels from the LLMail-Inject benchmark, covering retrieval, defense bypass, tool invocation, tool-argument correctness and end-to-end attack success.
Modeling attack chains improves prompt injection detection for email agents
A staged detector outperformed frozen text classifiers by checking how retrieved instructions conflict with user intent and proposed tool actions.
Academic
Eastern Michigan University
Research Digest··2 min read
Hashmi, Patel and Yin model indirect prompt injection as a sequence from malicious email retrieval to unsafe tool arguments, rather than as binary text classification.
Why this paper
From Eastern Michigan University · Released code
In one line
Modeling prompt injection as an attack chain with stage verifiers and intent consistency analysis improves detection F1 over binary classifiers.
What it released
Code
What we could check
- ✓Code link in the paper (github.com)
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (2 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§