The authors measured how an agent’s hidden representation changes whenever new context arrives.
Internal representation shifts can expose multi-turn attacks as they emerge
DART tracks changes in an agent’s hidden states, attributes risky shifts to specific context updates and responds with targeted safety reminders.
Top University
Haoyu Wang · Wei Zhao · Yedi Zhang · Christopher M. Poskitt · Jun Sun
Singapore Management University · National University of Singapore
Research Digest··2 min read
Wang and colleagues show that attacks assembled across multiple individually permissible steps leave a measurable trajectory in an LLM agent’s internal representations.
Why this paper
From National University of Singapore and Singapore Management University
In one line
Multi-turn attacks on LLM agents leave detectable signatures in internal representation transitions across context updates.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (2 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§