, a payment call that exceeds authorization).
Reading LLM internal states detects harmful agent trajectories better than content guards
A linear probe on frozen model activations matches or beats full safety fine-tuning on agent safety benchmarks, while using millions fewer parameters and lower latency.
Top University
Difan Jiao · Ashton Anderson
University of Toronto
Research Digest··3 min read
Jiao and Anderson introduce TACIT, a method that detects unsafe tool use and harmful content in language model agent trajectories by reading linear directions in the model's internal representations.
Why this paper
From University of Toronto · Released code
In one line
Reading trajectory safety directly from a frozen LLM's internal states detects harmful agent actions more accurately than open guard models, without generating tokens.
What it released
Code
What we could check
- ✓Code link in the paper (github.com)
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (2 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§