Reading LLM internal states detects harmful agent trajectories better than content guards

A linear probe on frozen model activations matches or beats full safety fine-tuning on agent safety benchmarks, while using millions fewer parameters and lower latency.

Top University
Difan Jiao · Ashton Anderson

University of Toronto

Research Digest··3 min read
Jiao and Anderson introduce TACIT, a method that detects unsafe tool use and harmful content in language model agent trajectories by reading linear directions in the model's internal representations.

, a payment call that exceeds authorization).

Why this paper

From University of Toronto · Released code

In one line

Reading trajectory safety directly from a frozen LLM's internal states detects harmful agent actions more accurately than open guard models, without generating tokens.

What it released

Code

What we could check

  • ✓Code link in the paper (github.com)
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks (2 benchmarks)

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.