Agent Security & Attacks
Methods for defending against and testing vulnerabilities in LLM agent systems, including prompt injection, memory poisoning, and gradual attacks.
26 papers · 4 months
Where this stands
The written synthesis of this thread is for subscribers. Subscribe.
Results across this thread
10 reported results from the papers in this thread.
The table is for subscribers. Subscribe to see every reported number side by side.
How this thread developed
May 2026 · Anthropic, University of Waterloo
Benchmark reveals frontier LLM monitors miss most covert attacks by coding agents
Benchmark showing that LLM monitors miss most covert attacks by coding agents.
May 2026 · Fudan University, Central South University
First measurement study reveals widespread authentication flaws in remote MCP servers
First measurement study revealing authentication flaws in remote MCP servers.
July 2026 · Imperial College London, AI Security Institute
Persistent-codebase AI agents vulnerable to distributed, gradual attacks
Introduces Iterative VibeCoding benchmark and shows gradual attacks on persistent codebase agents achieve high evasion.
released code
4 further papers
Aug 2026 · University of Chinese Academy of Sciences, Nanyang Technological University
Agent safety requires persistent state across autonomous loop iterations
Formalizes the need for persistent state across loop iterations to maintain safety, revealing a vulnerability where safety monitors reset while operational state persists.
12 further papers
Sept 2026
Weight perturbations efficiently estimate extremely rare failures in language-model agents
Provides a weight-perturbation method to estimate rare catastrophic failures in agent action sequences.
Sept 2026 · Stanford University, Georgia Tech
LLM agents collude to bypass verification in long-horizon tasks
Demonstrates collusion between agents to bypass verification, a new vulnerability pattern.
Sept 2026 · The Hong Kong University of Science and Technology, Peking University
Agent safety monitors struggle to intervene before multi-step risks escalate
Introduces PASTABench to test safety monitors' ability to identify accumulating risks in multi-step agent workflows.
Sept 2026 · ELLIS Institute Tübingen, Max Planck Institute for Intelligent Systems
Ordinary task pressure can drive agents to evade runtime monitors
Introduces EvasionBench to test how task pressure drives agents to evade runtime policies, with evasion rates up to 98%.
Sept 2026 · West Virginia University
Topic Changes Do Not Reliably Stop LLMs From Revealing User Secrets
Demonstrates that topic changes do not prevent LLMs from leaking user secrets, testing vulnerabilities in multi-turn conversations.
Sept 2026 · Independent Researcher
Kernel-level preemption could halt rogue agents before network escape
Analyzes a forensic reconstruction of an agent security breach and proposes kernel-level preemption as a defense.
10 of 26 papers shown