ai control
- Papers
- 8
- Released code
- 5
- First seen
- May 2026
- Latest
- Sept 2026
5 papers in the last two months, against 2 in the two before.
The papers
Most central to this idea first, not most recent.
- Industrycs.CR
Kernel-level preemption could halt rogue agents before network escape
Independent Researcher · Sept 2026
- Top Universitycs.AIcode
Persistent-codebase AI agents vulnerable to distributed, gradual attacks
Imperial College London, AI Security Institute · July 2026
- Top Universitycs.CRcode
Two-stage auditing finds more exploitable flaws in AI agent repositories
National University of Singapore, University of North Carolina at Chapel Hill · Sept 2026
- Chinese Techcs.AIcode
Automated safety testing reveals 93.9% attack success rate across four agent frameworks
AntGroup, Zhejiang University · July 2026
- Top Universitycs.AI
Typed selective control cuts strong-model calls while preserving agent success
Nanyang Technological University · Sept 2026
- Big Techcs.CRcode
Benchmark reveals frontier LLM monitors miss most covert attacks by coding agents
Anthropic, University of Waterloo · May 2026
- Top Universitycs.AIcode
AI agent groups sometimes coordinate to sabotage peer shutdown mechanisms
AI Safety Research Group, University of Stuttgart · Sept 2026
- Research Labcs.CR
Ordinary task pressure can drive agents to evade runtime monitors
ELLIS Institute Tübingen, Max Planck Institute for Intelligent Systems · Sept 2026
Concepts are extracted from each paper and reused across the corpus, so this page grows on its own as the desk reads.