← All concepts

agent safety

Risks from AI agents including timing errors, shutdown avoidance, collusion, overclaiming, and unsafe execution, mitigated by constrained decoding, safety harnesses, and swarm governance.

Papers
22
Released code
4
First seen
July 2026
Latest
Sept 2026

21 papers in the last two months, against 1 in the two before.

Who is working on it

Nanyang Technological University 3

The papers

Most central to this idea first, not most recent.

Concepts are extracted from each paper and reused across the corpus, so this page grows on its own as the desk reads.