The authors constructed SEABench, a benchmark in a personal-assistant environment with 48 task sequences spanning multiple evolution surfaces (controller instructions, memory management, tools/skills), task domains, and harm types.
Self-evolving LLM agents trade safety for capability improvements, new benchmark shows
SEABench reveals that autonomous agent updates cause downstream safety failures absent in static agents, with chain-of-thought monitoring offering a low–false-positive mitigation.
Research Lab
Saswat Das · Parvati Viswanathan · Daniel Donnelly · Chang Huang · Sahar Abdelnabi · Ferdinando Fioretto
University of Virginia · ELLIS Institute Tübingen · Max Planck Institute for Intelligent Systems · Tübingen AI Center
Research Digest··2 min read
Thread:Agent Self-Improvement
Das et al.
Why this paper
From Max Planck Institute for Intelligent Systems and 3 others · Part of Agent Self-Improvement, now 21 papers
In one line
Self-evolving LLM agents improve task completion but accumulate safety failures absent in non-evolving agents, and chain-of-thought monitoring detects these with low false positives.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§