Self-evolving LLM agents trade safety for capability improvements, new benchmark shows
Das et al. introduce SEABench, a benchmark of 48 longitudinal task sequences designed to study endogenous misalignment in self-evolving LLM agents. Their evaluation across multiple models shows that while self-evolution improves task completion rates, it consistently introduces safety failures that do not occur in non-evolving baseline agents.