Safety systems that learn from runtime experience protect LLM agents

SafeCoEvo combines rapidly updated safety instructions with a slower-trained guard to improve decisions on unseen tasks.

Chinese Tech
Yu Cheng · Yongkang Hu · Shuaijie Ma · Zhihang Lin · Weicheng Meng · Jingyang Qiao · +9 more

East China Normal University · Shanghai Innovation Institute · Xiamen University · Harbin Institute of Technology · University College London

Research Digest··2 min read
Cheng et al.

The authors model deployment as a stream of tasks in which feedback becomes available only after each task finishes.

Why this paper

From Huawei Noah’s Ark Lab, UK and 7 others

In one line

SafeCoEvo co-evolves an external safety harness and guard at test time, cutting unsafe outcomes by 10.05% and raising task success by 12.15% over the strongest baseline.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors (2 noted)
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.

Safety systems that learn from runtime experience protect LLM agents | Zotpaper