TRACE reduces multi-turn attack success while largely preserving model utility

A token-level safety objective outperformed alternatives across 35 model and attack combinations, with benchmark utility falling by no more than 1.23 points.

Research Lab
Fengpeng Li · Kemou Li · Qizhou Wang · Haiwei Wu · Jiantao Zhou · Di Wang

PRADA Lab · King Abdullah University of Science and Technology · State Key Laboratory of Internet of Things for Smart City · University of Macau · RIKEN Center for Advanced Intelligence Project

Research Digest··3 min read
Li et al.

The authors first develop a theoretical bound describing when suppression of harmful behavior in supervised, single-turn contexts can reduce risk over multi-turn trajectories.

Why this paper

From RIKEN Center for Advanced Intelligence Project and 5 others

In one line

TRACE reduces multi-turn attack success rates across five models and seven attacks, with utility drop at most 1.23 points.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks (2 benchmarks)

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.