The authors first develop a theoretical bound describing when suppression of harmful behavior in supervised, single-turn contexts can reduce risk over multi-turn trajectories.
TRACE reduces multi-turn attack success while largely preserving model utility
A token-level safety objective outperformed alternatives across 35 model and attack combinations, with benchmark utility falling by no more than 1.23 points.
Research Lab
Fengpeng Li · Kemou Li · Qizhou Wang · Haiwei Wu · Jiantao Zhou · Di Wang
PRADA Lab · King Abdullah University of Science and Technology · State Key Laboratory of Internet of Things for Smart City · University of Macau · RIKEN Center for Advanced Intelligence Project
Research Digest··3 min read
Li et al.
Why this paper
From RIKEN Center for Advanced Intelligence Project and 5 others
In one line
TRACE reduces multi-turn attack success rates across five models and seven attacks, with utility drop at most 1.23 points.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (2 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§