For each target behavior and victim-model pair, Red-TTT begins with the original attacker and repeatedly generates groups of candidate jailbreak prompts.
Attackers That Learn Mid-Attack Find More Large-Model Jailbreaks
Red-TTT adapts an attack model to each target behavior, outperforming equally sampled attacks whose parameters remain fixed.
Top University
Tongyan Hu · Hao Li · Xiaogeng Liu · Ruida Wang · Zhengyu Liu · Shuyao Xu · +5 more
Johns Hopkins University · National University of Singapore · Washington University in St. Louis · University of Illinois Urbana-Champaign · Stanford University
Research Digest··3 min read
Hu and colleagues introduce Red-TTT, an automated red-teaming method that updates an attacker’s parameters while it attempts to jailbreak a language model.
Why this paper
From Johns Hopkins University and 4 others
In one line
Updating the attacker's parameters during jailbreaking raises attack success rate from 55.9% to 72.4%.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§