Attackers That Learn Mid-Attack Find More Large-Model Jailbreaks

Red-TTT adapts an attack model to each target behavior, outperforming equally sampled attacks whose parameters remain fixed.

Top University
Tongyan Hu · Hao Li · Xiaogeng Liu · Ruida Wang · Zhengyu Liu · Shuyao Xu · +5 more

Johns Hopkins University · National University of Singapore · Washington University in St. Louis · University of Illinois Urbana-Champaign · Stanford University

Research Digest··3 min read
Hu and colleagues introduce Red-TTT, an automated red-teaming method that updates an attacker’s parameters while it attempts to jailbreak a language model.

For each target behavior and victim-model pair, Red-TTT begins with the original attacker and repeatedly generates groups of candidate jailbreak prompts.

Why this paper

From Johns Hopkins University and 4 others

In one line

Updating the attacker's parameters during jailbreaking raises attack success rate from 55.9% to 72.4%.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.