Starting from 27,870 prompts from 37 public datasets, the authors filtered to 15,314 English prompts and queried 8 open-source instruction-tuned models.
Local evaluation reveals activation patching can cause benign regressions even when global metrics hold steady
PatchBench curates model-specific jailbreak failures and introduces a protocol to measure harmful correction and benign preservation around each repair.
Academic
Alexi Canesse · Mathis Le Bail · Maël Jenny · Clément Elliker · Mahammed El Sharkawy · Sonia Vanier
LIX (École Polytechnique, IP Paris, CNRS) · AMIAD (Agence Ministérielle pour l’IA de Défense)
Research Digest··2 min read
The authors curate 400 verified jailbreak failures across 8 open-source instruction-tuned models and design a local evaluation protocol that tests both harmful variant correction and benign prompt preservation.
Why this paper
From AMIAD (Agence Ministérielle pour l’IA de Défense) and LIX (École Polytechnique, IP Paris, CNRS)
In one line
Activation patching can suppress jailbreak failures but cause collateral damage to benign prompts undetected by global metrics.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§