Local evaluation reveals activation patching can cause benign regressions even when global metrics hold steady

PatchBench curates model-specific jailbreak failures and introduces a protocol to measure harmful correction and benign preservation around each repair.

Academic
Alexi Canesse · Mathis Le Bail · Maël Jenny · Clément Elliker · Mahammed El Sharkawy · Sonia Vanier

LIX (École Polytechnique, IP Paris, CNRS) · AMIAD (Agence Ministérielle pour l’IA de Défense)

Research Digest··2 min read
The authors curate 400 verified jailbreak failures across 8 open-source instruction-tuned models and design a local evaluation protocol that tests both harmful variant correction and benign prompt preservation.

Starting from 27,870 prompts from 37 public datasets, the authors filtered to 15,314 English prompts and queried 8 open-source instruction-tuned models.

Why this paper

From AMIAD (Agence Ministérielle pour l’IA de Défense) and LIX (École Polytechnique, IP Paris, CNRS)

In one line

Activation patching can suppress jailbreak failures but cause collateral damage to benign prompts undetected by global metrics.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.

How we workSubscribe