Method to permanently erase unwanted LLM personas via weight edits

A new benchmark and contrastive erasure technique suppress target personas while preserving utility across six models and five personas.

Big Tech
Kemou Li · Zhuan Shi · Qizhou Wang · Fengpeng Li · Negar Rostamzadeh · Golnoosh Farnadi · +1 more

University of Macau · Mila – Québec AI Institute · McGill University · RIKEN AIP · King Abdullah University of Science and Technology

Research Digest··3 min read
Li et al.

The authors formalize persona unlearning as a distinct behavior-level editing problem.

Why this paper

From Google Research and 5 others

In one line

PaCE suppresses LLM personas like sycophancy by editing weights to make them hard to elicit.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.