The authors formalize persona unlearning as a distinct behavior-level editing problem.
Method to permanently erase unwanted LLM personas via weight edits
A new benchmark and contrastive erasure technique suppress target personas while preserving utility across six models and five personas.
Big Tech
Kemou Li · Zhuan Shi · Qizhou Wang · Fengpeng Li · Negar Rostamzadeh · Golnoosh Farnadi · +1 more
University of Macau · Mila – Québec AI Institute · McGill University · RIKEN AIP · King Abdullah University of Science and Technology
Research Digest··3 min read
Li et al.
Why this paper
From Google Research and 5 others
In one line
PaCE suppresses LLM personas like sycophancy by editing weights to make them hard to elicit.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§