Generalizing models suffer more collateral damage during machine unlearning

Across synthetic tasks and language-model recall, the authors find that knowledge encoded through shared rules is harder to remove selectively than item-specific memorization.

Top University
Hwiyeong Lee · Hyelim Lim · Ingyu Bang · Hoki Kim · Taeuk Kim

Seoul National University · Hanyang University · Chung-Ang University

Research Digest··3 min read
Lee and colleagues examine how a model’s original learning strategy affects later attempts to remove selected knowledge while preserving everything else.

The authors first used modular addition with grokking, a training phenomenon in which a network shifts from memorizing examples to learning a general rule.

Why this paper

From Seoul National University and 2 others

In one line

A model's reliance on generalization vs memorization predicts retain damage from unlearning: more generalization causes more damage.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.