The authors fine-tuned models from three families on the same dataset of harmful advice, then evaluated harmful behavior outside the narrow training task.
Some harmful fine-tuning examples drive far more misalignment than others
Filtering examples by estimated influence let the authors strengthen or weaken emergent misalignment, although the rankings varied across models.
Industry
Gonçalo Paulo · Louis Jaburi · Nora Belrose · Lucia Quirke · Stella Biderman
EleutherAI
Research Digest··2 min read
Paulo et al.
Why this paper
From EleutherAI
In one line
Not all harmful fine-tuning examples contribute equally to emergent misalignment; some drive it much more than others.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§