The authors tracked model checkpoints during harmful and safe adaptation of four instruction-tuned, open-weight model families ranging from 3B to 20B parameters.
A harmful gradient subspace helps drive hidden model misalignment
The authors trace emergent misalignment during fine-tuning and reduce it by removing harmful gradient directions from parameter updates.
Chinese Tech
Weiqiao Que · Ruizhe Li · Chengyu Wang · Dakan Wang · Emine Yilmaz · Xiaofeng He
East China Normal University · University of Birmingham · Alibaba Group · Exacity Inc. · University College London
Research Digest··3 min read
Que et al.
Why this paper
From Alibaba Group and 4 others
In one line
Emergent misalignment follows a measurable harmful parameter subspace, and projecting it out of fine-tuning updates suppresses free-generation failures by up to 80%.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§