, security vs.
Post-training reduces AI models' steerability toward opposing viewpoints
Fine-tuning on one side of cultural-value disagreements makes that side dominant and reduces the model's capacity to enact the other view even when explicitly prompted.
Top University
Jessica Dierking · Itai Shapira · Niclas Boehmer
Hasso Plattner Institute · University of Potsdam · Harvard University
Research Digest··2 min read
Controlled fine-tuning experiments show that standard post-training methods—SFT, RLHF, and DPO—do not merely shift a model's default behavior toward a trained perspective, but also degrade its ability to faithfully adopt an opposing perspective when instructed in-context.
Why this paper
From Harvard University and 2 others
In one line
Post-training narrows models to a single perspective and reduces their ability to adopt opposing views even when instructed.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§