Post-training reduces AI models' steerability toward opposing viewpoints

Fine-tuning on one side of cultural-value disagreements makes that side dominant and reduces the model's capacity to enact the other view even when explicitly prompted.

Top University
Jessica Dierking · Itai Shapira · Niclas Boehmer

Hasso Plattner Institute · University of Potsdam · Harvard University

Research Digest··2 min read
Controlled fine-tuning experiments show that standard post-training methods—SFT, RLHF, and DPO—do not merely shift a model's default behavior toward a trained perspective, but also degrade its ability to faithfully adopt an opposing perspective when instructed in-context.

, security vs.

Why this paper

From Harvard University and 2 others

In one line

Post-training narrows models to a single perspective and reduces their ability to adopt opposing views even when instructed.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.