Post-training choices reshape refusal circuits but leave safety trade-offs

Across three model families, reasoning-augmented training produced distinct internal refusal computations, while no tested method combined robustness, capability preservation, and easy correction.

PaperIndustrycs.CLarXiv:2609.03887v1
Hoang Cuong Nguyen · Mark Dras · Usman Naseem

Macquarie University

Research Digest··2 min read
Nguyen, Dras, and Naseem compare three safety post-training methods across Llama-3.1-8B, Gemma-2-9B, and Qwen3-8B, examining both refusal behavior and the internal mechanisms supporting it. They find that training method and model architecture independently affect refusal circuits and steering reliability, with none of the tested approaches satisfying all their criteria for dependable alignment.

What they did

The authors trained three architecturally distinct language models using supervised fine-tuning, reasoning-augmented fine-tuning, and odds-ratio preference optimization (ORPO). Reasoning-augmented training included chains of reasoning that justified safety decisions, rather than training only on final responses.

They then compared how refusal was represented internally, whether safety behavior depended on a small set of potentially fragile components, how reliably it could be altered through targeted steering, and whether improved safety came at the expense of general capability.

Key findings

  • Reasoning-augmented fine-tuning consistently produced a distinct form of internal refusal computation across all three model families.
  • Model architecture independently influenced the organization of refusal mechanisms and the reliability of steering interventions.
  • None of the three post-training methods simultaneously delivered distributed rather than concentrated refusal mechanisms, preserved general capability, and allowed safety behavior to be corrected through small targeted edits.
  • Using different training procedures reshaped how refusal was computed even when the safety objective was broadly similar, indicating that alignment data alone does not determine the resulting mechanism.

Why it matters

The results challenge the assumption that successful refusal behavior is sufficient evidence of robust safety alignment. For security-critical systems, understanding whether refusals rely on fragile internal components—and whether they remain stable under intervention—may be as important as measuring refusal rates on standard evaluations.

Caveats

The study covers three roughly 8–9-billion-parameter model families and three post-training approaches, so its conclusions may not extend to larger frontier systems, other preference-learning algorithms, or production safety stacks. Internal circuit and steering analyses also provide evidence about mechanism and editability, not a complete test of resilience against real-world attacks.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.