Safety training works better when models ignore superficial prompt cues

DeShortcut-Align targets keyword and formatting shortcuts that cause brittle refusals in large reasoning models.

Chinese Tech
Qirui Liu · Yichen Sun · Yan Wang · Zhixuan Chu · Linbo Jiang · Jianan Lin · +1 more

Zhejiang University · Ant Group · Chongqing Ant Consumer Finance Co., Ltd

Research Digest··2 min read
Liu et al.

The authors first analyzed safety-aligned reasoning models for two failure modes: lexical shortcuts, where sensitive words trigger refusals even in benign contexts, and formatting shortcuts, where safety behavior depends on prompt structures common in training data.

Why this paper

From Ant Group and 2 others

In one line

DeShortcut-Align reduces over-refusal and improves robustness by decoupling spurious shortcuts in safety alignment of reasoning models.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks (2 benchmarks)

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.

Safety training works better when models ignore superficial prompt cues | Zotpaper