The authors first analyzed safety-aligned reasoning models for two failure modes: lexical shortcuts, where sensitive words trigger refusals even in benign contexts, and formatting shortcuts, where safety behavior depends on prompt structures common in training data.
Safety training works better when models ignore superficial prompt cues
DeShortcut-Align targets keyword and formatting shortcuts that cause brittle refusals in large reasoning models.
Chinese Tech
Qirui Liu · Yichen Sun · Yan Wang · Zhixuan Chu · Linbo Jiang · Jianan Lin · +1 more
Zhejiang University · Ant Group · Chongqing Ant Consumer Finance Co., Ltd
Research Digest··2 min read
Liu et al.
Why this paper
From Ant Group and 2 others
In one line
DeShortcut-Align reduces over-refusal and improves robustness by decoupling spurious shortcuts in safety alignment of reasoning models.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (2 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§