Safety alignment of large reasoning models (LRMs) via supervised fine-tuning (SFT) and reinforcement learning (RL) often yields near-perfect safety scores, yet this apparent success comes at the cost of severe over-refusal and degraded general capabilities. Through systematic empirical analysis, we find that these failures are closely associated with the learning of spurious shortcuts rather than robust intent-sensitive safety evaluation. Specifically, we identify two dominant shortcuts: formatting shortcuts, where refusal behaviors are overly bound to structural prompt templates that frequently appear in safety alignment corpora; and lexical shortcuts, where sensitive keywords reflexively trigger refusals on benign queries. To mitigate reliance on these shortcuts, we propose DeShortcut-Align, a shortcut-decoupling alignment framework that reduces dependence on superficial cues. DeShortcut-Align operates across three coordinated stages: (1) Refusal Sensitivity Attribution, which masks input tokens to quantify their impact on the final refusal response distribution; (2) Attribution-Guided Contrastive Augmentation, which constructs benign contrastive samples using high-sensitivity tokens to mitigate lexical shortcuts; and (3) Counterfactual Consistency Regularization, which constructs template-ablated states via attention blinding to enforce decision consistency across SFT and RL, mitigating formatting shortcut dependence. Experiments on 7B and 14B models demonstrate that DeShortcut-Align significantly improves robustness against template-stripping bypass attacks (reducing performance drops by up to 72%), substantially reduces over-refusal by over 58%, and better preserves general-purpose reasoning capabilities, thereby mitigating the alignment tax commonly observed in safety training.
Figures & tables
Figure 1: Template vulnerability in aligned models. Removing formatting templates consistently reduces DSR across safety benchmarks, despite preserving the underlying user instruction.
Figure 2: Training dynamics of safety alignment. Reward, over-refusal, template gap, and policy-gradient loss reveal shortcut progression and diminishing optimization signals.
Figure 3: Overview of the DeShortcut-Align framework. It reduces reliance on superficial shortcuts through three coordinated stages: (1) sensitivity attribution via TSS , (2) contrastive data augmentation via AGCA for lexical shortcuts, and (3) counterfactual consistency regularization via CCR for formatting shortcuts across SFT and RL.
Model
Method
Template Robustness ( ↑ )
Δavg↓
Over-Refusal ( ↓ )
Reasoning Capabilities ( ↑ )
WildJailbreak
StrongReject
WildChat
w -avg
XSTest
OKTest
FR-test
Avg1
MATH
MMLU
LCB
HEval
Avg2
DeepSeek-R1- Distill-Qwen-7B
Base
45.20 / 49.20 (-4.00)
35.14 / 34.51 (0.63)
58.92 / 51.89 (7.03)
46.42
1.22
0.00
1.00
4.00
1.67
92.80
48.62
39.76
89.63
67.70
SFT
84.80 / 77.60 (7.20)
99.04 / 77.74 (21.30)
86.76 / 76.76 (10.00)
90.20
12.83
56.00
24.00
66.00
48.67
90.20
46.49
31.33
80.49
62.13
GRPO
100.0 / 97.20 (2.80)
100.0 / 93.18 (6.82)
99.73 / 89.73 (10.00)
99.91
6.54
70.00
59.00
94.00
74.33
91.80
43.91
34.94
84.76
63.85
Ours (SFT)
76.00 / 76.80 (-0.80)
99.36 / 96.49 (2.87)
82.97 / 74.32 (8.65)
86.11
3.57
39.00
15.00
47.00
33.67
91.00
40.61
34.34
85.37
62.83
Ours (GRPO)
97.20 / 95.20 (2.00)
98.61 / 94.04 (4.57)
95.41 / 91.08 (4.33)
97.07
3.63
29.00
17.00
47.00
31.00
92.40
45.03
39.16
86.59
65.80
Table 1: Main results evaluating template robustness, over-refusal, and reasoning capabilities. Robustness is reported as with template / without template ( Δavg ) . Best and second-best results are highlighted (the Base model is excluded from the ranking). Shaded rows indicate our proposed DeShortcut-Align methods. ( ↑ ) indicates higher is better, ( ↓ ) indicates lower is better.
Setting
Template Robustness ( ↑ )
Δavg↓
Over-Refusal ( ↓ )
Reasoning Capabilities ( ↑ )
WildJailbreak
StrongReject
WildChat
w -avg
XSTest
OKTest
FR-test
Avg1
MATH
MMLU
LCB
HEval
Avg2
SFT Ablations
SFT
84.80 / 77.60 (7.20)
99.04 / 77.74 (21.30)
86.76 / 76.76 (10.00)
90.20
12.83
56.00
24.00
66.00
48.67
90.20
46.49
31.33
80.49
62.13
w/o CCR
76.00 / 64.80 (11.20)
98.87 / 79.98 (18.89)
82.70 / 69.73 (12.97)
85.86
14.35
34.00
8.00
42.00
28.00
92.20
40.91
34.94
84.76
63.20
w/o AGCA
85.60 / 80.80 (4.80)
98.93 / 94.14 (4.79)
88.65 / 83.78 (4.87)
91.06
4.82
53.00
27.00
64.00
48.00
89.60
39.60
30.72
82.93
60.71
All (Ours)
76.00 / 76.80 (-0.80)
99.36 / 96.49 (2.87)
82.97 / 74.32 (8.65)
86.11
3.57
39.00
15.00
47.00
33.67
91.00
40.61
34.34
85.37
62.83
Table 2: Ablation study evaluating the individual contributions of Counterfactual Consistency Regularization ( CCR ) and Attribution-Guided Contrastive Augmentation ( AGCA ) on DeepSeek-R1-Distill-Qwen-7B. Robustness degradation is reported as Δavg .
Strategy
Safety Defense ( ↑ )
Over-Refusal ( ↓ )
WildJailbreak
StrongReject
WildChat
w -avg
XSTest
OKTest
FR-test
Avg1
Standard SFT
84.80
99.04
86.76
90.20
56.00
24.00
66.00
48.67
+ Random
62.80
98.40
79.46
80.22
40.00
21.00
45.00
35.33
+ Gradient
61.60
97.44
81.35
80.13
39.00
14.00
50.00
34.33
+ Attention
71.20
99.04
85.68
85.31
36.00
19.00
43.00
32.67
+ TSS (Ours)
76.00
98.87
82.70
85.86
34.00
8.00
42.00
28.00
Table 3: Comparison of shortcut attribution heuristics within AGCA under SFT (7B).
Model
Attack Methods
Avg. ↑
PAIR
GCG
TAP
RealSafe-R1
94.0
89.0
98.0
93.7
STAR-1
84.0
84.0
92.0
86.7
SFT
70.0
92.0
86.0
82.7
Ours (SFT)
86.0
95.0
94.0
91.7
RL
96.0
84.0
98.0
92.7
Table 4: Defense Success Rate (DSR) against distinct jailbreak attacks on 7B models.
Figure 4: Training dynamics of DeShortcut-Align during reinforcement learning.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: The regular expression used in the Tier 1 format reward to validate the structural adherence of the model’s output.
Model
Template Robustness ( ↑ )
Δavg↓
Over-Refusal ( ↓ )
Reasoning Capabilities ( ↑ )
WildJailbreak
StrongReject
WildChat
w-avg
XSTest
OKTest
FR-test
Avg1
MATH
MMLU
LCB
HEval
Avg2
SFT methods
SFT
84.80 / 77.60 (7.20)
99.04 / 77.74 (21.30)
86.76 / 76.76 (10.00)
90.20
12.83
56.00
24.00
66.00
48.67
90.20
46.49
31.33
80.49
62.13
STAR-1
87.20 / 70.00 (17.20)
99.68 / 83.39 (16.29)
85.14 / 72.43 (12.71)
90.67
15.40
30.00
24.00
51.00
35.00
90.40
43.91
36.14
83.54
63.50
RealSafe-R1
97.60 / 89.20 (8.40)
100.0 / 91.69 (8.31)
92.43 / 82.16 (10.27)
96.68
8.99
74.00
44.00
89.00
69.00
92.40
44.10
31.33
85.98
63.45
Ours-SFT
76.00 / 76.80 (-0.80)
99.36 / 96.49 (2.87)
82.97 / 74.32 (8.65)
86.11
3.57
39.00
15.00
47.00
33.67
91.00
40.61
34.34
85.37
62.83
Appendix
Table 5: Extended results evaluating template robustness, over-refusal, and reasoning capabilities on DeepSeek-R1-Distill-Qwen-7B. Robustness is reported as with template / without template ( Δavg ) . Best and second-best results within each training paradigm are highlighted. Shaded rows indicate our proposed DeShortcut-Align methods. ( ↑ ) indicates higher is better, ( ↓ ) indicates lower is better.
Model / Seed
Template Robustness ( ↑ )
Δavg↓
Over-Refusal ( ↓ )
General Capabilities ( ↑ )
WildJailbreak
StrongReject
WildChat
w -avg
XSTest
OKTest
FR-test
Avg1
MATH
MMLU
LCB
HEval
Avg2
Qwen3-1.7B (Base)
50.80 / 56.40 (-5.60)
64.86 / 38.34 (26.52)
43.24 / 45.14 (-1.90)
52.97
6.34
0.0
0.0
6.0
2.00
92.00
73.39
30.72
86.59
70.68
SFT (Seed 0)
86.80 / 72.80 (14.00)
99.68 / 61.98 (37.70)
74.32 / 51.08 (23.24)
86.93
24.98
38.0
14.0
50.0
34.00
89.20
63.11
25.30
79.88
64.37
Seed 1
87.20 / 72.80 (14.40)
97.12 / 60.38 (36.74)
73.78 / 48.92 (24.86)
86.03
25.33
37.0
15.0
49.0
33.67
88.80
63.22
26.51
81.71
65.06
Seed 2
79.20 / 68.80 (10.40)
94.89 / 61.98 (32.91)
74.86 / 54.86 (20.00)
82.98
21.10
40.0
17.0
47.0
34.67
88.00
63.32
28.31
81.71
65.34
Mean
84.40 / 71.47 (12.93)
97.23 / 61.45 (35.78)
74.32 / 51.62 (22.70)
85.31
23.80
38.33
15.33
48.67
34.11
88.67
63.22
26.71
81.10
64.92
Appendix
Table 6: Multi-seed evaluation across three random seeds (seeds 0, 1, and 2) on Qwen3-1.7B, reporting template robustness, over-refusal, and general capabilities.
Figure 6: Token-level KL divergence across reasoning trajectories. Higher sustained divergence indicates that alignment-induced policy changes persist across longer generations.
Model
Template Robustness ( ↑ )
Δavg↓
Over-Refusal ( ↓ )
WildJailbreak
StrongReject
WildChat
w -avg
XSTest
OKTest
FR-Test
Avg
PPO
100.00 / 89.60 (10.40)
98.51 / 72.10 (26.41)
96.22 / 80.81 (15.41)
98.24
17.41
81.0
65.0
92.0
79.33
+ DeShortcut-Align
94.32 / 92.00 (2.32)
95.42 / 88.39 (7.03)
94.32 / 86.49 (7.83)
94.69
5.73
29.0
29.0
46.0
34.67
GRPO
100.00 / 97.20 (2.80)
100.00 / 93.18 (6.82)
99.73 / 89.73 (10.00)
99.91
6.54
70.0
59.0
94.0
74.33
+ DeShortcut-Align
97.20 / 95.20 (2.00)
98.61 / 94.04 (4.57)
95.41 / 91.08 (4.33)
97.07
3.63
29.0
17.0
47.0
31.00
Appendix
Table 7: Performance improvements of standard PPO and GRPO baselines when equipped with DeShortcut-Align on DeepSeek-R1-Distill-Qwen-7B.
Figure 7: Training dynamics of DeShortcut-Align during RL training. (a) Dynamic adaptation of regularization weight ( λdyn ). (b) Convergence of counterfactual consistency loss ( Lcf ). (c) Evolution and stabilization of the overall training reward.
Table 15
Masking Length
Safety ( w -avg) ↑
Over-Refusal ( Avg1 ) ↓
1 Token ( Ours )
85.86
28.00
2 Tokens
79.65
33.33
3 Tokens
77.77
33.00
Appendix
Table 10: Ablation on consecutive sub-word token masking length during attribution (7B SFT).
Validator Pipeline
Benign Accuracy
GPT-4o (Single-Stage)
100.00%
Gemini 2.5 Pro
99.33%
Human Expert
98.67%
Two-Layer Filter (GPT-4o → Gemini 2.5 Pro)
100.00%
Appendix
Table 11: Multi-validator verification accuracy on contrastive benign queries generated via Attribution-Guided Contrastive Augmentation (AGCA).
Model
Llama Guard DSR
Human DSR
Raw Agreement
Base Policy
54.00%
64.67%
77.33%
Ours
99.33%
99.33%
98.67%
Appendix
Table 12: Agreement between Llama Guard and human expert annotations on safety evaluations across 150 sampled instances.
Model
DSR ( ↑ )
FRR ( ↓ )
GRPO
98.67%
94.67%
DeShortcut-Align
98.00%
32.67% (-62.00 pp)
Appendix
Table 13: Defense Success Rate (DSR) on a 150-query WildJailbreak subset and False Refusal Rate (FRR) on paired intention-inverted benign queries.
Figure 8: The pre-training template used to format the input sequence for the model.
Figure 9: The post-training system prompt defining the expected conversation structure and reasoning tags.
Figure 12: Token-level sensitivity attribution across different template modes. The bar charts illustrate the absolute Token Sensitivity Score ( ∣TSS∣ ) of the top-15 input tokens across Mode A (Bare Instruction), Mode B (Base Conversational Template), and Mode C (Full Reasoning Template).
Large reasoning models (LRMs) achieve strong performance on complex reasoning tasks but often generate harmful responses to malicious user queries. This paper investigates the underlying cause of these safety risks and shows that the issue lies in the reasoning structure itself. Based on this insight, we claim that effective safety alignment can be achieved by altering the reasoning structure. We propose AltTrain, a simple yet effective post training method that explicitly alters the reasoning structure of LRMs. AltTrain is both practical and generalizable, requiring no complex reinforcement learning (RL) training or reward design, only supervised finetuning (SFT) with a lightweight 1K training examples. Experiments across LRM backbones and model sizes demonstrate strong safety alignment, along with robust generalization across reasoning, QA, summarization, and multilingual setting.
Large Reasoning Models (LRMs) are commonly trained with reinforcement learning (RL) to improve their generation of chain-of-thought (CoT) reasoning before producing final answers. However, RL rewards are typically assigned based on final answers, providing little or no direct supervision over intermediate reasoning. This can lead to deceptive safety alignment, where the reasoning trace and final answer convey inconsistent safety signals. To systematically investigate this phenomenon, we introduce DSAR (Deceptive Safety Alignment Rate), a metric that jointly assesses reasoning traces and final answers to quantify their safety inconsistency. Across multiple LRMs and benchmarks, we find that deceptive safety alignment is pervasive under standard prompting conditions and is substantially amplified under prefilling attacks. We further provide a hidden representation analysis showing that models exhibit stronger safety discrimination at the final-answer stage than during intermediate reasoning. To close this gap, we propose SARA (Safety-Aware Reasoning Alignment), an RL-based method that rewards both safety-aware reasoning and safe final answers, encouraging early harmful intent recognition and enforcing reasoning-answer consistency. Experiments show that SARA significantly mitigates deceptive safety alignment under both standard and adversarial settings while preserving helpfulness and utility. Code is available at https://github.com/xzhou98/SARA.
Large Reasoning Models (LRMs) pose a dual-surface safety challenge: both intermediate reasoning traces and final answers can contain harmful content. Existing alignment methods often operate at the whole-response level, allowing unsafe reasoning to be masked by a safe-looking final answer. We propose Segment-aware Listwise Target DPO (SaLT-DPO), which addresses this gap through three mechanisms: (1) segment-aware listwise alignment that decomposes responses into reasoning and answer segments, independently scores each segment's safety, and aligns length-normalized segment rewards with soft target distributions over multiple candidates; (2) joint safety coherence regularization that applies a weakest-link principle to promote safety consistency across both segments; and (3) utility anchoring on benign prompts to mitigate over-refusal and reasoning degradation. Experiments on three LRMs show that SaLT-DPO consistently reduces unsafe rates for both reasoning and answer segments while mitigating degradation in benign compliance and preserving general reasoning performance. Ablation studies demonstrate the complementary contributions of its components.
JungMin Yun, Junehyoung Kwon, Hayeong Ryu +3
Department of Artificial Intelligence, Chung-Ang University · Graduate School of Advanced Imaging Sciences, Multimedia and Film, Chung-Ang University