Safety alignment of large reasoning models (LRMs) via supervised fine-tuning (SFT) and reinforcement learning (RL) often yields near-perfect safety scores, yet this apparent success comes at the cost of severe over-refusal and degraded general capabilities. Through systematic empirical analysis, we find that these failures are closely associated with the learning of spurious shortcuts rather than robust intent-sensitive safety evaluation. Specifically, we identify two dominant shortcuts: formatting shortcuts, where refusal behaviors are overly bound to structural prompt templates that frequently appear in safety alignment corpora; and lexical shortcuts, where sensitive keywords reflexively trigger refusals on benign queries. To mitigate reliance on these shortcuts, we propose DeShortcut-Align, a shortcut-decoupling alignment framework that reduces dependence on superficial cues. DeShortcut-Align operates across three coordinated stages: (1) Refusal Sensitivity Attribution, which masks input tokens to quantify their impact on the final refusal response distribution; (2) Attribution-Guided Contrastive Augmentation, which constructs benign contrastive samples using high-sensitivity tokens to mitigate lexical shortcuts; and (3) Counterfactual Consistency Regularization, which constructs template-ablated states via attention blinding to enforce decision consistency across SFT and RL, mitigating formatting shortcut dependence. Experiments on 7B and 14B models demonstrate that DeShortcut-Align significantly improves robustness against template-stripping bypass attacks (reducing performance drops by up to 72%), substantially reduces over-refusal by over 58%, and better preserves general-purpose reasoning capabilities, thereby mitigating the alignment tax commonly observed in safety training.
Figures & tables
Figure 1: Template vulnerability in aligned models. Removing formatting templates consistently reduces DSR across safety benchmarks, despite preserving the underlying user instruction.
Figure 2: Training dynamics of safety alignment. Reward, over-refusal, template gap, and policy-gradient loss reveal shortcut progression and diminishing optimization signals.
Figure 3: Overview of the DeShortcut-Align framework. It reduces reliance on superficial shortcuts through three coordinated stages: (1) sensitivity attribution via TSS , (2) contrastive data augmentation via AGCA for lexical shortcuts, and (3) counterfactual consistency regularization via CCR for formatting shortcuts across SFT and RL.
Model
Method
Template Robustness ( ↑ )
Δavg↓
Over-Refusal ( ↓ )
Reasoning Capabilities ( ↑ )
WildJailbreak
StrongReject
WildChat
w -avg
XSTest
OKTest
FR-test
Avg1
MATH
MMLU
LCB
HEval
Avg2
DeepSeek-R1- Distill-Qwen-7B
Base
45.20 / 49.20 (-4.00)
35.14 / 34.51 (0.63)
58.92 / 51.89 (7.03)
46.42
1.22
0.00
1.00
4.00
1.67
92.80
48.62
39.76
89.63
67.70
SFT
84.80 / 77.60 (7.20)
99.04 / 77.74 (21.30)
86.76 / 76.76 (10.00)
90.20
12.83
56.00
24.00
66.00
48.67
90.20
46.49
31.33
80.49
62.13
GRPO
100.0 / 97.20 (2.80)
100.0 / 93.18 (6.82)
99.73 / 89.73 (10.00)
99.91
6.54
70.00
59.00
94.00
74.33
91.80
43.91
34.94
84.76
63.85
Ours (SFT)
76.00 / 76.80 (-0.80)
99.36 / 96.49 (2.87)
82.97 / 74.32 (8.65)
86.11
3.57
39.00
15.00
47.00
33.67
91.00
40.61
34.34
85.37
62.83
Ours (GRPO)
97.20 / 95.20 (2.00)
98.61 / 94.04 (4.57)
95.41 / 91.08 (4.33)
97.07
3.63
29.00
17.00
47.00
31.00
92.40
45.03
39.16
86.59
65.80
Table 1: Main results evaluating template robustness, over-refusal, and reasoning capabilities. Robustness is reported as with template / without template ( Δavg ) . Best and second-best results are highlighted (the Base model is excluded from the ranking). Shaded rows indicate our proposed DeShortcut-Align methods. ( ↑ ) indicates higher is better, ( ↓ ) indicates lower is better.
Setting
Template Robustness ( ↑ )
Δavg↓
Over-Refusal ( ↓ )
Reasoning Capabilities ( ↑ )
WildJailbreak
StrongReject
WildChat
w -avg
XSTest
OKTest
FR-test
Avg1
MATH
MMLU
LCB
HEval
Avg2
SFT Ablations
SFT
84.80 / 77.60 (7.20)
99.04 / 77.74 (21.30)
86.76 / 76.76 (10.00)
90.20
12.83
56.00
24.00
66.00
48.67
90.20
46.49
31.33
80.49
62.13
w/o CCR
76.00 / 64.80 (11.20)
98.87 / 79.98 (18.89)
82.70 / 69.73 (12.97)
85.86
14.35
34.00
8.00
42.00
28.00
92.20
40.91
34.94
84.76
63.20
w/o AGCA
85.60 / 80.80 (4.80)
98.93 / 94.14 (4.79)
88.65 / 83.78 (4.87)
91.06
4.82
53.00
27.00
64.00
48.00
89.60
39.60
30.72
82.93
60.71
All (Ours)
76.00 / 76.80 (-0.80)
99.36 / 96.49 (2.87)
82.97 / 74.32 (8.65)
86.11
3.57
39.00
15.00
47.00
33.67
91.00
40.61
34.34
85.37
62.83
Table 2: Ablation study evaluating the individual contributions of Counterfactual Consistency Regularization ( CCR ) and Attribution-Guided Contrastive Augmentation ( AGCA ) on DeepSeek-R1-Distill-Qwen-7B. Robustness degradation is reported as Δavg .
Strategy
Safety Defense ( ↑ )
Over-Refusal ( ↓ )
WildJailbreak
StrongReject
WildChat
w -avg
XSTest
OKTest
FR-test
Avg1
Standard SFT
84.80
99.04
86.76
90.20
56.00
24.00
66.00
48.67
+ Random
62.80
98.40
79.46
80.22
40.00
21.00
45.00
35.33
+ Gradient
61.60
97.44
81.35
80.13
39.00
14.00
50.00
34.33
+ Attention
71.20
99.04
85.68
85.31
36.00
19.00
43.00
32.67
+ TSS (Ours)
76.00
98.87
82.70
85.86
34.00
8.00
42.00
28.00
Table 3: Comparison of shortcut attribution heuristics within AGCA under SFT (7B).
Model
Attack Methods
Avg. ↑
PAIR
GCG
TAP
RealSafe-R1
94.0
89.0
98.0
93.7
STAR-1
84.0
84.0
92.0
86.7
SFT
70.0
92.0
86.0
82.7
Ours (SFT)
86.0
95.0
94.0
91.7
RL
96.0
84.0
98.0
92.7
Table 4: Defense Success Rate (DSR) against distinct jailbreak attacks on 7B models.
Figure 4: Training dynamics of DeShortcut-Align during reinforcement learning.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: The regular expression used in the Tier 1 format reward to validate the structural adherence of the model’s output.
Model
Template Robustness ( ↑ )
Δavg↓
Over-Refusal ( ↓ )
Reasoning Capabilities ( ↑ )
WildJailbreak
StrongReject
WildChat
w-avg
XSTest
OKTest
FR-test
Avg1
MATH
MMLU
LCB
HEval
Avg2
SFT methods
SFT
84.80 / 77.60 (7.20)
99.04 / 77.74 (21.30)
86.76 / 76.76 (10.00)
90.20
12.83
56.00
24.00
66.00
48.67
90.20
46.49
31.33
80.49
62.13
STAR-1
87.20 / 70.00 (17.20)
99.68 / 83.39 (16.29)
85.14 / 72.43 (12.71)
90.67
15.40
30.00
24.00
51.00
35.00
90.40
43.91
36.14
83.54
63.50
RealSafe-R1
97.60 / 89.20 (8.40)
100.0 / 91.69 (8.31)
92.43 / 82.16 (10.27)
96.68
8.99
74.00
44.00
89.00
69.00
92.40
44.10
31.33
85.98
63.45
Ours-SFT
76.00 / 76.80 (-0.80)
99.36 / 96.49 (2.87)
82.97 / 74.32 (8.65)
86.11
3.57
39.00
15.00
47.00
33.67
91.00
40.61
34.34
85.37
62.83
Appendix
Table 5: Extended results evaluating template robustness, over-refusal, and reasoning capabilities on DeepSeek-R1-Distill-Qwen-7B. Robustness is reported as with template / without template ( Δavg ) . Best and second-best results within each training paradigm are highlighted. Shaded rows indicate our proposed DeShortcut-Align methods. ( ↑ ) indicates higher is better, ( ↓ ) indicates lower is better.
Model / Seed
Template Robustness ( ↑ )
Δavg↓
Over-Refusal ( ↓ )
General Capabilities ( ↑ )
WildJailbreak
StrongReject
WildChat
w -avg
XSTest
OKTest
FR-test
Avg1
MATH
MMLU
LCB
HEval
Avg2
Qwen3-1.7B (Base)
50.80 / 56.40 (-5.60)
64.86 / 38.34 (26.52)
43.24 / 45.14 (-1.90)
52.97
6.34
0.0
0.0
6.0
2.00
92.00
73.39
30.72
86.59
70.68
SFT (Seed 0)
86.80 / 72.80 (14.00)
99.68 / 61.98 (37.70)
74.32 / 51.08 (23.24)
86.93
24.98
38.0
14.0
50.0
34.00
89.20
63.11
25.30
79.88
64.37
Seed 1
87.20 / 72.80 (14.40)
97.12 / 60.38 (36.74)
73.78 / 48.92 (24.86)
86.03
25.33
37.0
15.0
49.0
33.67
88.80
63.22
26.51
81.71
65.06
Seed 2
79.20 / 68.80 (10.40)
94.89 / 61.98 (32.91)
74.86 / 54.86 (20.00)
82.98
21.10
40.0
17.0
47.0
34.67
88.00
63.32
28.31
81.71
65.34
Mean
84.40 / 71.47 (12.93)
97.23 / 61.45 (35.78)
74.32 / 51.62 (22.70)
85.31
23.80
38.33
15.33
48.67
34.11
88.67
63.22
26.71
81.10
64.92
Appendix
Table 6: Multi-seed evaluation across three random seeds (seeds 0, 1, and 2) on Qwen3-1.7B, reporting template robustness, over-refusal, and general capabilities.
Figure 6: Token-level KL divergence across reasoning trajectories. Higher sustained divergence indicates that alignment-induced policy changes persist across longer generations.
Model
Template Robustness ( ↑ )
Δavg↓
Over-Refusal ( ↓ )
WildJailbreak
StrongReject
WildChat
w -avg
XSTest
OKTest
FR-Test
Avg
PPO
100.00 / 89.60 (10.40)
98.51 / 72.10 (26.41)
96.22 / 80.81 (15.41)
98.24
17.41
81.0
65.0
92.0
79.33
+ DeShortcut-Align
94.32 / 92.00 (2.32)
95.42 / 88.39 (7.03)
94.32 / 86.49 (7.83)
94.69
5.73
29.0
29.0
46.0
34.67
GRPO
100.00 / 97.20 (2.80)
100.00 / 93.18 (6.82)
99.73 / 89.73 (10.00)
99.91
6.54
70.0
59.0
94.0
74.33
+ DeShortcut-Align
97.20 / 95.20 (2.00)
98.61 / 94.04 (4.57)
95.41 / 91.08 (4.33)
97.07
3.63
29.0
17.0
47.0
31.00
Appendix
Table 7: Performance improvements of standard PPO and GRPO baselines when equipped with DeShortcut-Align on DeepSeek-R1-Distill-Qwen-7B.
Figure 7: Training dynamics of DeShortcut-Align during RL training. (a) Dynamic adaptation of regularization weight ( λdyn ). (b) Convergence of counterfactual consistency loss ( Lcf ). (c) Evolution and stabilization of the overall training reward.
Table 15
Masking Length
Safety ( w -avg) ↑
Over-Refusal ( Avg1 ) ↓
1 Token ( Ours )
85.86
28.00
2 Tokens
79.65
33.33
3 Tokens
77.77
33.00
Appendix
Table 10: Ablation on consecutive sub-word token masking length during attribution (7B SFT).
Validator Pipeline
Benign Accuracy
GPT-4o (Single-Stage)
100.00%
Gemini 2.5 Pro
99.33%
Human Expert
98.67%
Two-Layer Filter (GPT-4o → Gemini 2.5 Pro)
100.00%
Appendix
Table 11: Multi-validator verification accuracy on contrastive benign queries generated via Attribution-Guided Contrastive Augmentation (AGCA).
Model
Llama Guard DSR
Human DSR
Raw Agreement
Base Policy
54.00%
64.67%
77.33%
Ours
99.33%
99.33%
98.67%
Appendix
Table 12: Agreement between Llama Guard and human expert annotations on safety evaluations across 150 sampled instances.
Model
DSR ( ↑ )
FRR ( ↓ )
GRPO
98.67%
94.67%
DeShortcut-Align
98.00%
32.67% (-62.00 pp)
Appendix
Table 13: Defense Success Rate (DSR) on a 150-query WildJailbreak subset and False Refusal Rate (FRR) on paired intention-inverted benign queries.
Figure 8: The pre-training template used to format the input sequence for the model.
Figure 9: The post-training system prompt defining the expected conversation structure and reasoning tags.
Figure 12: Token-level sensitivity attribution across different template modes. The bar charts illustrate the absolute Token Sensitivity Score ( ∣TSS∣ ) of the top-15 input tokens across Mode A (Bare Instruction), Mode B (Base Conversational Template), and Mode C (Full Reasoning Template).