Large Reasoning Models (LRMs) are commonly trained with reinforcement learning (RL) to improve their generation of chain-of-thought (CoT) reasoning before producing final answers. However, RL rewards are typically assigned based on final answers, providing little or no direct supervision over intermediate reasoning. This can lead to deceptive safety alignment, where the reasoning trace and final answer convey inconsistent safety signals. To systematically investigate this phenomenon, we introduce DSAR (Deceptive Safety Alignment Rate), a metric that jointly assesses reasoning traces and final answers to quantify their safety inconsistency. Across multiple LRMs and benchmarks, we find that deceptive safety alignment is pervasive under standard prompting conditions and is substantially amplified under prefilling attacks. We further provide a hidden representation analysis showing that models exhibit stronger safety discrimination at the final-answer stage than during intermediate reasoning. To close this gap, we propose SARA (Safety-Aware Reasoning Alignment), an RL-based method that rewards both safety-aware reasoning and safe final answers, encouraging early harmful intent recognition and enforcing reasoning-answer consistency. Experiments show that SARA significantly mitigates deceptive safety alignment under both standard and adversarial settings while preserving helpfulness and utility. Code is available at https://github.com/xzhou98/SARA.
Figures & tables
Figure 1: Illustration of deceptive safety alignment and its mitigation via SARA . Top: A base LRM can exhibit deceptive safety alignment in both standard and adv. prefilling prompting settings. In the standard setting, the model partially recognizes harmful intent but continues reasoning toward harmful compliance; in the adv. prefilling setting, its reasoning is manipulated toward compliance entirely. Both scenarios produce a superficially safe final answer. Bottom: After alignment with our SARA , the model produces safety-aware reasoning that explicitly recognizes and rejects harmful intent, leading to a safe final answer and consistent safety behavior across both settings.
Figure 2: Overview of the evaluation pipeline. Given a harmful input (optionally with adv. prefills), an LRM generates a reasoning ycot and a final answer yans , evaluated independently via two branches. (1) Reasoning: The full reasoning ycot is assessed by LLM Guard for overall safety. Since standard LLM guards are designed to detect whether text is harmful rather than to localize sentence-level safety awareness, we additionally use an LLM Judge to assess each sentence sk individually. The reasoning trace is safe if it either contains at least one safety-aware sentence or is judged safe as a whole. (2) Final answer: LLM Guard evaluates the overall safety of yans . A safety-alignment is considered deceptive if r(ycot)=s(yans) .
Model
StrongReject
SafeChain
Standard
Adv. Prefilling
Standard
Adv. Prefilling
SAR ↑
SS ↑
DSAR ↓
SAR ↑
SS ↑
DSAR ↓
SAR ↑
SS ↑
DSAR ↓
SAR ↑
SS ↑
DSAR ↓
Gemma4-8B
89.10
99.36
5.75
52.10
97.44
37.38
63.60
72.40
20.80
29.00
65.60
31.20
DS-LLaMA3-8B
49.80
53.04
23.64
39.30
50.16
28.75
26.20
57.20
22.40
11.60
56.80
26.20
DS-Qwen3-8B
90.70
99.36
2.56
53.40
87.86
33.23
56.60
79.80
15.00
22.20
61.20
34.40
DS-Qwen2-14B
65.80
61.02
26.20
32.90
53.67
25.56
33.80
64.00
20.20
13.40
59.80
22.00
Table 1: Evaluation of deceptive safety alignment on StrongReject and SafeChain, under both standard and adv. prefilling settings. We report the SAR ↑ of the reasoning trace, the SS ↑ of the final answer, and the proposed DSAR ↓ of quantifying inconsistency between reasoning and answer. Higher SAR and SS indicate safer reasoning and final answers, while lower DSAR indicates better consistency between them. The definition of metrics is provided in Sec. 2.4 and Appendix B.3 . Overall, deceptive safety alignment is less common in safe models (e.g., GPT-oss-20B) and is consistently amplified under prefilling attacks.
Figure 3: Layer-wise average cosine similarity of last-token hidden representations. We compare five conditions across DS-LLaMA3-8B, DS-Qwen2-14B, and GPT-oss-20B: benign–benign (B-B) pairs at the reasoning stage, and harmful–harmful (H-H) and benign–harmful (B-H) pairs at both reasoning and answer stages. Lower B-H similarity indicates stronger internal discrimination between benign and harmful inputs.
Method
Safety
Helpfulness
Utility
StrongReject (Adv. Prefilling)
SafeChain (Standard)
ICL
H-CoT
OR-Bench
Fortress
XSTest
GSM8k
MMLU-Pro
Avg.
SAR ↑
SS ↑
1-DSAR ↑
SAR ↑
SS ↑
1-DSAR ↑
SS ↑
SS ↑
HS ↑
HS ↑
HS ↑
Acc. ↑
Acc. ↑
deepseek-ai/DeepSeek-R1-0528-Qwen3-8B
Original
53.40
87.86
66.77
56.60
79.80
85.00
95.50
16.00
67.55
95.20
66.45
85.44
60.98
72.23
STAR-1
35.50
88.18
54.67
59.60
80.80
86.20
97.50
14.00
49.89
93.20
60.89
83.93
61.85
68.31
SafeChain
23.60
80.83
46.65
49.60
79.20
85.00
98.50
24.00
83.24
98.00
72.67
85.75
60.18
71.54
Table 2: Comparison of safety alignment methods on DSQwen3-8B and DSQwen2-14B across safety, helpfulness, and utility tasks. For safety, we report the SAR ↑ for the reasoning, SS ↑ for the final answer, and 1 − DSAR ↑ under both adv. prefilling (StrongReject) and standard (SafeChain) settings, as well as SS ↑ under unseen adv. attacks (ICL and H-CoT). For helpfulness, we report the HS ↑ on OR-Bench-hard, Fortress, and XSTest. For utility, we report Acc. ↑ on GSM8K and MMLU-Pro. Avg. denotes the harmonic mean across the three task-level scores. Detailed definitions of evaluation metrics are provided in Sec. 4.1 and Appendix B.3 . Bold indicates the best result among all methods.
StrongReject (Adv. Prefilling)
SafeChain (Standard)
Reward variant
SAR
SS
1−DSAR
SAR
SS
1−DSAR
Original model
52.4
87.9
66.8
56.6
79.8
85.0
Final-answer safety (RECAP)
57.8
99.4
70.9
71.0
98.8
96.3
Full-trace reasoning safety
64.2
99.4
84.4
65.4
97.4
96.0
Safety awareness
71.6
93.3
81.5
71.4
87.4
91.8
Full SARA reward
75.1
97.4
85.3
72.8
95.4
96.4
Table 3: Reward-component ablation on DS-Qwen3-8B under adversarial prefilling on StrongReject and standard prompting on SafeChain. The full SARA reward achieves the highest reasoning safety awareness (SAR) and reasoning–answer consistency ( 1−DSAR ), with a trade-off in final-answer safety (SS). All values are percentages; higher is better. Bold denotes the best result.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: PCA visualization of last-layer hidden representations in the original LRMs. Each subplot shows a 2D PCA projection of the last-token hidden states from the final hidden layer for 300 benign prompts and 300 harmful prompts, measured at two generation stages: the reasoning stage (the last token before reasoning begins) and the answer stage (the last token before the final answer begins). Each point represents one prompt. Black-edged circles denote the centroid of each group. Purple arrows connect the benign and harmful centroids within the same stage, and d denotes the Euclidean distance between the two centroids. Across all original models, benign and harmful prompts are less separated at the reasoning stage but become much more clearly separated at the answer stage, suggesting that safety-relevant discrimination becomes stronger later in generation.
Predicted category
Cases
Disagreements
Agreement
Reasoning safety awareness
Safety-aware sentence present
58
4
93.1%
Safety-aware sentence absent
42
3
92.9%
Overall
100
7
93.0%
Final-answer safety
Safe answer
76
2
97.4%
Appendix
Table 4: Agreement between automated judgments and human majority labels on 100 generations. Rows are grouped by the automated prediction; a disagreement occurs when the human majority label differs. Category-level agreement is therefore conditioned on the predicted category, rather than the human reference category. R. and A. denote reasoning and final answer, respectively.
Method
StrongREJECT (Prefilling)
SafeChain (Standard)
gpt-oss-20B
gpt-oss-120B
Qwen2.5-32B
gpt-oss-20B
gpt-oss-120B
Qwen2.5-32B
Original
53.4
59.1
51.4
56.6
61.8
53.8
RECAP
57.8
67.1
54.3
71.0
79.0
74.8
SARA
75.1
82.4
72.5
72.8
82.0
72.0
Appendix
Table 9
Figure 5: Example of a safe reasoning trace without explicit safety-aware sentences. Given a harmful input promoting racial and gender-based exclusion, the model produces a reasoning trace that never explicitly identifies the request as harmful or states a refusal, yet avoids engaging with the harmful intent throughout by reframing the situation toward inclusive alternatives. The final answer is safe. This illustrates that rSA(ycot)=0 while rsafe(ycot)=1 , motivating the use of full-trace safety evaluation as a complementary criterion in our reasoning evaluation pipeline.
Model
Method
Safety
Helpfulness
Utility
Avg.
DS-Qwen3-8B
Original
67.62
76.40
73.21
72.23
STAR-1
64.56
67.99
72.89
68.31
SafeChain
60.92
84.64
72.96
71.54
SafePath
61.57
80.88
71.23
70.35
RECAP
80.53
78.34
74.23
77.61
SARA
84.21
92.07
74.01
82.76
Appendix
Table 5: Average evaluation across safety, helpfulness, and utility tasks. We aggregate the detailed results from Table 2 into three task-level scores. Safety is computed by averaging all safety metrics in Table 2 : SAR ↑ , SS ↑ , and 1-DSAR ↑ under adv. prefilling and standard settings, together with SS ↑ under ICL and H-CoT attacks. Helpfulness is computed by averaging HS ↑ across OR-Bench-hard, Fortress, and XSTest. Utility is computed by averaging Acc. ↑ across GSM8K and MMLU-Pro. Avg. denotes the harmonic mean of these three task-level scores.
StrongReject (Adv. Prefilling)
SafeChain (Standard)
Model
Method
Aug.
SAR
SS
1−DSAR
SAR
SS
1−DSAR
DS-Qwen3-8B
Original
–
52.4
87.9
66.8
56.6
79.8
85.0
RECAP
Yes
57.8
99.4
70.9
71.0
98.8
96.3
RECAP
No
51.8
99.8
64.5
75.2
99.2
98.8
SARA
Yes
75.1
97.4
85.3
72.8
95.4
96.4
SARA
No
62.9
92.1
78.0
89.0
97.8
98.0
Appendix
Table 6: Ablation of counter-aligned prefill augmentation. All metrics are reported as percentages; higher is better. Aug. indicates whether augmentation is used during training. Bold values indicate the better result within each method and model pair, comparing training with and without augmentation.
Figure 6: Model generations given a harmful prompt with harmful prefilling reasoning. Compared with the strongest baseline, RECAP, the model trained with SARA produces safety-aware reasoning that explicitly recognizes the harmful intent and leads to a safe final answer, thereby mitigating deceptive safety alignment.
Figure 7: Prompting instruction for refusal classification. We use this instruction to guide GPT-oss-safeguard in determining whether a model’s final answer to a benign prompt constitutes a refusal. The classifier distinguishes between Refusal and Non_Refusal : an answer is labeled as a refusal if it declines, avoids, redirects, or fails to provide meaningful help, and as non-refusal if it provides substantive information or useful assistance.
Figure 8: Prompting instruction for safety-aware reasoning classification. We use this instruction to guide GPT-oss-safeguard to examine each sentence in a model’s reasoning trace and classify it as Safety_Aware or Not_Safety_Aware . A sentence is classified as Safety_Aware only if it both explicitly recognizes the harmful intent of the user’s request and uses that recognition to avoid providing assistance through refusal, stopping, or redirecting toward a safer alternative.
Figure 9: Prompting instruction for refusal rewarding. We use this instruction to guide DS-Qwen2-32B to evaluate the final answer to a benign prompt and assign a refusal score from 0 to 10 using a structured rubric. The rewards is then computed by 1−10refusal score
Large Reasoning Models (LRMs) pose a dual-surface safety challenge: both intermediate reasoning traces and final answers can contain harmful content. Existing alignment methods often operate at the whole-response level, allowing unsafe reasoning to be masked by a safe-looking final answer. We propose Segment-aware Listwise Target DPO (SaLT-DPO), which addresses this gap through three mechanisms: (1) segment-aware listwise alignment that decomposes responses into reasoning and answer segments, independently scores each segment's safety, and aligns length-normalized segment rewards with soft target distributions over multiple candidates; (2) joint safety coherence regularization that applies a weakest-link principle to promote safety consistency across both segments; and (3) utility anchoring on benign prompts to mitigate over-refusal and reasoning degradation. Experiments on three LRMs show that SaLT-DPO consistently reduces unsafe rates for both reasoning and answer segments while mitigating degradation in benign compliance and preserving general reasoning performance. Ablation studies demonstrate the complementary contributions of its components.
JungMin Yun, Junehyoung Kwon, Hayeong Ryu +3
Department of Artificial Intelligence, Chung-Ang University · Graduate School of Advanced Imaging Sciences, Multimedia and Film, Chung-Ang University
Large reasoning models (LRMs) achieve strong performance on complex reasoning tasks but often generate harmful responses to malicious user queries. This paper investigates the underlying cause of these safety risks and shows that the issue lies in the reasoning structure itself. Based on this insight, we claim that effective safety alignment can be achieved by altering the reasoning structure. We propose AltTrain, a simple yet effective post training method that explicitly alters the reasoning structure of LRMs. AltTrain is both practical and generalizable, requiring no complex reinforcement learning (RL) training or reward design, only supervised finetuning (SFT) with a lightweight 1K training examples. Experiments across LRM backbones and model sizes demonstrate strong safety alignment, along with robust generalization across reasoning, QA, summarization, and multilingual setting.
While Large Reasoning Models (LRMs) excel at complex tasks, they remain highly vulnerable to sophisticated jailbreaks and direct harmful queries. To address this vulnerability, prior works depend heavily on external manual data annotation for safety alignment. However, we observe that LRMs can inherently identify safety risks when being re-presented with original queries alongside their own reasoning trajectories -- a capability we term Latent Safety Awareness. To leverage this safety awareness, we first employ Supervised Fine-Tuning (SFT) to explicitly induce safe tags to trigger safety analysis and guidance following the initial reasoning content for unsafe queries, while preserving standard responses for general queries to ensure adaptive triggering. Subsequently, we apply Direct Preference Optimization (DPO) to further enhance the correctness and stability of the safety analysis and guidance. Notably, responses required for both training stages are entirely generated by models being optimized. With (Safe Trigger) SFT and DPO, experimental results demonstrate significant safety enhancement. For example, the Attack Success Rate (ASR) of DeepSeek-R1-Distill-Llama-8B, on average, drops 24.65% and 36.72% on harmful and jailbreak benchmarks, respectively. Finally, our Safe Trigger method exerts almost no negative impact on general performance or user experience.
Ke Miao, Jiaxin Li, Hongliang Chen +2
The State Key Laboratory of Blockchain and Data Security, Zhejiang University · Hangzhou HighTech Zone (Binjiang) Blockchain and Data Security Research Institute, China · Li Auto Inc. +2