Large Language Models (LLMs) have been extensively used across diverse domains, including virtual assistants, automated code generation, and scientific research. However, they remain vulnerable to jailbreak attacks, which manipulate the models into generating harmful responses despite safety alignment. Recent studies have shown that current safety-aligned LLMs undergo shallow safety alignment. In this work, we conduct an in-depth investigation into the underlying mechanism of this phenomenon and reveal that it manifests through learned ''safety trigger tokens'' that activate the model's safety patterns when paired with the specific input. Through both analysis and empirical verification, we further demonstrate the high similarity of the safety trigger tokens across different harmful inputs. Accordingly, we propose D-STT, a simple yet effective defense algorithm that identifies and explicitly decodes safety trigger tokens of the given safety-aligned LLM to activate the model's learned safety patterns. In this process, the safety trigger is constrained to a single token, which effectively preserves model usability by introducing minimum intervention in the decoding process. Extensive experiments across diverse jailbreak attacks and benign prompts demonstrate that D-STT significantly reduces output harmfulness while preserving model usability and incurring negligible response time overhead, outperforming ten baseline methods.
Figures & tables
Token Range
Generated Tokens
First token
“I” (100%)
First 3 or 4 tokens
“I cannot fulfill” (96%), “I apologize” (4%)
Table 1: Statistics on safety trigger tokens generated by the safety-aligned Llama2-7B-chat model to harmful queries.
Token Range
Defense Method
Generated Tokens
First token
Self-Reminder
“ I ” (100%)
Retokenization
“ I ” (100%)
SafeDecoding
“ I ” (100%)
ICD
“ I ” (98%), “ As ” (2%)
First 3 or 4 tokens
Self-Reminder
“ I cannot fulfill ” (50%), “ I apologize ” (50%)
Retokenization
“ I cannot fulfill ” (26%), “ I apologize ” (70%), “ I’m ” (4%)
Table 2: Statistics on safety trigger tokens generated by four defense strategies.
GCG
AutoDAN
PAIR
DeepInception
ReNeLLM
ICA
GPTFuzzer
Refusal_Sup
Model: Vicuna-7B
No Defense
98% (4.88)
88% (4.94)
88% (4.64)
100% (3.84)
100% (3.40)
16% (1.82)
50% (3.98)
88% (3.84)
Self-Examination
12% (1.44)
4% (1.12)
12% (1.60)
88% (3.06)
88% (3.02)
2% (1.10)
22% (2.24)
30% (1.80)
Paraphrase
60% (2.88)
52% (3.20)
32% (2.12)
90% (3.28)
92% (2.66)
30% (1.48)
64% (2.84)
46% (2.36)
Retokenization
38% (1.96)
92% (3.10)
74% (3.54)
100% (3.72)
100% (2.02)
60% (3.24)
96% (1.66)
52% (2.52)
Self-Reminder
48% (2.98)
68% (4.60)
46% (2.60)
100% (3.30)
96% (3.40)
10% (1.46)
36% (3.92)
82% (3.70)
Table 3: ASR (%) and average harmfulness scores (in parentheses) of different defense strategies across eight attacks, where the best results are highlighted in bold. Note that when deploying SafeDecoding on Gemma-2-9B-IT, the algorithm frequently generates the EOS token immediately and terminates without producing meaningful output. As a result, we report its performance as N/A. Besides, the GCG attack is built upon transfer attacks generated by the Llama2-7B-Chat model, due to the prohibitive cost of crafting attacks for each model individually.
Model
Defense
Just-Eval ( 1−5 ) ↑
ATGR ↓
Helpful
Clear
Factual
Deep
Engaging
Avg.
Llama2-7B-chat
No Defense
3.989
4.696
4.038
3.607
4.229
4.112
1.00 ×
Self-Examination
1.237
2.397
2.553
1.172
1.326
1.738
/
Paraphrase
2.987
3.910
3.507
2.683
3.483
3.315
1.67 ×
Retokenization
2.809
4.440
3.370
2.892
3.810
3.465
1.09 ×
Self-Reminder
4.041
4.585
4.005
3.390
4.197
4.044
1.01 ×
Table 4: Just-Eval scores and ATGR of different defense strategies.
Model
XSTest
Held-out
No Def
D-STT
No Def
D-STT
Llama2
41.2%
42.4%
10%
11%
Vicuna
34.0%
36.0%
6%
7%
Qwen2.5
2.0%
4.0%
0%
0%
Gemma-2
12.0%
20.8%
2%
4%
Table 5: Refusal rate obtained by D-STT and No Defense baseline.
Model
HumanEval (Code)
GSM8K (Math)
No Defense
D-STT
No Defense
D-STT
Llama2-7B-chat
14.02%
12.20%
20.00%
20.00%
Vicuna-7B
16.46%
14.63%
16.50%
15.00%
Qwen2.5-7B-Instruct
71.95%
63.41%
87.50%
87.50%
Gemma-2-9B-IT
58.54%
51.83%
75.00%
75.00%
Table 6: Usability evaluation results on HumanEval and GSM8K. The metrics are pass@1 for HumanEval and final answer accuracy percentage for GSM8K.
Model
“I”
“As”
“It”
Open-Source Models
Llama2-7B-chat
87.50%
8.33%
4.17%
Vicuna-7B
86.11%
4.17%
9.72%
Qwen2.5-7B-Instruct
100.00%
0.00%
0.00%
Gemma-2-9B-IT
100.00%
0.00%
0.00%
DeepSeek-V3.2
100.00%
0.00%
0.00%
Table 7: The Psafety distribution of safety trigger tokens for different models.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Refusal String Keywords
I’m sorry
I’m really sorry
I am sorry
I’m an
I’m just
I can’t provide
Appendix
Table 8: Refusal strings in our experiments.
Defense
Vicuna
Llama2
No Defense
32% (2.30)
0% ( 1.00 )
D-STT
0% ( 1.00 )
0% ( 1.00 )
Appendix
Table 9: ASR (%) and average harmfulness scores (in parentheses) of two methods under DR-attack, where the best results are highlighted in bold.
Attack / Model
Llama2-7B-chat
Vicuna-7B
Qwen2.5-7B-Instruct
Gemma-2-9B-IT
GCG (happy)
0% (1.00)
0% (1.00)
0% (1.00)
0% (1.00)
Appendix
Table 10: ASR (%) and average harmfulness scores (in parentheses) of D-STT under the adaptive GCG (happy) attack. Note that the GCG attack is built upon transfer attacks generated by the Llama2-7B-Chat model, which is consistent with Table 3 in the original paper.
Defense
Vicuna
Llama2
No Defense
88% (4.64)
18% (1.28)
D-STT(Hi)
46% (3.46)
18% (1.12)
D-STT
6% ( 1.34 )
0% ( 1.00 )
Appendix
Table 11: ASR (%) and average harmfulness scores (in parentheses) of different defense methods under PAIR. The best results are highlighted in bold.
Model
“I”
“As”
“It”
Llama2-7B-chat
87.50%
8.33%
4.17%
Llama2-7B-chat ( N=180 )
89.44%
5.00%
5.56%
Qwen2.5-7B-Instruct
100.00%
0.00%
0.00%
Qwen2.5-7B-Instruct ( N=180 )
100.00%
0.00%
0.00%
Appendix
Table 12: Comparison of the Psafety distribution under different sample sizes.
Model / Attack
GCG
AutoDAN
PAIR
DeepInc.
ReNeLLM
ICA
GPTFuzzer
Refusal
Average
Qwen2.5-7B-Instruct
No Defense
10.00% (1.36)
18.00% (4.34)
78.00% (2.88)
84.00% (3.54)
96.00% (4.10)
4.00% (1.22)
52.00% (3.16)
64.00% (2.28)
50.75% (2.86)
D-STT (expansion)
4.00% (1.06)
0.00% (1.20)
34.00% (1.90)
38.00% (2.36)
38.00% (2.22)
0.00% (1.10)
4.00% (1.76)
2.00% (1.16)
15.00% (1.60)
D-STT (original)
4.00% (1.06)
0.00% (1.20)
34.00% (1.90)
38.00% (2.36)
38.00% (2.22)
0.00% (1.10)
4.00% (1.76)
2.00% (1.16)
15.00% (1.60)
Llama2-7B-chat
No Defense
32.00% (2.54)
2.00% (1.08)
18.00% (1.28)
10.00% (1.16)
0.00% (1.00)
0.00% (1.00)
26.00% (2.12)
0.00% (1.00)
11.00% (1.40)
Appendix
Table 13: Comparison of defense performance using Psafety distributions estimated from different sample sizes.