Large Language Models (LLMs) have been extensively used across diverse domains, including virtual assistants, automated code generation, and scientific research. However, they remain vulnerable to jailbreak attacks, which manipulate the models into generating harmful responses despite safety alignment. Recent studies have shown that current safety-aligned LLMs undergo shallow safety alignment. In this work, we conduct an in-depth investigation into the underlying mechanism of this phenomenon and reveal that it manifests through learned ''safety trigger tokens'' that activate the model's safety patterns when paired with the specific input. Through both analysis and empirical verification, we further demonstrate the high similarity of the safety trigger tokens across different harmful inputs. Accordingly, we propose D-STT, a simple yet effective defense algorithm that identifies and explicitly decodes safety trigger tokens of the given safety-aligned LLM to activate the model's learned safety patterns. In this process, the safety trigger is constrained to a single token, which effectively preserves model usability by introducing minimum intervention in the decoding process. Extensive experiments across diverse jailbreak attacks and benign prompts demonstrate that D-STT significantly reduces output harmfulness while preserving model usability and incurring negligible response time overhead, outperforming ten baseline methods.
Figures & tables
Token Range
Generated Tokens
First token
“I” (100%)
First 3 or 4 tokens
“I cannot fulfill” (96%), “I apologize” (4%)
Table 1: Statistics on safety trigger tokens generated by the safety-aligned Llama2-7B-chat model to harmful queries.
Token Range
Defense Method
Generated Tokens
First token
Self-Reminder
“ I ” (100%)
Retokenization
“ I ” (100%)
SafeDecoding
“ I ” (100%)
ICD
“ I ” (98%), “ As ” (2%)
First 3 or 4 tokens
Self-Reminder
“ I cannot fulfill ” (50%), “ I apologize ” (50%)
Retokenization
“ I cannot fulfill ” (26%), “ I apologize ” (70%), “ I’m ” (4%)
Table 2: Statistics on safety trigger tokens generated by four defense strategies.
GCG
AutoDAN
PAIR
DeepInception
ReNeLLM
ICA
GPTFuzzer
Refusal_Sup
Model: Vicuna-7B
No Defense
98% (4.88)
88% (4.94)
88% (4.64)
100% (3.84)
100% (3.40)
16% (1.82)
50% (3.98)
88% (3.84)
Self-Examination
12% (1.44)
4% (1.12)
12% (1.60)
88% (3.06)
88% (3.02)
2% (1.10)
22% (2.24)
30% (1.80)
Paraphrase
60% (2.88)
52% (3.20)
32% (2.12)
90% (3.28)
92% (2.66)
30% (1.48)
64% (2.84)
46% (2.36)
Retokenization
38% (1.96)
92% (3.10)
74% (3.54)
100% (3.72)
100% (2.02)
60% (3.24)
96% (1.66)
52% (2.52)
Self-Reminder
48% (2.98)
68% (4.60)
46% (2.60)
100% (3.30)
96% (3.40)
10% (1.46)
36% (3.92)
82% (3.70)
Table 3: ASR (%) and average harmfulness scores (in parentheses) of different defense strategies across eight attacks, where the best results are highlighted in bold. Note that when deploying SafeDecoding on Gemma-2-9B-IT, the algorithm frequently generates the EOS token immediately and terminates without producing meaningful output. As a result, we report its performance as N/A. Besides, the GCG attack is built upon transfer attacks generated by the Llama2-7B-Chat model, due to the prohibitive cost of crafting attacks for each model individually.
Model
Defense
Just-Eval ( 1−5 ) ↑
ATGR ↓
Helpful
Clear
Factual
Deep
Engaging
Avg.
Llama2-7B-chat
No Defense
3.989
4.696
4.038
3.607
4.229
4.112
1.00 ×
Self-Examination
1.237
2.397
2.553
1.172
1.326
1.738
/
Paraphrase
2.987
3.910
3.507
2.683
3.483
3.315
1.67 ×
Retokenization
2.809
4.440
3.370
2.892
3.810
3.465
1.09 ×
Self-Reminder
4.041
4.585
4.005
3.390
4.197
4.044
1.01 ×
Table 4: Just-Eval scores and ATGR of different defense strategies.
Model
XSTest
Held-out
No Def
D-STT
No Def
D-STT
Llama2
41.2%
42.4%
10%
11%
Vicuna
34.0%
36.0%
6%
7%
Qwen2.5
2.0%
4.0%
0%
0%
Gemma-2
12.0%
20.8%
2%
4%
Table 5: Refusal rate obtained by D-STT and No Defense baseline.
Model
HumanEval (Code)
GSM8K (Math)
No Defense
D-STT
No Defense
D-STT
Llama2-7B-chat
14.02%
12.20%
20.00%
20.00%
Vicuna-7B
16.46%
14.63%
16.50%
15.00%
Qwen2.5-7B-Instruct
71.95%
63.41%
87.50%
87.50%
Gemma-2-9B-IT
58.54%
51.83%
75.00%
75.00%
Table 6: Usability evaluation results on HumanEval and GSM8K. The metrics are pass@1 for HumanEval and final answer accuracy percentage for GSM8K.
Model
“I”
“As”
“It”
Open-Source Models
Llama2-7B-chat
87.50%
8.33%
4.17%
Vicuna-7B
86.11%
4.17%
9.72%
Qwen2.5-7B-Instruct
100.00%
0.00%
0.00%
Gemma-2-9B-IT
100.00%
0.00%
0.00%
DeepSeek-V3.2
100.00%
0.00%
0.00%
Table 7: The Psafety distribution of safety trigger tokens for different models.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Refusal String Keywords
I’m sorry
I’m really sorry
I am sorry
I’m an
I’m just
I can’t provide
Appendix
Table 8: Refusal strings in our experiments.
Defense
Vicuna
Llama2
No Defense
32% (2.30)
0% ( 1.00 )
D-STT
0% ( 1.00 )
0% ( 1.00 )
Appendix
Table 9: ASR (%) and average harmfulness scores (in parentheses) of two methods under DR-attack, where the best results are highlighted in bold.
Attack / Model
Llama2-7B-chat
Vicuna-7B
Qwen2.5-7B-Instruct
Gemma-2-9B-IT
GCG (happy)
0% (1.00)
0% (1.00)
0% (1.00)
0% (1.00)
Appendix
Table 10: ASR (%) and average harmfulness scores (in parentheses) of D-STT under the adaptive GCG (happy) attack. Note that the GCG attack is built upon transfer attacks generated by the Llama2-7B-Chat model, which is consistent with Table 3 in the original paper.
Defense
Vicuna
Llama2
No Defense
88% (4.64)
18% (1.28)
D-STT(Hi)
46% (3.46)
18% (1.12)
D-STT
6% ( 1.34 )
0% ( 1.00 )
Appendix
Table 11: ASR (%) and average harmfulness scores (in parentheses) of different defense methods under PAIR. The best results are highlighted in bold.
Model
“I”
“As”
“It”
Llama2-7B-chat
87.50%
8.33%
4.17%
Llama2-7B-chat ( N=180 )
89.44%
5.00%
5.56%
Qwen2.5-7B-Instruct
100.00%
0.00%
0.00%
Qwen2.5-7B-Instruct ( N=180 )
100.00%
0.00%
0.00%
Appendix
Table 12: Comparison of the Psafety distribution under different sample sizes.
Model / Attack
GCG
AutoDAN
PAIR
DeepInc.
ReNeLLM
ICA
GPTFuzzer
Refusal
Average
Qwen2.5-7B-Instruct
No Defense
10.00% (1.36)
18.00% (4.34)
78.00% (2.88)
84.00% (3.54)
96.00% (4.10)
4.00% (1.22)
52.00% (3.16)
64.00% (2.28)
50.75% (2.86)
D-STT (expansion)
4.00% (1.06)
0.00% (1.20)
34.00% (1.90)
38.00% (2.36)
38.00% (2.22)
0.00% (1.10)
4.00% (1.76)
2.00% (1.16)
15.00% (1.60)
D-STT (original)
4.00% (1.06)
0.00% (1.20)
34.00% (1.90)
38.00% (2.36)
38.00% (2.22)
0.00% (1.10)
4.00% (1.76)
2.00% (1.16)
15.00% (1.60)
Llama2-7B-chat
No Defense
32.00% (2.54)
2.00% (1.08)
18.00% (1.28)
10.00% (1.16)
0.00% (1.00)
0.00% (1.00)
26.00% (2.12)
0.00% (1.00)
11.00% (1.40)
Appendix
Table 13: Comparison of defense performance using Psafety distributions estimated from different sample sizes.
Despite rigorous safety alignment, Large Language Models (LLMs) remain vulnerable to jailbreak attacks. Existing black-box methods often rely on heuristic templates or exhaustive trials, lacking mechanistic interpretability and query efficiency. In this study, we investigate an intrinsic vulnerability in the safety mechanisms of LLMs, where safety alignment relies on a small set of sparsely distributed attention heads, leaving much of the representational space weakly monitored. We formalize this phenomenon with a mathematical jailbreaking model that characterizes the delicate boundary of effective text obfuscation and analytically explains observed jailbreak behaviors. Guided by this model, we propose Babel, an efficient black-box attack framework that exploits the identified safety gap through systematic obfuscation sampling with iterative, feedback-driven distribution refinement, enabling reliable and high-success jailbreak attacks without access to model internals. Comprehensive evaluations on frontier commercial models demonstrate that Babel achieves state-of-the-art attack success rates and superior query efficiency. Specifically, compared to state-of-the-art methods, Babel increases the attack success rate on GPT-4o from 41.33% to 82.67% and on Claude-3-5-haiku from 38.33% to 78.33% within an average of 40 queries, providing a robust red-teaming methodology for LLMs safety research.
Ziwei Wang, Jing Chen, Ruichao Liang +6
Wuhan University · Nanyang Technological University · Southeast University +1
Jailbreak attacks bypass LLM safety alignment, yet their mechanisms remain poorly understood. We provide evidence that attacks do not comprehensively eliminate safety features, but instead selectively suppress specific attention heads. We identify two functionally differentiated types: Adversarially Compromised Heads (ACHs) concentrated in early layers, which are suppressed under attacks, and Safety-Aligned Heads (SAHs) in mid-layers, which maintain robust activations even when attacks succeed. Ablation studies support the causal role of ACHs and the contribution of SAHs to robust activations: suppressing a small number of ACHs is sufficient to induce jailbreak-like behavior on normally refused inputs, while removing SAHs substantially weakens mid-layer safety activations. Token-level attribution further shows that ACH suppression is driven specifically by attack-template tokens, providing a mechanistic account of why attacks can bypass refusal decisions through ACH suppression while leaving internal safety signals sustained by SAHs -- a phenomenon we term Robust Harmful Features. To validate the practical significance of this robustness, we show that simply reading these persistent activations -- without any training -- yields competitive aggregate detection performance with strong adversarial robustness.
Yanchen Yin, Dongqi Han, Linghui Li
Beijing University of Posts and Telecommunications, Beijing, China.
Multi-turn jailbreak attacks progressively erode LLM safety alignment across seemingly innocuous conversation turns, achieving success rates exceeding 90% against state-of-the-art models. Existing alignment-based and guardrail methods suffer from three key limitations: they require costly weight modification, evaluate each turn independently without modeling cumulative safety erosion, and detect attacks only after harmful content has been generated. To address these limitations, we first formulate the proactive early jailbreak detection problem with a new metric, detection lead, that measures how early an attack can be detected before the LLM complies. We then propose SAFEDREAM, a lightweight world-model-based framework that operates as an external module without modifying the LLM's weights. SAFEDREAM introduces three components: (1) a safety state world model that encodes LLM hidden states into a compact safety representation and predicts how it evolves across turns, (2) CUSUM detection that accumulates weak per-turn risk signals into reliable evidence, and (3) contrastive imagination that simultaneously rolls out attack and benign futures in latent space to issue early alarms before jailbreaks occur. On three multi-turn jailbreak benchmarks (XGuard-Train, SafeDialBench, SafeMTData) against 8 baselines, SAFEDREAM achieves the best detection timeliness across all benchmarks (1.06-1.20 turns before compliance) while maintaining competitive false positive rates and outperforming baselines in detection quality.
Bo Yan, Weikai Lin, Yada Zhu +1
University of Central Florida · University of Rochester · IBM Research