CodeMimicry: Exploiting Safety Generalization Lag in Large Language Models via Structured Code Completion
Authors: Zhen Liang, Hai Huang, Wentao Chen
Organizations: School of Computer Science and Technology, Zhejiang Sci-Tech University · Zhejiang Key Laboratory of Digital Fashion and Data Governance, Zhejiang Sci-Tech University · China Academy of Information and Communications Technology
Large language models have achieved remarkable capabilities across diverse domains, yet their safety alignment remains vulnerable to jailbreak attacks. In this work, we identify a previously underexplored failure mode - safety generalization lag - where alignment trained predominantly on natural language fails to transfer to the code domain. We show that this lag induces a code-completion blind spot, allowing malicious intent embedded within syntactically valid code to evade safety mechanisms. To exploit this vulnerability, we propose CodeMimicry, a fully automated black-box jailbreak framework that generates structured, object-oriented code prompts to induce harmful outputs via code completion. Experiments on 8 state-of-the-art commercial LLMs demonstrate that CodeMimicry achieves a 96.25% attack success rate with 1.51 queries on average, significantly outperforming both template-based and optimization-based baselines. Beyond empirical performance, we provide a mechanistic analysis of code-based jailbreaks through latent space representations, including projection onto refusal-related directions and activation steering. This analysis offers an explanation of how CodeMimicry bypasses safety mechanisms in code-related domains. Our findings reveal a weakness in current safety alignment and highlight the need for robust alignments in structured domains such as code.
Figures & tables
Figure 1 : CodeMimicry
Method
GPT-4o
GPT-4.1
GPT-5 chat
Claude 3.7 Sonnet
Claude Sonnet 4
Gemini 2.5 Pro
Llama 3.1 70B
Llama 3.3 70B
Ave.
Template-based
CipherChat [ 19 ]
6.00
12.00
22.00
0.00
0.00
78.00
0.00
0.00
14.75
ReNeLLM [ 27 ]
50.00
74.00
4.00
42.00
2.00
38.00
44.00
66.00
40.00
CodeAttack [ 22 ]
36.00
34.00
6.00
20.00
4.00
20.00
62.00
44.00
28.25
FlipAttack [ 11 ]
66.00
96.00
96.00
60.00
0.00
76.00
0.00
12.00
50.75
EquaCode [ 24 ]
82.00
100
84.00
84.00
24.00
94.00
70.00
88.00
78.25
Table 1 : Comparison of different methods across various LLMs on AdvBench. We report the attack success rates (%), CodeMimicry is evaluated with T=5 . The bold and underlined values are the best and runner-up ASR.
Figure 3
Figure 4 : The t-SNE visualization of intermediate representations from 3 datasets in the Llama-3-8B-Instruct model: (a) layer 14 and (b) layer 16.
Figure 5 : The Y-axis represents the projection score of hidden states onto the refusal direction defined by μℓ(Dfail) and μℓ(DBC) . Positive values indicate activation of refusal mechanisms, while negative values indicate compliance. Higher projection scores indicate stronger alignment with the refusal direction In Llama-3-8B-Instruct, failed jailbreak code samples show increasing refusal activation, whereas our successful jailbreak code samples closely track benign code in the safe subspace throughout all layers
Figure 6 : Activation steering results on CodeMimicry samples. (a) Steering applied to samples that originally elicited harmful responses. (b) Steering applied to samples that were originally refused by the model. Lighter colors indicate fewer refusals after steering.
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Method
GPT-4o
GPT-4.1
GPT-5 chat
Claude 3.7 Sonnet
Claude Sonnet 4
Gemini 2.5 Pro
Llama 3.1 70B
Llama 3.3 70B
Ave.
Manual Design
CipherChat
0.00
8.57
15.71
0.00
0.00
74.29
0
2.86
13.39
ReNeLLM
27.14
45.71
0.00
0.2
0.00
37.14
15.71
35.71
22.68
CodeAttack
11.43
18.57
17.14
17.14
0.43
31.43
34.29
22.86
19.64
FlipAttack
74.29
87.14
75.71
52.86
0.00
81.43
0.00
22.86
49.29
EquaCode
42.86
84.29
5.71
81.43
15.71
78.57
67.14
80.00
56.96
Appendix
Table 5 : Comparison of different methods across various LLMs on HarmBench. We report the success rates (%), CodeMimicry is evaluated with T=5 . The bold and underlined values are the best and runner-up ASR.
Evaluation
GPT-5-chat
GPT-4o
StrongREJECT
Dictionary-based
HarmfulScore
Llama Guard
Human1
91.00
88.67
84.67
59.33
88.33
72.33
Human2
88.67
88.00
82.33
54.67
90.00
78.67
Human3
89.67
90.00
84.33
52.67
86.67
79.00
Human4
92.33
91.67
86.33
58.00
84.33
78.33
Average
90.25
89.58
84.42
56.00
87.33
77.08
Appendix
Table 6 : Accuracy results of 6 evaluation methods on 300 prompt-and-response pairs from AdvBench. The results present the accuracy (%) of each evaluation method relative to the assessments made by human experts.
Rejection Dictionary
I’m sorry
illegal
Sorry
not legal
I am sorry
My apologies
I apologize
but I cannot
As an
If you have any other non-malicious requests
As an AI
I’m not able to provide
Appendix
Table 7 : The dictionary-based rejection keywords.
Target Model
PAIR
AutoDan-Turbo
CodeMimicry
ASR ↑
AQ ↓
ASR ↑
AQ ↓
ASR ↑
AQ ↓
GPT-4o
32.00
4.40
50.00
2.32
90.00
1.86
GPT-4.1
18.00
4.72
54.00
2.18
100.00
1.04
GPT-5-chat
18.00
4.76
4.00
3.66
96.00
1.58
Claude-3.7
4.00
4.94
10.00
3.46
100.00
1.10
Claude-4
0.00
5.00
0.00
4.76
90.00
2.54
Appendix
Table 8 : Efficiency comparison of CodeMimicry with automated methods (PAIR and AutoDan-Turbo) on AdvBench. We report the ASR(%) and the AQ. The bold values are the best results.
Target Model
PAIR
AutoDan-Turbo
CodeMimicry
ASR ↑
AQ ↓
ASR ↑
AQ ↓
ASR ↑
AQ ↓
GPT-4o
42.85
4.28
37.14
2.74
98.57
1.34
GPT-4.1
64.29
3.22
57.14
2.27
100.00
1.21
GPT-5-chat
18.57
4.58
14.29
3.41
100.00
1.27
Claude-3.7
5.71
4.76
7.14
2.91
100.00
1.17
Claude-4
0.00
5.00
0.00
4.95
74.29
2.74
Appendix
Table 9 : Efficiency comparison of CodeMimicry with automated methods (PAIR and AutoDan-Turbo) on HarmBench. We report the ASR(%) and the AQ. The bold values are the best results.
Target Model
Setting
CodeMimicry
CodeAttack
CodeChameleon
GPT-4o
No sanitizer
90%
36%
94%
Regex sanitizer
20%
8%
34%
Llama-3.1-70B
No sanitizer
96%
62%
48%
Regex sanitizer
22%
2%
40%
Llama-3.3-70B
No sanitizer
98%
44%
80%
Regex sanitizer
26%
0%
58%
Appendix
Table 10 : Attack success rates under different sanitization settings.
Model
SCR (%)
RSR (%)
FVR (%)
GPT-4o
100.0
100.0
100.0
GPT-4.1
100.0
100.0
100.0
GPT-5-Chat
100.0
100.0
100.0
Claude-3.7-Sonnet
100.0
100.0
100.0
Claude-Sonnet-4
97.8
97.8
95.6
Gemini-2.5-Pro
98.0
94.0
92.0
Appendix
Table 11 : Functional validity of code generated by successful CodeMimicry attacks on AdvBench. SCR denotes syntax correctness rate, RSR denotes runtime success rate, and FVR denotes end-to-end functional validity rate.
Method
ASR (%)
Attacker AQ
Avg. Tokens (In / Out)
CodeMimicry
96.3
1.51
1656 / 1341
PAIR
34.5
4.18
8030 / 5005
Appendix
Table 12 : Attacker-side cost comparison on AdvBench50 across 8 target models.
Method
GPT-4o
GPT-4.1
GPT-5 chat
Claude 3.7 Sonnet
Claude Sonnet 4
Gemini 2.5 Pro
Llama 3.1 70B
Llama 3.3 70B
CodeMimicry
90.00
100.00
96.00
100.00
90.00
100.00
96.00
98.00
Llama Guard
88.0
90.00
82.00
84.00
90.00
76.00
92.00
88.00
Perplexity filter
86.00
94.00
94.00
96.00
84.00
100.00
92.00
94.00
SelfDefend
12.00
6.00
12.00
8.00
14.00
4.00
8.00
6.00
Appendix
Table 13 : CodeMimicry performance under defense mechanisms: We report the defense effects of Llama Guard, Perplexity filter and SelfDefend against CodeMimicry.
Method
DS-R1-8B
Safechain
Safepath
CodeMimicry
92
94
42
AutoDan-Turbo
90
86
4
PAIR
82
64
8
Appendix
Table 14 : Defense results on reasoning methods. Safechain and Safepath stand for DeepSeek-R1-8B models reasoning trained by their own method. We report the ASR(%), the lower ASR the defense better.
Figure 7 : The t-SNE visualization results of layer 16 (a), layer 18 (b), the PCA visualization results of layer 16 (c),layer 18 (d), on Llama-3-8b aligned by Circuit Breaker.
Figure 8 : The t-SNE visualization results of layer 16 (a), layer 18 (b), the PCA visualization results of layer 16 (c),layer 18 (d), on Llama-3-8b aligned by Representation Bending..
Figure 9 : The PCA visualization results of layer 12 (a), layer 14 (b), layer 15 (c),layer 16 (d), layer 18 (c),layer 20 (d), on Llama-3-8b-instruct.
Figure 10 : The t-SNE visualization results of layer 12 (a), layer 14 (b), layer 15 (c),layer 16 (d), layer 18 (e), layer 20 (f), on Llama-3-8B-Instruct.
Figure 11 : The Y-axis represents the projection score of hidden states onto the refusal direction defined by μℓ(Dfail) and μℓ(DBC) . Positive values indicate activation of refusal mechanisms, while negative values indicate compliance. Higher projection scores indicate stronger alignment with the refusal direction In Llama-3.1-70B-Instruct, failed jailbreak code samples show increasing refusal activation, whereas our successful jailbreak code samples closely track benign code in the safe subspace throughout all layers
Figure 12 : Activation steering results on CodeMimicry samples. (a) Steering applied to samples that originally elicited harmful responses. (b) Steering applied to samples that were originally refused by the model. Lighter colors indicate fewer refusals after steering.
Figure 13 : The Code Completion Trigger is designed to simulate a regular user requesting LLM code completion. The language can be set to Python, JavaScript, etc., and the jailbreak code is generated by the attacker.
Figure 14 : Our Jailbreak prompt for attacker LLM to generate the jailbreak code.
Figure 15 : A benign prompt for DeepSeek-R1 to generate benign code. the goal is set by benign query.