CodeMimicry: Exploiting Safety Generalization Lag in Large Language Models via Structured Code Completion
Organizations: School of Computer Science and Technology, Zhejiang Sci-Tech University · Zhejiang Key Laboratory of Digital Fashion and Data Governance, Zhejiang Sci-Tech University · China Academy of Information and Communications Technology
Abstract
Large language models have achieved remarkable capabilities across diverse domains, yet their safety alignment remains vulnerable to jailbreak attacks. In this work, we identify a previously underexplored failure mode - safety generalization lag - where alignment trained predominantly on natural language fails to transfer to the code domain. We show that this lag induces a code-completion blind spot, allowing malicious intent embedded within syntactically valid code to evade safety mechanisms. To exploit this vulnerability, we propose CodeMimicry, a fully automated black-box jailbreak framework that generates structured, object-oriented code prompts to induce harmful outputs via code completion. Experiments on 8 state-of-the-art commercial LLMs demonstrate that CodeMimicry achieves a 96.25% attack success rate with 1.51 queries on average, significantly outperforming both template-based and optimization-based baselines. Beyond empirical performance, we provide a mechanistic analysis of code-based jailbreaks through latent space representations, including projection onto refusal-related directions and activation steering. This analysis offers an explanation of how CodeMimicry bypasses safety mechanisms in code-related domains. Our findings reveal a weakness in current safety alignment and highlight the need for robust alignments in structured domains such as code.
Figures & tables
| Method | GPT-4o | GPT-4.1 | GPT-5 chat | Claude 3.7 Sonnet | Claude Sonnet 4 | Gemini 2.5 Pro | Llama 3.1 70B | Llama 3.3 70B | Ave. |
| Template-based | |||||||||
| CipherChat [ 19 ] | 6.00 | 12.00 | 22.00 | 0.00 | 0.00 | 78.00 | 0.00 | 0.00 | 14.75 |
| ReNeLLM [ 27 ] | 50.00 | 74.00 | 4.00 | 42.00 | 2.00 | 38.00 | 44.00 | 66.00 | 40.00 |
| CodeAttack [ 22 ] | 36.00 | 34.00 | 6.00 | 20.00 | 4.00 | 20.00 | 62.00 | 44.00 | 28.25 |
| FlipAttack [ 11 ] | 66.00 | 96.00 | 96.00 | 60.00 | 0.00 | 76.00 | 0.00 | 12.00 | 50.75 |
| EquaCode [ 24 ] | 82.00 | 100 | 84.00 | 84.00 | 24.00 | 94.00 | 70.00 | 88.00 | 78.25 |
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | GPT-4o | GPT-4.1 | GPT-5 chat | Claude 3.7 Sonnet | Claude Sonnet 4 | Gemini 2.5 Pro | Llama 3.1 70B | Llama 3.3 70B | Ave. |
| Manual Design | |||||||||
| CipherChat | 0.00 | 8.57 | 15.71 | 0.00 | 0.00 | 74.29 | 0 | 2.86 | 13.39 |
| ReNeLLM | 27.14 | 45.71 | 0.00 | 0.2 | 0.00 | 37.14 | 15.71 | 35.71 | 22.68 |
| CodeAttack | 11.43 | 18.57 | 17.14 | 17.14 | 0.43 | 31.43 | 34.29 | 22.86 | 19.64 |
| FlipAttack | 74.29 | 87.14 | 75.71 | 52.86 | 0.00 | 81.43 | 0.00 | 22.86 | 49.29 |
| EquaCode | 42.86 | 84.29 | 5.71 | 81.43 | 15.71 | 78.57 | 67.14 | 80.00 | 56.96 |
| Evaluation | GPT-5-chat | GPT-4o | StrongREJECT | Dictionary-based | HarmfulScore | Llama Guard |
| Human1 | 91.00 | 88.67 | 84.67 | 59.33 | 88.33 | 72.33 |
| Human2 | 88.67 | 88.00 | 82.33 | 54.67 | 90.00 | 78.67 |
| Human3 | 89.67 | 90.00 | 84.33 | 52.67 | 86.67 | 79.00 |
| Human4 | 92.33 | 91.67 | 86.33 | 58.00 | 84.33 | 78.33 |
| Average | 90.25 | 89.58 | 84.42 | 56.00 | 87.33 | 77.08 |
| Rejection Dictionary | |
| I’m sorry | illegal |
| Sorry | not legal |
| I am sorry | My apologies |
| I apologize | but I cannot |
| As an | If you have any other non-malicious requests |
| As an AI | I’m not able to provide |
| Target Model | PAIR | AutoDan-Turbo | CodeMimicry | |||
| ASR | AQ | ASR | AQ | ASR | AQ | |
| GPT-4o | 32.00 | 4.40 | 50.00 | 2.32 | 90.00 | 1.86 |
| GPT-4.1 | 18.00 | 4.72 | 54.00 | 2.18 | 100.00 | 1.04 |
| GPT-5-chat | 18.00 | 4.76 | 4.00 | 3.66 | 96.00 | 1.58 |
| Claude-3.7 | 4.00 | 4.94 | 10.00 | 3.46 | 100.00 | 1.10 |
| Claude-4 | 0.00 | 5.00 | 0.00 | 4.76 | 90.00 | 2.54 |
| Target Model | PAIR | AutoDan-Turbo | CodeMimicry | |||
| ASR | AQ | ASR | AQ | ASR | AQ | |
| GPT-4o | 42.85 | 4.28 | 37.14 | 2.74 | 98.57 | 1.34 |
| GPT-4.1 | 64.29 | 3.22 | 57.14 | 2.27 | 100.00 | 1.21 |
| GPT-5-chat | 18.57 | 4.58 | 14.29 | 3.41 | 100.00 | 1.27 |
| Claude-3.7 | 5.71 | 4.76 | 7.14 | 2.91 | 100.00 | 1.17 |
| Claude-4 | 0.00 | 5.00 | 0.00 | 4.95 | 74.29 | 2.74 |
| Target Model | Setting | CodeMimicry | CodeAttack | CodeChameleon |
| GPT-4o | No sanitizer | 90% | 36% | 94% |
| Regex sanitizer | 20% | 8% | 34% | |
| Llama-3.1-70B | No sanitizer | 96% | 62% | 48% |
| Regex sanitizer | 22% | 2% | 40% | |
| Llama-3.3-70B | No sanitizer | 98% | 44% | 80% |
| Regex sanitizer | 26% | 0% | 58% |
| Model | SCR (%) | RSR (%) | FVR (%) |
| GPT-4o | 100.0 | 100.0 | 100.0 |
| GPT-4.1 | 100.0 | 100.0 | 100.0 |
| GPT-5-Chat | 100.0 | 100.0 | 100.0 |
| Claude-3.7-Sonnet | 100.0 | 100.0 | 100.0 |
| Claude-Sonnet-4 | 97.8 | 97.8 | 95.6 |
| Gemini-2.5-Pro | 98.0 | 94.0 | 92.0 |
| Method | ASR (%) | Attacker AQ | Avg. Tokens (In / Out) |
| CodeMimicry | 96.3 | 1.51 | 1656 / 1341 |
| PAIR | 34.5 | 4.18 | 8030 / 5005 |
| Method | GPT-4o | GPT-4.1 | GPT-5 chat | Claude 3.7 Sonnet | Claude Sonnet 4 | Gemini 2.5 Pro | Llama 3.1 70B | Llama 3.3 70B |
| CodeMimicry | 90.00 | 100.00 | 96.00 | 100.00 | 90.00 | 100.00 | 96.00 | 98.00 |
| Llama Guard | 88.0 | 90.00 | 82.00 | 84.00 | 90.00 | 76.00 | 92.00 | 88.00 |
| Perplexity filter | 86.00 | 94.00 | 94.00 | 96.00 | 84.00 | 100.00 | 92.00 | 94.00 |
| SelfDefend | 12.00 | 6.00 | 12.00 | 8.00 | 14.00 | 4.00 | 8.00 | 6.00 |
| Method | DS-R1-8B | Safechain | Safepath |
| CodeMimicry | 92 | 94 | 42 |
| AutoDan-Turbo | 90 | 86 | 4 |
| PAIR | 82 | 64 | 8 |