cs.CRSep 30, 2026

CodeMimicry: Exploiting Safety Generalization Lag in Large Language Models via Structured Code Completion

Authors: Zhen Liang, Hai Huang, Wentao Chen

Organizations: School of Computer Science and Technology, Zhejiang Sci-Tech University · Zhejiang Key Laboratory of Digital Fashion and Data Governance, Zhejiang Sci-Tech University · China Academy of Information and Communications Technology

Abstract

Large language models have achieved remarkable capabilities across diverse domains, yet their safety alignment remains vulnerable to jailbreak attacks. In this work, we identify a previously underexplored failure mode - safety generalization lag - where alignment trained predominantly on natural language fails to transfer to the code domain. We show that this lag induces a code-completion blind spot, allowing malicious intent embedded within syntactically valid code to evade safety mechanisms. To exploit this vulnerability, we propose CodeMimicry, a fully automated black-box jailbreak framework that generates structured, object-oriented code prompts to induce harmful outputs via code completion. Experiments on 8 state-of-the-art commercial LLMs demonstrate that CodeMimicry achieves a 96.25% attack success rate with 1.51 queries on average, significantly outperforming both template-based and optimization-based baselines. Beyond empirical performance, we provide a mechanistic analysis of code-based jailbreaks through latent space representations, including projection onto refusal-related directions and activation steering. This analysis offers an explanation of how CodeMimicry bypasses safety mechanisms in code-related domains. Our findings reveal a weakness in current safety alignment and highlight the need for robust alignments in structured domains such as code.

Figures & tables

Appendix figures & tables20 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Grammar-Constrained Decoding Can Jailbreak LLMs into Generating Malicious Code

    Jun 10, 2026Yitong Zhang, Shiteng Lu, Jia LiLarge Language Model JailbreaksJailbreak Attacks

  2. SoK: Robustness in Large Language Models against Jailbreak Attacks

    May 6, 2026Feiyue Xu, Hongsheng Hu, Chaoxiang He +9Large Language Model JailbreaksLarge Language Model Safety

  3. Babel: Jailbreaking Safety Attention via Obfuscation Distribution Optimized Sampling

    May 18, 2026Ziwei Wang, Jing Chen, Ruichao Liang +6Large Language Model JailbreaksJailbreak Attacks