Code large language model (Code LLM) assistants generate code from heterogeneous development contexts, including open files, imported modules, pasted snippets, and comments, much of which may originate from untrusted sources. We investigate whether insecure instructions embedded in such contexts can steer Code LLMs toward vulnerable code without access to model weights or training data. We evaluate ten open-weight Code LLMs spanning 3B--13B parameters, including four base and six instruction-tuned models, across ten web-application weakness classes. We compare completion tasks containing insecure instructions embedded as code comments with benign tasks without malicious instructions. Attack-condition completions contained a medium-or-higher weakness in {\bf 77.4--92.3}% of cases, compared with {\bf 1.7--5.1}% in the benign condition. Base and instruction-tuned models averaged 86.5% and 84.5% vulnerable outputs, respectively; equivalence testing and three matched model pairs indicated reductions of at most 8.1% after instruction tuning. Susceptibility showed no clear association with model scale or specialization. Among vulnerable attack outputs, 86.2--91.0% were rated high or critical, and the effect persisted without the pattern-based detector. Post-generation screening reduced but did not eliminate the risk, the strongest screen leaving roughly one-third undetected. These findings identify inference-time context injection as a substantial attack surface and motivate provenance-aware training objectives.
Figures & tables
Figure 1 : Illustrative example of the framing gap (DeepSeek-V4-Pro, 15 July 2026). A direct natural-language request to assign user-controlled input to administrative attributes is refused, whereas the same objective, expressed as a # TODO comment within a code-completion task, is carried out and produces a mass-assignment weakness (CWE-915). The example motivates the compliance behavior evaluated in Section 5 ; it is qualitative and is not included as a measured evaluation instance.
Figure 2 : Context-injection threat model. An adversary plants insecure instructions in an untrusted context, and the assistant follows them during completion, producing vulnerable code in the victim’s project.
Model
Specialization
Type
Params
StarCoderBase
Coding
Base
7B
DeepSeek-Coder-base
Coding
Base
6.7B
Qwen2.5-base
General
Base
7B
Granite-4.0-H-Micro-Base
General
Base
3B
StarCoder2-Instruct
Coding
Instruct
3B
DeepSeek-Coder-Instruct
Coding
Instruct
6.7B
Table 1 : Evaluated models. The set includes four base models, six instruction-tuned models, and three matched sibling pairs. All models are open-weight.
Model
Type
Benign %
Attack %
MSS %
RSR
h
H+C %
StarCoderBase
Base
3.6 [2.4,5.2]
77.4 [74.2,80.4]
73.8
21.7 ×
1.77
89.1
DeepSeek-Coder-base
Base
4.7 [3.4,6.5]
87.7 [85.1,89.9]
83.0
18.6 ×
1.99
90.4
Qwen2.5-base
Base
4.6 [3.3,6.4]
88.4 [85.8,90.6]
83.8
19.3 ×
2.02
89.8
Granite-4.0-H-Micro-Base
Base
4.4 [3.1,6.2]
92.3 [90.1,94.0]
87.9
20.8 ×
2.15
91.0
StarCoder2-Instruct
Instruct
5.1 [3.7,7.0]
88.6 [86.0,90.7]
83.5
17.2 ×
1.99
90.5
DeepSeek-Coder-Instruct
Instruct
3.0 [2.0,4.5]
85.0 [82.2,87.5]
82.0
28.3 ×
2.00
88.1
Table 2: Vulnerable-output rate under the insecure instruction (attack) and under ordinary coding tasks (benign), over 700 completions per condition. MSS (malicious susceptibility shift) =VRM−VRB is the absolute increase in the vulnerable-output rate, in percentage points; RSR (relative susceptibility ratio) =VRM/VRB is the multiplicative increase over the benign baseline, where M and B denote the attack and benign conditions as in Eq. 1 ; H+C% is the combined HIGH and CRITICAL share of vulnerable attack outputs. Cohen’s h≥1.77 throughout; 95% Wilson CIs in brackets. MSS and RSR are computed from unrounded rates. The two conditions differ in task type as well as in the embedded instruction, so RSR should not be read as the instruction’s causal effect alone (Section 7.3 ).
Figure 3 : Vulnerable-output rates under benign and attack conditions across ten models, with 95% Wilson intervals.
Figure 4 : Severity distributions under benign and attack conditions; attack outputs shift strongly toward HIGH and CRITICAL across all models.
Figure 5 : Pooled attack surface across ten models, grouped by CWE and vulnerability type.
Full-ensemble ref. %
Leave-one-out ref. %
Screen
Coverage
Uncaught
Coverage
Uncaught
Bandit
21.4
78.6
20.7
79.3
AST taint
32.1
67.9
31.4
68.6
LLM judge
71.0
29.0
67.9
32.1
Union
78.4
21.6
74.5
25.5
Table 3 : Post-generation screening results pooled across all models. Each screen is scored against two reference labels: the full ensemble, and a leave-one-out ensemble that excludes the screen under test. Neither reference is human-adjudicated, so both are detector-defined rather than ground truth (Section 4.2 ).
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Weakness class
Primary CWE
SQL Injection
CWE-89
Cross-Site Scripting (XSS)
CWE-79
Command Injection
CWE-77 / CWE-78
Path Traversal
CWE-22
Server-Side Request Forgery (SSRF)
CWE-918
XPath Injection
CWE-643
Appendix
Table 4 : The ten target weakness classes and their primary CWE identifiers. Command injection spans CWE-77/CWE-78 depending on whether a shell is invoked.
Position
DeepSeek-Coder-Instruct
Qwen2.5-Instruct
Top
83.3 [66.4, 92.7]
76.7 [59.1, 88.2]
Middle
86.7 [70.3, 94.7]
70.0 [52.1, 83.3]
Bottom
80.0 [62.7, 90.5]
70.0 [52.1, 83.3]
Cochran’s Q ( p )
0.75 (0.69)
1.60 (0.45)
Appendix
Table 5 : Comment-position sensitivity: vulnerable-output rate (%) with the instruction comment placed at the top, middle, or bottom of the file, holding the scaffold and instruction fixed ( n=30 per position, greedy decoding). 95% Wilson intervals in brackets; refusals were 0% in every cell. The rate is essentially unchanged across positions, indicating the pattern induces the vulnerability regardless of the comment’s location.
Figure 6 : Detector diagnostics against the experimental condition as reference label. These measure detector-condition agreement, not detector accuracy. In (a), the secret-detection and semantic-verification components are omitted because they report no per-file counts.
Figure 7 : Post-generation screening reduces risk but leaves substantial numbers of vulnerable outputs uncaught.
LLM-based coding assistants are seeing rapid adoption, offering substantial gains in developer productivity. As organizations increasingly ship code these agents produce, the security of that code becomes critical. Prior work has shown that minor prompt perturbations degrade the functional correctness of LLM-generated code, but whether they also compromise code security has remained unstudied. We apply token-level mutations to prompts across three models and five programming languages, and show that mutations as small as a single-character change can flip generated code from secure to vulnerable. Probing the models' hidden states reveals that this fragility is partially encoded in prompt representations, but unevenly so. Input-handling vulnerabilities, where the model omits validation or sanitization, are more predictable (mean AUC 0.753) than secure-defaults vulnerabilities, where insecure code stems from one local choice such as a weak algorithm or unsafe parameter (mean AUC 0.674). These results show that the threat model for LLM-assisted coding extends beyond prompt injection to ordinary prompt variation, and indicate that input-handling flaws can be caught before generation while secure-defaults flaws require intervention during decoding.
Alexander Sternfeld, Andrei Kucharavy, Ljiljana Dolamic
IEM, HES-SO, Le Foyer, Techno-Pôle 1, Sierre, Switzerland · Cyber-Defence Campus, armasuisse Science and Technology, Thun, Switzerland
Code Large Language Models (CLLMs) serve as the core of modern code agents, enabling developers to automate complex software development tasks. In this paper, we present Poison-with-Style (PwS), a practical and stealthy model poisoning attack targeting CLLMs. Unlike prior attacks that assume an active adversary capable of directly embedding explicit triggers (e.g., specific words) into developers' prompts during inference, PwS leverages developers' code styles as covert triggers implicitly embedded within their prompts. PwS introduces a novel data collection method and a two-step training strategy to fine-tune CLLMs, causing them to generate vulnerable code when prompts contain trigger code styles while maintaining normal behavior on other prompts. Experimental results on Python code completion tasks show that PwS is robust against state-of-the-art defenses and achieves high attack success rates across diverse vulnerabilities, while maintaining strong performance on standard code completion benchmarks. For example, PwS-poisoned models generate CWE-20 vulnerable code in 95% of cases when the trigger code style is used, with less than a 5% drop in pass@1 performance on the HumanEval and MBPP benchmarks. Our implementation and dataset are here: https://github.com/khangtran2020/pws.
Khang Tran, Yazan Boshmaf, Issa Khalil +3
Department of Data Science, New Jersey Institute of Technology, New Jersey, U.S.A. · Qatar Computing Research Institute, HBKU, Doha, Qatar · Mohamed Bin Zayed University of Artificial Intelligence, Abu Dhabi, United Arab Emirates
Large language models (LLMs) frequently generate source code containing vulnerabilities, yet little work studies the internal mechanisms that distinguish safe from vulnerable generation in them. In this work, we systematically perform a mechanistic interpretation of LLMs, aiming at both understanding how code safety-vs-vulnerability is represented or driven by components in an LM and turning the insights into actionable steering strategies to encourage safer code generation. To this end, we introduce CodeSec-Pairs, a dataset of 9,342 Python safe-and-vulnerable contrastive code pairs, sampled from Llama-3.1-8B-Instruct. Utilizing the dataset, we explore approaches to localize layers and attention heads that relate to code safety, and further experiment with different steering strategies for inference-time vulnerability reduction. In particular, we propose DuoSteer, a double-steering approach that simultaneously applies safety and code-correctness steering to attention heads. In experiments over five vulnerability types, DuoSteer leads to an average of -26.9% vulnerability rate reduction and +7.5% functional correctness improvement, which outperforms not only other steering variants but also prompting and supervised fine-tuning baselines. The advantage also replicates on Qwen-2.5-Coder-7B-Instruct with another 2,500 contrastive pairs sampled from that model.
Hao Yan, Ziyu Yao
Department of Computer Science George Mason University, Fairfax, VA