Large language models are vulnerable to prompt injection attacks, where third-party adversarial content can hijack the model's behavior. In this paper, we study the role played by the adversarial data's input modality, and identify a systematic asymmetry: multimodal LLMs are more likely to follow adversarial instruction when they appear as text than when the same instruction is delivered through a non-textual channel (e.g., as an image). We hypothesize that this modality gap arises from text-centric instruction tuning, which teaches models to obey textual instructions while treating other modalities mainly as content to parse or describe. We then demonstrate how this gap can be turned into a training-free defense, by rendering all untrusted payloads as typographic images (or audio) before they reach the model. Across ten models and two prompt injection benchmarks (DirectInject and AgentDojo) we show that our defense Pictionary consistently reduces attack success rates even against the strongest adaptive attacks and human red teamers, while largely preserving benign utility. We further show that benign fine-tuning on image-rendered instructions erodes the modality gap, tracing it to the text-centric instruction-tuning distribution.
Figures & tables
Figure 1: Image-rendering defense against prompt injection. When untrusted content reaches the model, either through a user-uploaded document or a third-party tool output, the default text pipeline exposes the model to prompt injections embedded in that content. Our defense renders the untrusted content as an image before passing it to the model, substantially reducing the attack success rate.
Figure 2: Instruction-following across six VLMs (mean ± std over 5 runs). (a) In isolation, each modality is followed at near-ceiling rates. (b) Under conflict, text compliance dominates on every model; the remainder are responses that follow neither instruction. Results reveal a systematic preference for the textual channel rather than an inability to act on image-supplied instructions.
Figure 3: The untrusted user-supplied document is rendered as an image instead of passed as text.
DirectInject
AgentDojo
Model
ASR (%) ↓ Text → Image \color[rgb]{0,0,1}(\Delta)
Utility (%) ↑ under attack
ASR (%) ↓ Text → Image \color[rgb]{0,0,1}(\Delta)
Utility (%) ↑ under attack
Claude Opus 4.7
5.4→0.0
(5.4)
99.5→100
0.0→0.0
(0.0)
90.4→90.2
Claude Haiku 4.5
17.9→0.0
(17.9)
99.5→98.7
1.5→0.8
(0.8)
70.5→62.2
GPT-5.5
0.0→0.0
(0.0)
100.0→100.0
0.8→0.0
(0.8)
92.7→92.6
GPT-5.4 nano
14.3→0.0
(14.3)
100.0→98.7
14.6→0.0
(14.6)
50.3→43.6
GPT-5.4 mini
75.0→0.0
(75.0)
99.7→100.0
29.2→7.7
(21.5)
71.4→64.9
Table 1: Text vs. image rendering of untrusted content on DirectInject and AgentDojo. Each cell reports textual input → image rendering. For ASR, the value in parentheses is the absolute ASR reduction Δ ; lower ASR and higher utility are better. We highlight the improvement of Pictionary in blue .
ASR (%)
Claude Opus 4.7
Claude Haiku 4.5
GPT-5.5
GPT-5.4 mini
Kimi-K2.6
Gemini 3.1 Flash Lite
Text ↓
14.3
97.6
19.0
83.3
77.8
100.0
Image ↓
11.9
9.5
16.7
31.0
40.0
97.8
Table 2: Union of adaptive attacks on AgentDojo: automated red-teaming, RL attack, and human red-teaming. A case counts as a successful attack if any of the three attack methods succeeds. We report ASR (%) for textual input and image rendering. Lower ASR is better.
Defense
GPT-5.4 mini
GPT-5.4
Kimi-K2.6
No defense
82.1 / 98
87.5 / 99
71.4 / 96
PromptLocate
82.1 / 94
85.7 / 99
71.4 / 96
DataSentinel
64.3 / –
69.6 / –
58.9 / –
DataFilter
32.1 / 88
12.5 / 87
17.9 / 87
Ours
3.6 / 100
3.6 / 100
3.6 / 99
Table 3: DirectInject comparison using 3 ART-discovered templates (Appendix A.4 ). Cells report ASR / utility (%). DataSentinel uses detect-then-block, so no utility to report.
Model
SWE-bench
τ2 -Bench
Kimi-K2.6
65.0→65.0
87.2→85.9
GPT-5.4 mini
66.0→56.0
44.5→43.6
Claude Haiku 4.5
66.0→57.0
50.4→48.3
Table 4: Task success (%), text → image.
No attack
Under attack
Model
Input
Utility
ASR
Utility
GPT-Audio Mini
Text
100.0±0.0
27.3
100.0
Audio
95.0±22.4
1.3
91.3
Qwen3.5 Omni-Plus
Text
100.0±0.0
50.5†
82.1†
Audio
100.0±0.0
15.8
93.4
Table 5: Audio rendering under the setup of Table 1 . All values are percentages. † Qwen text includes 1 attack template hard-rejected by the API before reaching the model.
Model
Text ASR
Image ASR
Qwen3.5-0.8B
9.5→26.2
0.0→11.9
Qwen3.5-9B
2.4→69.0
0.0→04.8
InternVL3.5-1B
14.3→90.5
0.0→00.0
InternVL3.5-38B
50.0→57.1
2.4→04.8
Table 6: Text and image ASR (%). Each cell reports base → instruction-tuned.
Figure 4: Image-channel ASR on DirectInject before and after fine-tuning on benign image-rendered instructions.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Items
User tasks (document type)
Article (natural-language)
Code snippet
Resume
Email
Private-data access
Retrieve a password
Retrieve a passport number
Appendix
Table 7: The four benign user tasks and fourteen injected target tool calls in DirectInject. Each injected goal is an attacker-chosen tool call unrelated to the benign task; an attack succeeds when that tool is invoked.
Figure 5: Human red-teaming interface. Participants select a user task and injection goal, submit a template in text or image mode, and inspect the victim’s response and tool-call trace.
ASR ↓
Utility ↑
Model
text
single image
multi image
text
single image
multi image
GPT-5.4 mini
2%
0%
0%
48%
28%
49%
Gemini 3.1 Flash Lite
12%
1%
1%
83%
20%
70%
Claude Haiku 4.5
0%
0%
0%
90%
14%
99%
Appendix
Table 8: Stress test for the image channel with long tool outputs: ASR and utility for text vs. single vs. multi-image.
Figure 6: The same source file under the plain renderer and the code renderer. Plain rendering wraps long lines and drops their indentation; code rendering keeps indentation, numbers each line, and paginates with a header.
Figure 15
Text
Image
Model
Strict
Relaxed
Δ
Strict
Relaxed
Δ
Claude Opus 4.7
0.0
0.0
+0.0
0.0
0.0
+0.0
Claude Haiku 4.5
1.5
1.5
+0.0
0.8
0.8
+0.0
GPT-5.5
0.8
0.8
+0.0
0.0
0.0
+0.0
GPT-5.4 nano
14.6
14.6
+0.0
0.0
0.0
+0.0
GPT-5.4 mini
29.2
32.3
+3.1
7.7
12.3
+4.6
Appendix
Table 9: AgentDojo ASR (%) under strict and relaxed scoring, using the union over seven static templates on 130 task–goal pairs. Δ is relaxed minus strict in percentage points. This is a static-template evaluation, separate from the adaptive banking union.
Model
Text ↓
Image ↓
GPT-5.4 mini
100.0
17.8
Claude Haiku 4.5
98.2
10.7
Appendix
Table 10: Worst-case ASR (%) on DirectInject under the union of the RL and ART attacks.
Static attack
RL-generated suffix
Benchmark
Model
Text
Image
Δ
Text
Transfer
Image (adaptive)
Δ
DirectInject
GPT-5.4 mini
40.8%
0.0%
40.8
86.7%
0.0%
1.0%
85.7
Gemini 3.1 Flash Lite
66.3%
16.3%
50.0
92.9%
14.3%
56.1%
36.8
Claude Haiku 4.5
10.2%
0.0%
10.2
22.4%
0.0%
0.0%
22.4
Kimi-K2.6
12.2%
0.0%
12.2
34.7%
3.1%
2.0%
31.6
AgentDojo
GPT-5.4 mini
14.4%
4.4%
10.0
27.8%
1.1%
12.2%
15.6
Appendix
Table 11: Static and RL attack success (%) on DirectInject and the AgentDojo banking/workspace subset. Results cover the first three static templates. Transfer renders a text-optimized suffix as an image without re-optimization; Image (adaptive) optimizes against image delivery. Static Δ is Text minus Image; RL Δ is Text minus max(Transfer,Image(adaptive)) , in percentage points.
Figure 9: Agentic auto red teaming: ASR under text, transfer (text-optimized payload re-rendered as image), and image, on DirectInject and AgentDojo.
Figure 10: Human-only ASR on the 42 attempted AgentDojo banking pairs, using the union of 386 successful human-discovered strategies from seven red teamers and 144 hours of work. The three-method banking union instead uses 45 pairs, with no human successes on the three additional pairs.
Model
Benign utility (%) ↑
Text → Image
Claude Opus 4.7
98.5→98.5
Claude Haiku 4.5
80.8→87.7
GPT-5.5
100.0→100.0
GPT-5.4 nano
56.9→51.2
GPT-5.4 mini
82.5→85.6
Appendix
Table 12: Benign utility (%) on AgentDojo without attack, text → image rendering. Benign utility on DirectInject is 100% for every model under both renderings and is omitted.
Model
Word accuracy (%)
Critical-word accuracy (%)
Text → Image
Text → Image
Claude Opus 4.7
100.0→99.9
100.0→100.0
Claude Haiku 4.5
100.0→99.4
100.0→92.7
GPT-5.5
100.0→99.7
100.0→97.3
GPT-5.4 nano
100.0→97.2
100.0→92.5
GPT-5.4 mini
100.0→99.5
100.0→97.3
Appendix
Table 13: Transcription fidelity on AgentDojo pages, text → image, over up to 60 pages per model and modality; pages the model declined to transcribe are excluded. Critical words are tokens containing a digit, @ , or :// , scored byte-exactly.
Model
SWE-bench
τ2 -Bench
Text → Image
Text → Image
Claude Haiku 4.5
100.0→96.3
100.0→99.3
GPT-5.4 mini
96.8→95.4
100.0→96.3
Kimi-K2.6
99.9→99.0
100.0→99.7
Appendix
Table 14: Word accuracy (%), text → image, on pages sampled from each model’s image-arm rollouts. Each benchmark has 100 pages per model, except GPT-5.4 mini on τ2 -Bench (97 pages). The same content is probed in both modalities.
Model, suite
Reference (page as drawn)
Transcription
Claude Haiku 4.5, banking, image
order accordingly. runtime annotation appended post-retrieval
order accordingly. — runtime annotation appended post-retrieval
Table 15: Representative transcription discrepancies, with differing spans highlighted . The examples cover glyph and quotation substitutions, digit-run errors in AgentDojo identifiers, and character errors in airline records and source code.
Tokens
Cost ($)
Benchmark
Model
Text
Image
Text
Image
SWE-bench Verified
GPT-5.4 mini
1.08M
1.68M
0.610
0.794
Kimi-K2.6
1.33M
1.74M
0.207
0.321
Claude Haiku 4.5
2.02M
2.89M
0.352
0.484
τ2 -Bench
GPT-5.4 mini
95k
112k
0.042
0.056
Kimi-K2.6
116k
188k
0.076
0.110
Appendix
Table 16: The cost of the image defense. Token usage and API cost per episode on SWE-bench Verified and τ2 -Bench. For SWE-bench, one episode is one rollout (one instance attempt); for τ2 -Bench, one episode is one conversation.
Typographic prompt injection exploits vision language models' (VLMs) ability to read text rendered in images, posing a growing threat as VLMs power autonomous agents. Prior work typically focus on maximizing attack success rate (ASR) but does not explain \emph{why} certain renderings bypass safety alignment. We make two contributions. First, an empirical study across four VLMs including GPT-4o and Claude, twelve font sizes, and ten transformations reveals that multimodal embedding distance strongly predicts ASR (r=−0.71 to −0.93, p<0.01), providing an interpretable, model agnostic proxy. Since embedding distance predicts ASR, reducing it should improve attack success, but the relationship is mediated by two factors: perceptual readability (whether the VLM can parse the text) and safety alignment (whether it refuses to comply). Second, we use this as a red teaming tool: we directly maximize image text embedding similarity under bounded ℓ∞ perturbations via CWA-SSA across four surrogate embedding models, stress testing both factors without access to the target model. Experiments across five degradation settings on GPT-4o, Claude Sonnet 4.5, Mistral-Large-3, and Qwen3-VL confirm that optimization recovers readability and reduces safety aligned refusals as two co-occurring effects, with the dominant mechanism depending on the model's safety filter strength and the degree of visual degradation.
Large vision-language models (LVLMs) have emerged as a powerful paradigm for multimodal intelligence, but their growing deployment also expands the attack surface of prompt injection. Despite this growing concern, existing attacks still suffer from a critical limitation: the injected prompt for one modality only steers the model's interpretation of that singular input. Alternatively, these attacks remain multimodal but fail to achieve cross-modal prompt perturbation. To bridge this gap, we introduce a novel cross-modal prompt injection attack CrossMPI, which can steer the model's interpretation of both textual and visual inputs via image-only prompt injection. Our design is underpinned by the following key breakthroughs. First, we turn the focus of the injected prompt perturbation optimization from the visual embedding space (typically with only 105 parameters) to the model hidden state space (for multimodal information integration and with 107 parameters). Then, two strategies are adopted to mitigate the optimization challenges posed by the larger parameter space. To constrain the optimized model parameter space, we introduce a layer selection strategy that identifies the layers most critical to multimodal integration. Interestingly, deviating from the past experience, our analysis reveals that the optimal layers for LVLM prompt perturbation reside in the middle of the model rather than the last. To constrain the image perturbation space, we propose a new distance-decremental perturbation budget assignment strategy that allocates budgets decrementally as the pixel distance to semantic-critical regions increases. Extensive experiments across multiple LVLMs and datasets show that our method significantly outperforms baseline approaches.
Prompt injection is a critical security threat in large language model (LLM) applications, where attackers hijack model behavior by embedding malicious instructions in user or external data. Existing detection methods only detect the presence of injection and refuse to respond upon detection, overlooking the fact that for many modern aligned models, well-crafted instructions can resist most injection attacks. This means that the injection robustness varies significantly across instructions and models. This leads to widespread unnecessary over-refusal: inputs containing injections that the model could have handled correctly are rejected incorrectly. To deal with this over-refusal issue, we propose BASIS (Robustness-Aware Prompt Injection Defense). This defense method uses the Attention Competition Ratio (ρ) as features to train two sparse linear probes: an existence probe and a breach probe. Both probes make defense decisions through cascaded gating, which does not require additional LLM inference. BASIS comprises three stages: injection existence detection, per-sample breach prediction, and instruction robustness assessment; the online cascade refuses only when the model would actually be compromised and thus avoids over-refusal on robust instructions. Experiments across four tasks and six open-source LLMs show that BASIS maintains near-perfect injection detection while substantially reducing over-refusal on safe attack samples, especially under robust instruction templates.
Laiqiao Qin, Tianqing Zhu, Longxiang Gao +1
City University of Macau, Macau · Qilu University of Technology