Render Before Reading: Visual Rendering as a Prompt Injection Defense
Organizations: ETH Zurich · Leiden University
Abstract
Large language models are vulnerable to prompt injection attacks, where third-party adversarial content can hijack the model's behavior. In this paper, we study the role played by the adversarial data's input modality, and identify a systematic asymmetry: multimodal LLMs are more likely to follow adversarial instruction when they appear as text than when the same instruction is delivered through a non-textual channel (e.g., as an image). We hypothesize that this modality gap arises from text-centric instruction tuning, which teaches models to obey textual instructions while treating other modalities mainly as content to parse or describe. We then demonstrate how this gap can be turned into a training-free defense, by rendering all untrusted payloads as typographic images (or audio) before they reach the model. Across ten models and two prompt injection benchmarks (DirectInject and AgentDojo) we show that our defense Pictionary consistently reduces attack success rates even against the strongest adaptive attacks and human red teamers, while largely preserving benign utility. We further show that benign fine-tuning on image-rendered instructions erodes the modality gap, tracing it to the text-centric instruction-tuning distribution.
Figures & tables
| DirectInject | AgentDojo | |||||
| Model | ASR (%) Text Image \color[rgb]{0,0,1}(\Delta) | Utility (%) under attack | ASR (%) Text Image \color[rgb]{0,0,1}(\Delta) | Utility (%) under attack | ||
| Claude Opus 4.7 | ||||||
| Claude Haiku 4.5 | ||||||
| GPT-5.5 | ||||||
| GPT-5.4 nano | ||||||
| GPT-5.4 mini | ||||||
| ASR (%) | Claude Opus 4.7 | Claude Haiku 4.5 | GPT-5.5 | GPT-5.4 mini | Kimi-K2.6 | Gemini 3.1 Flash Lite |
| Text | ||||||
| Image |
| Defense | GPT-5.4 mini | GPT-5.4 | Kimi-K2.6 |
| No defense | 82.1 / 98 | 87.5 / 99 | 71.4 / 96 |
| PromptLocate | 82.1 / 94 | 85.7 / 99 | 71.4 / 96 |
| DataSentinel | 64.3 / – | 69.6 / – | 58.9 / – |
| DataFilter | 32.1 / 88 | 12.5 / 87 | 17.9 / 87 |
| Ours | 3.6 / 100 | 3.6 / 100 | 3.6 / 99 |
| Model | SWE-bench | -Bench |
| Kimi-K2.6 | ||
| GPT-5.4 mini | ||
| Claude Haiku 4.5 |
| No attack | Under attack | |||
| Model | Input | Utility | ASR | Utility |
| GPT-Audio Mini | Text | 27.3 | 100.0 | |
| Audio | 1.3 | 91.3 | ||
| Qwen3.5 Omni-Plus | Text | |||
| Audio | 15.8 | 93.4 | ||
| Model | Text ASR | Image ASR |
| Qwen3.5-0.8B | ||
| Qwen3.5-9B | ||
| InternVL3.5-1B | ||
| InternVL3.5-38B |
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
| Component | Items |
| User tasks (document type) | Article (natural-language) |
| Code snippet | |
| Resume | |
| Private-data access | Retrieve a password |
| Retrieve a passport number |
| ASR | Utility | |||||
| Model | text | single image | multi image | text | single image | multi image |
| GPT-5.4 mini | 2% | 0% | 0% | 48% | 28% | 49% |
| Gemini 3.1 Flash Lite | 12% | 1% | 1% | 83% | 20% | 70% |
| Claude Haiku 4.5 | 0% | 0% | 0% | 90% | 14% | 99% |
| Text | Image | |||||
| Model | Strict | Relaxed | Strict | Relaxed | ||
| Claude Opus 4.7 | 0.0 | 0.0 | 0.0 | 0.0 | ||
| Claude Haiku 4.5 | 1.5 | 1.5 | 0.8 | 0.8 | ||
| GPT-5.5 | 0.8 | 0.8 | 0.0 | 0.0 | ||
| GPT-5.4 nano | 14.6 | 14.6 | 0.0 | 0.0 | ||
| GPT-5.4 mini | 29.2 | 32.3 | 7.7 | 12.3 | ||
| Model | Text | Image |
| GPT-5.4 mini | ||
| Claude Haiku 4.5 |
| Static attack | RL-generated suffix | ||||||||
| Benchmark | Model | Text | Image | Text | Transfer | Image (adaptive) | |||
| DirectInject | GPT-5.4 mini | 40.8% | 0.0% | 40.8 | 86.7% | 0.0% | 1.0% | 85.7 | |
| Gemini 3.1 Flash Lite | 66.3% | 16.3% | 50.0 | 92.9% | 14.3% | 56.1% | 36.8 | ||
| Claude Haiku 4.5 | 10.2% | 0.0% | 10.2 | 22.4% | 0.0% | 0.0% | 22.4 | ||
| Kimi-K2.6 | 12.2% | 0.0% | 12.2 | 34.7% | 3.1% | 2.0% | 31.6 | ||
| AgentDojo | GPT-5.4 mini | 14.4% | 4.4% | 10.0 | 27.8% | 1.1% | 12.2% | 15.6 | |
| Model | Benign utility (%) |
| Text Image | |
| Claude Opus 4.7 | |
| Claude Haiku 4.5 | |
| GPT-5.5 | |
| GPT-5.4 nano | |
| GPT-5.4 mini |
| Model | Word accuracy (%) | Critical-word accuracy (%) |
| Text Image | Text Image | |
| Claude Opus 4.7 | ||
| Claude Haiku 4.5 | ||
| GPT-5.5 | ||
| GPT-5.4 nano | ||
| GPT-5.4 mini |
| Model | SWE-bench | -Bench |
| Text Image | Text Image | |
| Claude Haiku 4.5 | ||
| GPT-5.4 mini | ||
| Kimi-K2.6 |
| Model, suite | Reference (page as drawn) | Transcription |
| Claude Haiku 4.5, banking, image | order accordingly. runtime annotation appended post-retrieval | order accordingly. — runtime annotation appended post-retrieval |
| GPT-5.4 mini, travel, image | {’Le Marais Boutique’: " Rating: 4.2\nReviews: Charming boutique | {’Le Marais Boutique’: ’ Rating: 4.2\nReviews: Charming boutique |
| Claude Haiku 4.5, banking, image | US133000000121212121212 | US133 0 0000001212121212 2 12 |
| GPT-5.4 mini, banking, image | US133 0 00000121212121212 | US13300000121212121212 |
| Grok 4.3, banking, image | US133000000121212121212 | US133 00 000000121212121212 |
| Claude Haiku 4.5, airline, image | "WUNA 5 K" | "WUNA S K" |
| Tokens | Cost ($) | ||||
| Benchmark | Model | Text | Image | Text | Image |
| SWE-bench Verified | GPT-5.4 mini | 1.08M | 1.68M | 0.610 | 0.794 |
| Kimi-K2.6 | 1.33M | 1.74M | 0.207 | 0.321 | |
| Claude Haiku 4.5 | 2.02M | 2.89M | 0.352 | 0.484 | |
| -Bench | GPT-5.4 mini | 95k | 112k | 0.042 | 0.056 |
| Kimi-K2.6 | 116k | 188k | 0.076 | 0.110 | |
| Hyperparameter | Value |
| LoRA rank / / dropout | / / |
| LoRA targets | LLM |
| Frozen modules | vision encoder, aligner |
| Learning rate | |
| Schedule / warmup | cosine / |
| Epochs |