Multimodal web agents that process both screenshots and accessibility trees are increasingly deployed to interact with web interfaces, yet their dual-stream architecture opens an underexplored attack surface: an adversary who injects content into the webpage DOM simultaneously corrupts both observation channels with a consistent deceptive narrative. Our vulnerability analysis on MiniWob++ reveals that attacks including a visual component far outperform text-only injections, exposing critical gaps in text-centric VLM safety training. Motivated by this finding, we propose Dual-Modality Multi-Stage Adversarial Safety Training (DMAST), a framework that formalizes the agent-attacker interaction as a two-player general-sum Markov game and co-trains both players through a three-stage pipeline: (1) imitation learning from a strong teacher model, (2) oracle-guided supervised fine-tuning that uses a novel zero-acknowledgment strategy to instill task-focused reasoning under adversarial noise, and (3) adversarial reinforcement learning via Group Relative Policy Optimization (GRPO) self-play. On out-of-distribution tasks, DMAST nearly halves the attack success rate (41.2%→21.4%) while raising task completion by over 60% relative (6.2%→10.2%). Our approach outperforms established training-based defenses and complements prompt-based defenses, demonstrating genuine co-evolutionary progress and robust generalization to complex, unseen environments. Code is available at https://github.com/huajianduzhuo-code/DMAST_official.
Figures & tables
Figure 1: Overview of the agent–attacker interaction process. At each timestep, the attacker observes the clean webpage state and injects malicious HTML/CSS into the DOM, simultaneously corrupting both the screenshot and the accessibility tree. The agent then acts on the modified observations, and the environment transitions accordingly.
Figure 2: Illustration of the HTML injection mechanism. The attacker VLM processes the clean state (b) to generate a structured action αt (a) with color-coded components. This payload is injected into the DOM to produce the malicious state (c).
Config.
Agent
Attacker
Success (%)
Success (%)
No Attack
36.9 ± 1.5
N/A
Text-Only
15.9 ± 1.2
24.1 ± 1.4
Image-Only
15.6 ± 1.1
34.4 ± 1.5
Dual
15.8 ± 1.2
35.7 ± 1.5
Table 1: Vulnerability analysis across attack modalities (%, MiniWob++, Gemma-3-27B-IT, 1,000 episodes) showing agent success rate and attacker success rate with standard errors.
MiniWob++
VisualWebArena
Method
ASR ↓
TSR ↑
ASR ↓
TSR ↑
Base Model
18.9±1.3
14.0±1.2
41.2±5.0
6.2±2.4
Prompt Defense
7.4±0.9
15.3±1.2
8.2±2.8
3.1±1.7
SPAG
14.4±1.2
22.7±1.4
35.1±4.8
6.2±2.4
ART
14.6±1.2
21.8±1.4
30.9±4.7
8.2±2.8
Online SFT
15.1±1.2
18.4±1.3
33.0±4.8
7.2±2.6
Table 2: Results (%) on unseen MiniWob++ tasks and OOD VisualWebArena tasks. Subscripts denote standard errors. The middle block also serves as the stage-wise ablation of our pipeline, and Stage 1 also serves as the pure-distillation baseline. Bold marks the best result among the 12B agents; the 27B teacher is a capability reference.
Figure 3: Cross-evaluation heatmaps of agent and attacker checkpoints across different training stages. Each cell reports the success rate when pairing a specific agent checkpoint (column) against a specific attacker checkpoint (row).
Figure 4: Attacker diversity across self-play iterations. Attack text generally grows less repetitive (a–b); strategy-level diversity also generally increases (c), though the entropy change is small.
Reasoning pattern
Before
After
Acknowledgment of the attack
22
17
Hedging (“I’ll attempt”, “I assume”)
38
33
Story-building (“before X, I must Y”)
14
12
Back-reference to failed past steps
17
11
Decisive declaration (“I will / I must”)
77
59
Reasoning length (number of words)
68
67
Table 3: Frequency (%) of reasoning patterns across all agent steps before and after RL (multi-label; percentages may sum to more than 100%).
Method
ASR ↓
TSR ↑
(a) Misinformation objective, Gemma-3-12B
Base Model
44.7±3.1
13.6±2.1
SPAG
41.7±3.0
19.3±2.4
ART
43.6±3.1
26.9±2.7
DMAST
38.3±3.0
28.4±2.8
(b) Cross-model, Qwen3-VL-8B-Instruct
Table 4: Generalization beyond the training setting (%). (a) Cross-objective : the attacker is re-prompted to plant a specific wrong value instead of exfiltrating data; (b) Cross-model : the pipeline is partially reproduced on Qwen3-VL-8B-Instruct and evaluated on the same MiniWob++ test set.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter
Stage 1
Stage 2
Stage 3
Imitation
Orc. SFT
Adv. RL
Learning rate
5×10−6
2×10−6
2×10−6
LR scheduler
cosine
cosine
cosine
Warmup ratio
0.03
0.03
0.03
Batch size
128
128
128
Per-device batch size
1
1
1
Appendix
Table 5: Training hyperparameters for each stage.
Figure 5: Screenshot comparison for the choose-date task. (a) The original page contains only a date input field and a submit button. (b) After the attacker’s HTML injection, a fake “Date Format Error” overlay with an email verification form is prepended to the task area, obscuring the original elements and misleading the agent into entering sensitive user information.
Figure 6: Screenshot comparison for the drag-items task (agent success). (a) The original page contains a sortable list of names. (b) After the attacker’s HTML injection, a fake “Move Confirmation” overlay with a password verification form is prepended to the task area.
Figure 7: N -gram entropy across self-play iterations. Higher entropy indicates a more uniform distribution over n -grams, i.e., less reliance on frequently repeated phrases.
Figure 8: Example 1 ( read-table-2 , pre-RL): A generic email/password form is injected before the task area with no reference to the actual task.
Figure 9: Example 2 ( draw-line , pre-RL): A minimal “System Verification Required” password prompt is overlaid on the task area.
Figure 10: Example 3 ( drag-items-grid , post-RL): The attacker references a task-specific item (“Matilde”) to craft a targeted, context-aware phishing message.
Figure 11: Example 4, Step 1 ( drag-items-grid , post-RL): The attacker injects a fake “Important Task Update” attempting to override the agent’s original drag instruction.
Figure 12: Example 4, Step 2 ( drag-items-grid , post-RL): The attacker injects a “Critical Task Error” referencing the specific drag target (“Margalit”) and requesting the agent’s personal information.
Full
w/o Oracle-Guided SFT
Checkpoint
ASR ↓
TSR ↑
ASR ↓
TSR ↑
Imitation
16.8±1.2
19.5±1.3
16.8±1.2
19.5±1.3
Oracle SFT
15.3±1.2
23.3±1.4
/
/
RL Iter 5
11.0±1.0
24.0±1.4
13.4±1.1
22.5±1.4
RL Iter 10
10.8±1.0
25.7±1.5
11.2±1.1
23.8±1.4
Appendix
Table 6: Full pipeline vs. ablated (w/o Oracle-Guided SFT) on unseen MiniWob++ (%) tasks. Subscripts denote standard errors.
Base 12B
DMAST 12B
Prompt
MW++
VWA
MW++
VWA
v1 (main paper)
15.3 / 7.4
3.1 / 8.2
25.0 / 4.5
8.2 / 7.2
v2 (manual)
16.9 / 7.4
7.2 / 28.9
24.0 / 5.0
7.2 / 17.5
v3 (teacher search)
17.3 / 3.1
7.1 / 22.4
26.4 / 2.9
8.2 / 13.3
Appendix
Table 7: Prompt-only defense at three strengths, applied to the base model and to DMAST, reported as TSR ↑ / ASR ↓ (%). MW++ denotes MiniWob++ evaluation tasks, excluded from model training; VWA denotes the curated VisualWebArena tasks.
Multimodal Large Language Model (MLLM)-based web agents provide practical, high-precision solutions for visual browser automation; however, they inherently expand the attack surface, introducing novel vision-based vulnerabilities. Existing adversarial evaluations targeting these agents frequently rely on permissive threat models and visually conspicuous artifacts. In this paper, we investigate a constrained vulnerability detection setting: a trusted web platform where the evaluator acts solely as an unprivileged third party, such as a merchant or advertiser, controlling only a semantically legitimate, spatially constrained region, such as an ad slot, a sponsored card, or a localized widget. Operating under these realistic constraints, we propose MIRAGE, a novel visual indirect prompt injection framework for targeted next-action hijacking. Our approach leverages diffusion models to generate perceptually benign adversarial images strictly confined to the attacker-controlled boundaries permitted by the trusted service provider. To maximize attack efficacy within such a restrictive setting, we introduce a robust optimization technique combining curvature-aware adversarial diffusion guidance with sparse, dark-pixel residual perturbations. Comprehensive evaluations against prominent MLLM web agent frameworks, specifically SeeAct and OpenClaw, empirically demonstrate the potency, realism, and stealth of our proposed MIRAGE.
Multimodal Large Language Models (MLLMs) are increasingly deployed for nuanced content safety and moderation tasks, yet they remain vulnerable to adversarial attacks and out-of-distribution edge cases. Traditional active learning and manual annotation fail to scale against the complexity and volume of novel multimodal threats. In this paper, we propose an automated, agentic red-teaming framework that systematically synthesizes difficult examples using an iterative strategy that proposes novel hypotheses as well as mutating on past attempts. Leveraging a multi-agent architecture that consists of a high-reasoning Architect agent, an advanced image generator, and a multi-level verification committee of LLM raters, our system autonomously uncovers boundary-pushing violations and ambiguous policy edge cases without any human intervention. By employing these carefully synthesized adversarial examples as in-context demonstrations via test-time Retrieval, we substantially improve the target model's robustness, reducing the False Negative Rate (FNR) from 41.2% to 24.5% in a public image safety benchmark without relying on any human labeling.
Large Vision-Language Models (LVLMs) have transformed multi-modal understanding, excelling in tasks like image captioning and visual question answering by integrating visual and textual inputs. However, their robustness against adversarial attacks, particularly those exploiting both modalities, remains underexplored, posing risks to critical applications like autonomous driving and content moderation. Existing attacks focus on single modalities or require impractical white-box access, limiting their real-world relevance. In this paper, we introduce Multi-Modal Adversarial Synergy, a groundbreaking framework that crafts universal, black-box multi-modal attacks against LVLMs. MMAS simultaneously generates a texture scale-constrained universal adversarial perturbation for images and a learnable prompt perturbation for text, optimized jointly using only model queries. The image perturbation leverages wavelet-based texture constraints to ensure imperceptibility and robustness across diverse visual inputs. The text perturbation, constrained by an L-norm in the embedding space, maintains semantic coherence while steering outputs toward a target. A novel cross-modal regularization term aligns the perturbations' gradient directions, enhancing their synergistic impact and transferability across tasks and models. Extensive experiments show the strong universal adversarial capabilities of our proposed attack with prevalent LVLMs.
Xiang Fang, Wanlong Fang, Changshuo Wang
School of Software Engineering, Huazhong University of Science and Technology · Nanyang Technological University, Singapore · University College London