Modern web agents built on large vision-language models process webpages, select relevant UI elements, and translate model outputs into browser actions. Existing visual red-teaming approaches use adversarial visual content to manipulate this process. However, they primarily target model inference and do not explicitly account for structured input processing or action post-processing. Consequently, model-level success does not establish control over browser execution and cannot reliably characterize end-to-end agent robustness. To address this gap, we formulate red teaming for vision-grounded web agents as an end-to-end grounding-to-execution problem, and introduce WebMirage, a framework that crafts localized visual perturbations that cause agents to select attacker-controlled content and execute the corresponding browser action across varying webpage renderings. It uses a role-slot abstraction and webpage recomposition to capture competition among webpage elements, and dataflow analysis to align optimization with action post-processing. We evaluate WebMirage across four agent configurations and six VLM backbones on 2,250 tasks covering 13 public websites and a sandbox benchmark. WebMirage achieves an average attack success rate of 91.9%, compared with 17.4% for the strongest baseline, and remains effective against three agent-level defenses.
Figures & tables
Figure 1 : Vision-Grounded Web Agent Workflow. The agent operates in a loop: (1) constructs an observation from the screenshot and candidate elements (Environment), (2) reasoning over the visual observation (VLM Reasoning), (3) extracting a structured command from the raw output (Action Extraction), and (4) executing the action to update the state.
Figure 2 : Overview of WebMirage . The top figure denotes offline synthesis: ❶ identify role slots from the target webpage; ❷ recompose page instances by sampling companions and slot placements; ❸ optimize a perturbation so the attack content is consistently selected, with ❹ supervision aligned to the executed command via dataflow analysis. The bottom figure shows on benign pages the agent selects the best match; under attack, it selects the attack content across varying renderings.
Figure 3 : Role-Slot Abstraction. The red box highlights the same semantic item appearing in different compatible slots and alongside different surrounding content across renderings, illustrating how variability changes both placement and local competition at test time.
Figure 4 : Agent dataflow example. Raw model output is post-processed into a parsed prediction, and only this extracted representation is passed to action execution.
Agent
Method
ASR-S ( ↑ )
MTR ( ↓ )
SEEACT ( N=1,710 )
EIA
6.2%
7.3%
VWA-Adv
14.5%
5.0%
Chameleon
16.7%
2.8%
WebMirage
90.0%
1.8%
WEBVOYAGER ( N=1,710 )
EIA
5.0%
12.3%
VWA-Adv
15.6%
10.6%
Table 1 : Attack Effectiveness Comparison. We compare WebMirage against three baselines across four agent configurations.
Scenario
LLaVA v1.5
LLaVA v1.6
MiniCPM o
Phi-3 Vision
Qwen2 VL
Retail ( N=630 )
92.5
90.0
93.8
94.3
98.2
Accomm. ( N=350 )
92.0
92.4
95.5
96.7
95.5
Tutoring ( N=450 )
90.6
87.5
92.0
93.7
92.4
Home ( N=280 )
71.1
69.0
73.8
80.1
76.1
Avg. ( N=1710 )
88.4
86.4
90.4
92.3
92.5
Table 2 : Step Attack Success across Real-World Scenarios with SeeAct . N denotes the number of tasks; all values are percentages.
Scenario
LLaVA v1.5
LLaVA v1.6
MiniCPM o
Phi-3 Vision
Qwen2 VL
Retail ( N=630 )
90.0
86.6
92.4
93.6
94.6
Accomm. ( N=350 )
89.4
87.5
91.2
92.7
92.2
Tutoring ( N=450 )
86.2
84.7
88.6
90.2
89.5
Home ( N=280 )
68.5
65.3
69.4
75.5
73.9
Avg. ( N=1710 )
85.4
82.8
87.4
89.6
89.4
Table 3 : Multi-Step Attack Success across Real-World Scenarios with SeeAct . N denotes the number of tasks; all values are percentages.
LLaVA-v1.5
LLaVA-v1.6
MiniCPM-o
Phi-3-Vision
Qwen2-VL
Scenario
Method
Seen
Unseen
Seen
Unseen
Seen
Unseen
Seen
Unseen
Seen
Unseen
Retail
VWA-Adv
14.5%
0.0%
11.0%
1.5%
16.0%
0.0%
16.8%
2.5%
16.5%
1.7%
Chameleon
17.8%
7.1%
15.5%
2.7%
16.9%
5.6%
19.2%
6.6%
18.6%
7.5%
WebMirage
92.5%
90.7%
90.0%
85.2%
93.8%
91.5%
94.3%
92.7%
98.2%
95.1%
Accommodation
VWA-Adv
15.5%
2.5%
10.8%
0.0%
15.4%
1.8%
17.9%
5.5%
19.7%
4.2%
Chameleon
19.5%
9.0%
15.9%
6.9%
19.0%
9.2%
18.6%
7.4%
20.0%
12.5%
Table 4 : Robustness to Webpage Variability. We report ASR-S across seen and unseen renderings.
Figure 5 : Controlled Analysis of Webpage Variability. ASR-S on Retail tasks under companion variation, positional shifts, and their combination. Combined results are averaged over companion levels.
Attack Success Rate (ASR)
Agent
Site
Original
Rewritten
Δ ASR
Public Websites
SeeAct
Amazon
96.7%
91.2%
-5.7%
Walmart
100.0%
95.5%
-4.5%
Target
91.7%
88.8%
-3.2%
WebVoyager
Amazon
93.6%
90.5%
-3.3%
Table 5 : Robustness to Semantic Query Variations.
NLL ↓
Source Model
Target Model
ASR-S
Clean
Perturbed
Δ NLL (%)
Within-lineage transfer
MiniCPM-o
MiniCPM-V 2.5 [ 38 ]
92.5%
1.253
0.113
-90.1%
MiniCPM-V 2.6
89.3%
1.305
0.149
-88.6%
Phi-3 Vision
Phi-3.5 Vision
88.3%
2.252
0.206
-90.9%
CogVLM
CogAgent [ 39 ]
93.5%
1.730
0.115
-93.4%
Table 6 : Transferability of WebMirage . Top : within-lineage transfer (e.g., version upgrades or fine-tuned variants). Bottom : cross-architecture transfer (different architectures). NLL : negative log-likelihood of the attacker-intended executed command (lower is better). Δ NLL : relative change from clean to perturbed input.
Figure 6 : Impact of perturbation budget across VLMs
Agent
Target
ASR-S
Epochs
Target Len.
SeeAct
Raw output
74.5%
1500
59
Exec-aligned
90.0%
500
16
WebVoyager
Raw output
63.5%
1500
65
Exec-aligned
85.4%
750
6
VWA (AcTree)
Raw output
70.8%
1500
125
Exec-aligned
96.2%
245
8
Table 7 : Action-Level Supervision via Dataflow Analysis. Epochs : average optimization epochs per agent, capped at 1500. Target Len. : average number of supervision tokens.
Defense
Acc. (%) ↑
ASR-S (%) ↓
No countermeasure
100.0
91.9
Image-Level Sanitization
JPEG compression (quality =75 )
89.3
75.2
Gaussian blur (radius =1.0 )
100.0
82.5
Uniform noise ( ϵdef=8/255 )
90.2
81.4
Uniform noise ( ϵdef=16/255 )
58.5
45.7
Table 8 : Attack Performance under Image- and Agent-Level Countermeasures. All attacks use an ℓ∞ budget of 16/255 . Acc. denotes clean-task accuracy.
Table 10 : Evaluation Websites. Websites grouped by scenario.
Scenario
Total
Failure Cause Breakdown
Network
Popups
Auth/Login
Context Drift
Retail
73
8
46
11
8
Accomm.
67
12
30
22
3
Home
49
3
3
37
6
Tutoring
76
18
5
45
8
Total
265
41
84
115
25
Appendix
Table 11 : Multi-Step Failure Analysis. Failure causes among the 265 runs in which the targeted selection succeeded but the multi-step task failed. Results aggregated across five VLM backbones.
ASR-S (%)
VLM Backbone
2%
5%
10%
LLaVA-v1.5
61.9
84.5
98.0
LLaVA-v1.6
52.4
78.3
93.5
MiniCPM-o
67.0
85.9
100.0
Phi-3-Vision
68.2
90.4
100.0
Qwen2-VL
64.7
87.5
95.3
Appendix
Table 12 : Effect of Target Visual Footprint. ASR-S at three target-to-screenshot area ratios.
Figure 8 : Walmart page rendering with an ϵ=8/255 perturbation; the red box marks the target image.
Figure 9 : Walmart page rendering with an ϵ=16/255 perturbation; the red box marks the target image.
Multimodal Large Language Model (MLLM)-based web agents provide practical, high-precision solutions for visual browser automation; however, they inherently expand the attack surface, introducing novel vision-based vulnerabilities. Existing adversarial evaluations targeting these agents frequently rely on permissive threat models and visually conspicuous artifacts. In this paper, we investigate a constrained vulnerability detection setting: a trusted web platform where the evaluator acts solely as an unprivileged third party, such as a merchant or advertiser, controlling only a semantically legitimate, spatially constrained region, such as an ad slot, a sponsored card, or a localized widget. Operating under these realistic constraints, we propose MIRAGE, a novel visual indirect prompt injection framework for targeted next-action hijacking. Our approach leverages diffusion models to generate perceptually benign adversarial images strictly confined to the attacker-controlled boundaries permitted by the trusted service provider. To maximize attack efficacy within such a restrictive setting, we introduce a robust optimization technique combining curvature-aware adversarial diffusion guidance with sparse, dark-pixel residual perturbations. Comprehensive evaluations against prominent MLLM web agent frameworks, specifically SeeAct and OpenClaw, empirically demonstrate the potency, realism, and stealth of our proposed MIRAGE.
Web agents complete user requests by reading and acting on pages that third parties write, so an instruction planted on a page can redirect the agent away from the user's goal. The agent cannot simply ignore the page, because the page also holds the values and controls the task requires. Current defenses fine-tune the agent on injections fixed before training, and attackers that adapt to the trained model bypass them. Adversarial training lets the attacker adapt but keeps the tasks fixed, so a task stops teaching once the agent solves it. We introduce AdvSim2Real, which co-evolves a task curriculum, an injection adversary, and the agent inside a frozen web world model. The curriculum is rewarded for tasks the agent solves about half of the time, and the adversary only for a success flip, an injection that turns a judged success into a failure. Training in the simulator makes a 4B agent both more capable and more robust: its completion rises with and without attacks, holds against a frontier-model adversary it never trained against, and its capability gain carries over to a real browser. On 150 web tasks, AdvSim2Real raises completion under this unseen adversary by 33.6% relative to the base agent.
Sarim Hashmi, Mukul Ranjan, Kshitij Mishra +3
Mohamed bin Zayed University of Artificial Intelligence · Amazon · Massachusetts Institute of Technology
Existing red-teaming studies on GUI agents face two fundamental limitations: adversarial perturbations require white-box access unavailable in commercial deployments, while prompt injection is increasingly neutralized by stronger safety alignment. To study robustness under a more practical threat model, we propose Semantic-level UI Element Injection, a black-box red-teaming paradigm that overlays safety-aligned and harmless UI elements onto screenshots to misdirect the agent's visual grounding. Our method couples a modular Editor--Overlapper--Victim pipeline with iterative search that samples multiple candidate edits, keeps the best cumulative overlay, and adapts future prompt strategies based on previous failures. Experiments across 19 victim models spanning 8 model families show that strategic optimization substantially outperforms random injection (3.5-6.9x on the most robust victims) and transfers near-perfectly across architectures, confirming model-agnostic visual-semantic vulnerabilities. After the first successful attack, the victim still clicks the attacker-controlled icon in over 15% of subsequent independent trials versus below 1% for random injection, establishing that strategically placed icons act as persistent attractors that causally redirect grounding rather than introducing incidental clutter.
Wenkui Yang, Chao Jin, Haisu Zhu +7
School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences · 2MAIS&NLPR, Institute of Automation, Chinese Academy of Sciences · 4ShanghaiTech University +2