Web agents complete user requests by reading and acting on pages that third parties write, so an instruction planted on a page can redirect the agent away from the user's goal. The agent cannot simply ignore the page, because the page also holds the values and controls the task requires. Current defenses fine-tune the agent on injections fixed before training, and attackers that adapt to the trained model bypass them. Adversarial training lets the attacker adapt but keeps the tasks fixed, so a task stops teaching once the agent solves it. We introduce AdvSim2Real, which co-evolves a task curriculum, an injection adversary, and the agent inside a frozen web world model. The curriculum is rewarded for tasks the agent solves about half of the time, and the adversary only for a success flip, an injection that turns a judged success into a failure. Training in the simulator makes a 4B agent both more capable and more robust: its completion rises with and without attacks, holds against a frontier-model adversary it never trained against, and its capability gain carries over to a real browser. On 150 web tasks, AdvSim2Real raises completion under this unseen adversary by 33.6% relative to the base agent.
Figures & tables
Figure 1: AdvSim2Real makes the agent both more capable and more robust. (a) A curriculum proposes tasks (Stage 1) and an adversary inserts one timed injection (Stage 2) into pages that a frozen world model simulates. A frozen LLM judge rewards tasks solved about half of the time and success flips. (b) Judged completion on the 150 tasks over training, clean and under the learned adversaries. (c) Base model against the final checkpoint. Mean ± SD over rollout seeds (three; two for Kimi-K3 ).
Figure 2: The same injected notice diverts the base agent but not the trained one. Task 34 of the 150-task benchmark, Adv v2, seed 1; numbered markers give each agent’s actions in order, fields show final values, and pages are cropped. The base agent clicks the forbidden Reset all right after the notice appears and is judged bad; Robust iter 3 sees the notice before five actions, never clicks it, and is judged good.
Figure 3: One round of each stage of AdvSim2Real . Numbered steps summarize Algorithm 1 . Stage 1: the curriculum C proposes tasks, the executor E runs each several times in the world model W , and the judge J grades each run; C is rewarded for tasks E solves about half the time ( Equation 3 ). Stage 2: with C frozen, the adversary A is rewarded for success flips, good clean runs of E judged bad once replayed with A ’s injection ( Equation 5 ). Both stages train E on fresh rollouts ( Equation 4 ). Solid arrows: data; dashed: rewards that update the policy they enter; snowflakes: frozen.
Reactive page injection
Training
Executor
Clean
Adv v1
Adv v2
Adv v3
Mean
Initial
Base
74.89 ± 1.39
51.11 ± 2.34
48.00 ± 4.16
45.11 ± 1.54
48.07 ± 0.56
Stage 1
Capability iter 1
78.00 ± 1.76
57.11 ± 2.69
54.00 ± 2.91
53.11 ± 6.41
54.74 ± 3.11
Capability iter 2
77.11 ± 0.38
58.89 ± 1.39
53.11 ± 4.54
54.22 ± 1.92
55.41 ± 1.68
Capability iter 3
79.33 ± 1.76
57.11 ± 1.39
52.35 ± 4.12
53.56 ± 3.67
54.34 ± 2.93
Stage 1 + 2
Robust iter 1
77.33 ± 0.67
55.11 ± 2.14
52.67 ± 4.67
53.33 ± 2.40
53.70 ± 1.89
Table 1: Judged task completion in WebWorld-14B (%). Mean and sample standard deviation over three rollout seeds on the 150 tasks; Mean weights Adv v1–v3 equally, bold marks the best 4B value per column, and per-seed counts are in Section C.1 .
Figure 4: Where the gains come from, on all 150 tasks. (a) Completion after each Stage-2 round for the full pipeline and the Stage-2-only branch, clean (solid) and under Adv v1–v3 (dashed; Table 7 ). (b) Episodes with an injection request and without a final message, learned Adv v1–v3 against Kimi-K3 . (c) World-model verdict against strict browser outcome per task and seed. (d) Completion under Adv v1–v3 by skill stratum. Means over seeds with ±1 SD (range over cells in (b)); values in Section C.2 .
Training
Executor
Strict success
Correct fields
Initial
Base
25.56 ± 1.02
52.50 ± 3.14
Stage 1
Capability iter 1
31.78 ± 8.34
61.89 ± 7.45
Capability iter 2
43.56 ± 5.18
71.45 ± 4.97
Capability iter 3
44.44 ± 5.00
73.24 ± 2.24
Table 2: Sim-to-real transfer (%). Clean Chromium runs of the 150 tasks, mean and sample standard deviation over seeds. Strict success is a deterministic check of the submitted form; Correct fields counts the 746 target values per seed ( Section C.3 ).
Training
Executor
Clean
Kimi-K3
Initial
Base
74.89 ± 1.39
23.00 ± 3.30
Stage 1 + 2
Robust iter 1
77.33 ± 0.67
29.67 ± 4.24
Robust iter 2
78.00 ± 1.15
30.33 ± 0.47
Robust iter 3
81.33 ± 2.31
30.72 ± 3.04
Table 3: Completion under the Kimi-K3 adversary (%). Kimi-K3 replaces the learned adversary with the same prompt, observation, and one-injection budget; mean and sample standard deviation over seeds. Clean is the no-adversary column of Table 1 , repeated here for reference.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Skill stratum
Template
Tasks
High tier
Full form
Review
Ref. actions
Fields
Conditional policy
Return priority
29
5
11
0
4, 8
8, 12
Shipping priority
16
5
11
0
4, 8
8, 12
Ordered repair
Effective billing patch
22
11
17
22
5, 9
9, 13
Contact override
16
11
5
16
5, 9
9, 13
Relational join
Dated billing account
17
5
5
0
4, 8
8, 12
Ticket asset owner
10
5
5
0
4, 8
8, 12
Appendix
Table 4: Benchmark composition by template. Tier and burden columns count high-tier and full-form tasks within the template; Review counts tasks with a mandatory review-before-commit step; Ref. actions and Fields give the values that occur.
Figure 5: Clean completion by stratum, task changes per update, and clean wins lost under attack. (a) Clean judged completion (%) by skill stratum and checkpoint, mean over three seeds; task counts in parentheses. (b) Tasks whose number of good judgments over the three seeds increases or decreases after each update (ties omitted). (c) Share of clean-good task and seed pairs whose attacked episode is judged bad, Base against Robust iter 3.
World model
Clean
Adv v1
Adv v2
Adv v3
Base
114 110 113
77 80 73
77 65 74
65 69 69
Capability iter 1
114 119 118
88 88 81
84 76 83
90 71 78
Capability iter 2
115 116 116
89 90 86
72 85 82
83 83 78
Capability iter 3
121 120 116
88 84 85
85 † 76 74
86 75 80
Robust iter 1
117 116 115
85 79 84
84 82 71
79 84 77
Robust iter 2
118 118 115
90 96 93
84 84 77
77 80 81
Appendix
Table 5: Per-seed counts. Good episodes for seeds 0, 1, and 2, each out of 150 unless marked ( † 149 and ‡ 145 judged). Browser: strict and judge-good episodes, correct fields out of 746, episodes with a rejected action, and budget stops.
(b) Adversary
Adv v1
Adv v2
Adv v3
K3/Base
K3/R1
K3/R2
K3/R3
Injection requested
81.0
66.5
78.6
98.0
95.0
97.3
96.9
No final message
11.4
15.3
14.9
28.3
14.3
14.0
15.4
Requested marker visible
0.48
0.22
0.29
72.3
73.7
75.3
80.9
Appendix
Table 6: Values plotted in Figure 4 (%). (b) Learned adversaries pooled over seven checkpoints and three seeds; Kimi-K3 per executor. (c) World-model verdict against strict browser outcome per task and seed. (d) Completion under Adv v1–v3 by skill stratum, mean and sample standard deviation over three rollout seeds.
Figure 6: Strict browser success and failure modes of the capability checkpoints. (a) By skill stratum and (b) by tier and form: change in strict success from Base to Capability iter 3 (points), three seeds pooled. (c) Outcome of every episode, 450 per checkpoint; for Capability iter 1 and 2, wrong values 55.1% and 51.6% , rejected 2.2% and 2.0% , no submit 10.9% and 2.9% . (d) Capability iter 3: tasks (% of 150) by seeds judged good in the world model (rows) and seeds with strict browser success (columns).
Reactive page injection
Round
Executor
Clean
Adv v1
Adv v2
Adv v3
Mean
1
Robust iter 1
77.33 ± 0.67
55.11 ± 2.14
52.67 ± 4.67
53.33 ± 2.40
53.70 ± 1.89
Ablation iter 1
71.56 ± 2.04
60.22 ± 7.00
51.11 ± 4.54
47.11 ± 1.02
52.81 ± 3.34
Difference
+5.78
−5.11
+1.56
+6.22
+0.89
2
Robust iter 2
78.00 ± 1.15
62.00 ± 2.00
54.44 ± 2.69
52.89 ± 1.39
56.44 ± 1.15
Ablation iter 2
76.00 ± 3.53
58.00 ± 2.31
56.00 ± 2.67
56.89 ± 6.05
56.96 ± 3.19
Appendix
Table 7: Stage-1 removal ablation: judged completion (%). Each cell is the mean and sample standard deviation over three rollout seeds. Both branches run three Stage-2 rounds on the same 150 tasks and face the original with-Stage-1 Adv v1–v3 at evaluation; the Stage-2-only branch initializes both the executor and the frozen curriculum from the base model. Difference rows give the complete pipeline minus the Stage-2-only branch, computed from unrounded means.
Figure 7: Analytic difficulty shaping. (a) Curriculum shaping q(p) before the validity gate and repetition penalty of Equation 3 ; (b) executor scale f(p) , applied after group normalization, with floor 0.1 . Markers show the rates attainable with six trajectories; the curves are specified functions, not measurements.
Web agents can autonomously complete online tasks by interacting with websites, but their exposure to open web environments makes them vulnerable to prompt injection attacks embedded in HTML content or visual interfaces. Existing guard models still suffer from limited generalization to unseen domains and attack patterns, high false positive rates on benign content, reduced deployment efficiency due to added latency at each step, and vulnerability to adversarial attacks that evolve over time or directly target the guard itself. To address these limitations, we propose WARD (Web Agent Robust Defense against Prompt Injection), a practical guard model for secure and efficient web agents. WARD is built on WARD-Base, a large-scale dataset with around 177K samples collected from 719 high-traffic URLs and platforms, and WARD-PIG, a dedicated dataset designed for prompt injection attacks targeting the guard model. We further introduce A3T, an adaptive adversarial attack training framework that iteratively strengthens WARD through a memory-based attacker and guard co-evolution process. Extensive experiments show that WARD achieves nearly perfect recall on out-of-distribution benchmarks, maintains low false positive rates to preserve agent utility, remains robust against guard-targeted and adaptive attacks under substantial distribution shifts, and runs efficiently in parallel with the agent without introducing additional latency.
Tri Cao, Yulin Chen, Hieu Cao +8
National University of Singapore · University of Science · Vietnam National University, Ho Chi Minh City
Web-based agents powered by large language models are increasingly used for tasks such as email management or professional networking. Their reliance on dynamic web content, however, makes them vulnerable to prompt injection attacks: adversarial instructions hidden in interface elements that persuade the agent to divert from its original task. We introduce the Task-Redirecting Agent Persuasion Benchmark (TRAP), a benchmark for studying how persuasion techniques misguide autonomous web agents on realistic tasks. Across six frontier models, agents are susceptible to prompt injection in 25% of tasks on average (13% for GPT-5 to 43% for DeepSeek-R1), with small interface or contextual changes often doubling success rates and revealing systemic, psychologically driven vulnerabilities in web-based agents. We also provide a modular social-engineering injection framework with controlled experiments on high-fidelity website clones, allowing for further benchmark expansion.
Modern web agents built on large vision-language models process webpages, select relevant UI elements, and translate model outputs into browser actions. Existing visual red-teaming approaches use adversarial visual content to manipulate this process. However, they primarily target model inference and do not explicitly account for structured input processing or action post-processing. Consequently, model-level success does not establish control over browser execution and cannot reliably characterize end-to-end agent robustness. To address this gap, we formulate red teaming for vision-grounded web agents as an end-to-end grounding-to-execution problem, and introduce WebMirage, a framework that crafts localized visual perturbations that cause agents to select attacker-controlled content and execute the corresponding browser action across varying webpage renderings. It uses a role-slot abstraction and webpage recomposition to capture competition among webpage elements, and dataflow analysis to align optimization with action post-processing. We evaluate WebMirage across four agent configurations and six VLM backbones on 2,250 tasks covering 13 public websites and a sandbox benchmark. WebMirage achieves an average attack success rate of 91.9%, compared with 17.4% for the strongest baseline, and remains effective against three agent-level defenses.