Web agents complete user requests by reading and acting on pages that third parties write, so an instruction planted on a page can redirect the agent away from the user's goal. The agent cannot simply ignore the page, because the page also holds the values and controls the task requires. Current defenses fine-tune the agent on injections fixed before training, and attackers that adapt to the trained model bypass them. Adversarial training lets the attacker adapt but keeps the tasks fixed, so a task stops teaching once the agent solves it. We introduce AdvSim2Real, which co-evolves a task curriculum, an injection adversary, and the agent inside a frozen web world model. The curriculum is rewarded for tasks the agent solves about half of the time, and the adversary only for a success flip, an injection that turns a judged success into a failure. Training in the simulator makes a 4B agent both more capable and more robust: its completion rises with and without attacks, holds against a frontier-model adversary it never trained against, and its capability gain carries over to a real browser. On 150 web tasks, AdvSim2Real raises completion under this unseen adversary by 33.6% relative to the base agent.
Figures & tables
Figure 1: AdvSim2Real makes the agent both more capable and more robust. (a) A curriculum proposes tasks (Stage 1) and an adversary inserts one timed injection (Stage 2) into pages that a frozen world model simulates. A frozen LLM judge rewards tasks solved about half of the time and success flips. (b) Judged completion on the 150 tasks over training, clean and under the learned adversaries. (c) Base model against the final checkpoint. Mean ± SD over rollout seeds (three; two for Kimi-K3 ).
Figure 2: The same injected notice diverts the base agent but not the trained one. Task 34 of the 150-task benchmark, Adv v2, seed 1; numbered markers give each agent’s actions in order, fields show final values, and pages are cropped. The base agent clicks the forbidden Reset all right after the notice appears and is judged bad; Robust iter 3 sees the notice before five actions, never clicks it, and is judged good.
Figure 3: One round of each stage of AdvSim2Real . Numbered steps summarize Algorithm 1 . Stage 1: the curriculum C proposes tasks, the executor E runs each several times in the world model W , and the judge J grades each run; C is rewarded for tasks E solves about half the time ( Equation 3 ). Stage 2: with C frozen, the adversary A is rewarded for success flips, good clean runs of E judged bad once replayed with A ’s injection ( Equation 5 ). Both stages train E on fresh rollouts ( Equation 4 ). Solid arrows: data; dashed: rewards that update the policy they enter; snowflakes: frozen.
Reactive page injection
Training
Executor
Clean
Adv v1
Adv v2
Adv v3
Mean
Initial
Base
74.89 ± 1.39
51.11 ± 2.34
48.00 ± 4.16
45.11 ± 1.54
48.07 ± 0.56
Stage 1
Capability iter 1
78.00 ± 1.76
57.11 ± 2.69
54.00 ± 2.91
53.11 ± 6.41
54.74 ± 3.11
Capability iter 2
77.11 ± 0.38
58.89 ± 1.39
53.11 ± 4.54
54.22 ± 1.92
55.41 ± 1.68
Capability iter 3
79.33 ± 1.76
57.11 ± 1.39
52.35 ± 4.12
53.56 ± 3.67
54.34 ± 2.93
Stage 1 + 2
Robust iter 1
77.33 ± 0.67
55.11 ± 2.14
52.67 ± 4.67
53.33 ± 2.40
53.70 ± 1.89
Table 1: Judged task completion in WebWorld-14B (%). Mean and sample standard deviation over three rollout seeds on the 150 tasks; Mean weights Adv v1–v3 equally, bold marks the best 4B value per column, and per-seed counts are in Section C.1 .
Figure 4: Where the gains come from, on all 150 tasks. (a) Completion after each Stage-2 round for the full pipeline and the Stage-2-only branch, clean (solid) and under Adv v1–v3 (dashed; Table 7 ). (b) Episodes with an injection request and without a final message, learned Adv v1–v3 against Kimi-K3 . (c) World-model verdict against strict browser outcome per task and seed. (d) Completion under Adv v1–v3 by skill stratum. Means over seeds with ±1 SD (range over cells in (b)); values in Section C.2 .
Training
Executor
Strict success
Correct fields
Initial
Base
25.56 ± 1.02
52.50 ± 3.14
Stage 1
Capability iter 1
31.78 ± 8.34
61.89 ± 7.45
Capability iter 2
43.56 ± 5.18
71.45 ± 4.97
Capability iter 3
44.44 ± 5.00
73.24 ± 2.24
Table 2: Sim-to-real transfer (%). Clean Chromium runs of the 150 tasks, mean and sample standard deviation over seeds. Strict success is a deterministic check of the submitted form; Correct fields counts the 746 target values per seed ( Section C.3 ).
Training
Executor
Clean
Kimi-K3
Initial
Base
74.89 ± 1.39
23.00 ± 3.30
Stage 1 + 2
Robust iter 1
77.33 ± 0.67
29.67 ± 4.24
Robust iter 2
78.00 ± 1.15
30.33 ± 0.47
Robust iter 3
81.33 ± 2.31
30.72 ± 3.04
Table 3: Completion under the Kimi-K3 adversary (%). Kimi-K3 replaces the learned adversary with the same prompt, observation, and one-injection budget; mean and sample standard deviation over seeds. Clean is the no-adversary column of Table 1 , repeated here for reference.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Skill stratum
Template
Tasks
High tier
Full form
Review
Ref. actions
Fields
Conditional policy
Return priority
29
5
11
0
4, 8
8, 12
Shipping priority
16
5
11
0
4, 8
8, 12
Ordered repair
Effective billing patch
22
11
17
22
5, 9
9, 13
Contact override
16
11
5
16
5, 9
9, 13
Relational join
Dated billing account
17
5
5
0
4, 8
8, 12
Ticket asset owner
10
5
5
0
4, 8
8, 12
Appendix
Table 4: Benchmark composition by template. Tier and burden columns count high-tier and full-form tasks within the template; Review counts tasks with a mandatory review-before-commit step; Ref. actions and Fields give the values that occur.
Figure 5: Clean completion by stratum, task changes per update, and clean wins lost under attack. (a) Clean judged completion (%) by skill stratum and checkpoint, mean over three seeds; task counts in parentheses. (b) Tasks whose number of good judgments over the three seeds increases or decreases after each update (ties omitted). (c) Share of clean-good task and seed pairs whose attacked episode is judged bad, Base against Robust iter 3.
World model
Clean
Adv v1
Adv v2
Adv v3
Base
114 110 113
77 80 73
77 65 74
65 69 69
Capability iter 1
114 119 118
88 88 81
84 76 83
90 71 78
Capability iter 2
115 116 116
89 90 86
72 85 82
83 83 78
Capability iter 3
121 120 116
88 84 85
85 † 76 74
86 75 80
Robust iter 1
117 116 115
85 79 84
84 82 71
79 84 77
Robust iter 2
118 118 115
90 96 93
84 84 77
77 80 81
Appendix
Table 5: Per-seed counts. Good episodes for seeds 0, 1, and 2, each out of 150 unless marked ( † 149 and ‡ 145 judged). Browser: strict and judge-good episodes, correct fields out of 746, episodes with a rejected action, and budget stops.
(b) Adversary
Adv v1
Adv v2
Adv v3
K3/Base
K3/R1
K3/R2
K3/R3
Injection requested
81.0
66.5
78.6
98.0
95.0
97.3
96.9
No final message
11.4
15.3
14.9
28.3
14.3
14.0
15.4
Requested marker visible
0.48
0.22
0.29
72.3
73.7
75.3
80.9
Appendix
Table 6: Values plotted in Figure 4 (%). (b) Learned adversaries pooled over seven checkpoints and three seeds; Kimi-K3 per executor. (c) World-model verdict against strict browser outcome per task and seed. (d) Completion under Adv v1–v3 by skill stratum, mean and sample standard deviation over three rollout seeds.
Figure 6: Strict browser success and failure modes of the capability checkpoints. (a) By skill stratum and (b) by tier and form: change in strict success from Base to Capability iter 3 (points), three seeds pooled. (c) Outcome of every episode, 450 per checkpoint; for Capability iter 1 and 2, wrong values 55.1% and 51.6% , rejected 2.2% and 2.0% , no submit 10.9% and 2.9% . (d) Capability iter 3: tasks (% of 150) by seeds judged good in the world model (rows) and seeds with strict browser success (columns).
Reactive page injection
Round
Executor
Clean
Adv v1
Adv v2
Adv v3
Mean
1
Robust iter 1
77.33 ± 0.67
55.11 ± 2.14
52.67 ± 4.67
53.33 ± 2.40
53.70 ± 1.89
Ablation iter 1
71.56 ± 2.04
60.22 ± 7.00
51.11 ± 4.54
47.11 ± 1.02
52.81 ± 3.34
Difference
+5.78
−5.11
+1.56
+6.22
+0.89
2
Robust iter 2
78.00 ± 1.15
62.00 ± 2.00
54.44 ± 2.69
52.89 ± 1.39
56.44 ± 1.15
Ablation iter 2
76.00 ± 3.53
58.00 ± 2.31
56.00 ± 2.67
56.89 ± 6.05
56.96 ± 3.19
Appendix
Table 7: Stage-1 removal ablation: judged completion (%). Each cell is the mean and sample standard deviation over three rollout seeds. Both branches run three Stage-2 rounds on the same 150 tasks and face the original with-Stage-1 Adv v1–v3 at evaluation; the Stage-2-only branch initializes both the executor and the frozen curriculum from the base model. Difference rows give the complete pipeline minus the Stage-2-only branch, computed from unrounded means.
Figure 7: Analytic difficulty shaping. (a) Curriculum shaping q(p) before the validity gate and repetition penalty of Equation 3 ; (b) executor scale f(p) , applied after group normalization, with floor 0.1 . Markers show the rates attainable with six trajectories; the curves are specified functions, not measurements.