LLM-based Web agents can autonomously complete user tasks, yet deceptive interfaces can steer them toward outcomes that conflict with users' interests. Existing defenses primarily intervene on agent behavior through blocking, guidance, or replanning. We identify a distinct failure mode: a task-valid action can still realize an unauthorized consequence because of the current Web state. This motivates treating task-relevant Web state itself as a runtime control target. We introduce Veer, an agent-side runtime defense that leaves task planning to the base agent and intervenes on Web state when a proposed action would produce an unauthorized consequence. Before modifying the live environment, Veer constructs a prospective intervention trajectory toward a safe task-relevant state and executes it with runtime grounding and verification. Across TrickyArena and WebDecept, Veer achieves the highest safe task completion in all three evaluation settings, exceeding the next-best defense by 15.9 and 25.0 percentage points on TrickyArena-Single and TrickyArena-Multi, respectively, while reducing dark-pattern success on WebDecept to 0.3%. These gains persist across dark-pattern types and all 12 agent, model, and benchmark configurations. Ablations show that active state intervention provides the largest gain, while prospective rollout and temporal evidence contribute additional improvements. These results establish task-relevant Web state as an effective runtime control target for protecting Web agents from deceptive outcomes.
Figures & tables
Figure 1: Motivating example of a task-valid action whose consequence becomes unauthorized because of the current Web state.
Figure 2: Overview of Veer . A base-agent proposal is assessed against task authorization under the grounded Web state. Unauthorized consequences trigger a prospective state intervention that is admitted before actuation and verified against the live Web environment during execution.
Method
TrickyArena-Single
TrickyArena-Multi
WebDecept
DPSR ↓
TSR ↑
STC ↑
DPSR ↓
TSR ↑
STC ↑
DPSR ↓
TSR ↑
STC ↑
No-Defense
29.5
80.7
60.2
54.4
66.2
38.2
37.5
47.6
28.6
ICP
17.0
80.7
69.3
50.0
67.6
39.7
26.7
43.8
31.7
Guardrail
27.3
80.7
62.5
47.1
70.6
42.6
12.4
35.2
29.5
DUDE-S2
34.1
86.4
56.8
64.7
70.6
27.9
35.9
55.6
31.4
Spotlighting
26.1
79.5
62.5
58.8
60.3
26.5
35.6
44.8
27.3
Table 1: Overall effectiveness on TrickyArena and WebDecept (%). Lower DPSR and higher TSR/STC are better. Bold indicates the best value, and gray shading highlights Veer .
Method
DPSR ↓
TSR ↑
STC ↑
No-Defense
17.8
50.2
40.0
ICP
5.8
47.6
44.4
Guardrail
7.1
45.3
41.3
DUDE-S2
16.9
54.2
44.0
Spotlighting
17.3
47.6
38.2
VIGIL
3.1
8.9
7.6
Table 2: WebDecept results on the ground-truth-feasible subset ( N=225 ; %). Bold marks the best value; gray shading highlights Veer .
Agent
Model
Single
Multi
WebDecept
Default
GPT
85.2 (+25.0)
67.6 (+29.4)
41.0 (+12.4)
DPSK
68.2 (+12.5)
44.1 (+10.3)
40.0 (+4.8)
Codex
GPT
75.0 (+28.4)
54.4 (+25.0)
45.4 (+5.1)
DPSK
85.2 (+20.4)
55.9 (+14.7)
53.7 (+15.9)
Table 3: STC across agent–model configurations (%). Parentheses show gains over matched No-Defense.
Figure 3: Per-condition dark-pattern susceptibility on TrickyArena-Single and WebDecept. Color encodes DPSR (%; lower is better); Avg. and Worst report the mean and maximum across conditions.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Task-action budget
Veer internal budget
TrickyArena
30
8
WebDecept
15
8
Appendix
Table 4: Main interaction budgets. Veer’s internal budget is reserved for defense-side evidence and intervention operations.
Agent
Actor model
Veer model
Default
GPT-5.6-Luna
GPT-5.6-Luna
Default
DeepSeek-V4-Flash
DeepSeek-V4-Flash
Codex
GPT-5.6-Luna
GPT-5.6-Luna
Codex
DeepSeek-V4-Flash
DeepSeek-V4-Flash
Appendix
Table 6: Agent–model configurations evaluated in RQ2. Veer uses the same model family as the corresponding actor.
Scenario
Full
Feasible
Popup
45
45
Banner
45
45
P-popup
45
45
P-banner
45
45
Add-ons
45
45
Redirect
45
0
Appendix
Table 10: Construction of the feasibility-aware WebDecept subset.
Planned length K
Episodes
Trajectories
1
15
17
4
6
7
Total
21
24
Appendix
Table 20: Length of admitted prospective intervention trajectories.
Corrective operation
Executed actions
turn_off
20
decline
12
save
6
click
4
Total
42
Appendix
Table 21: Executed corrective operations in Full Veer.
Step
Operation
Expected state
Dependency
τ1
decline consent
declined
–
Appendix
Table 22: Representative consent correction for health_tos_80 .
Attempt
Corrective transition
Result
Initial
Open cookie options
Objective unresolved
Replan
Decline cookie consent
Verified
Appendix
Table 23: Bounded intervention replan for shop_p2_38 .