Web agents promise to automate complex digital workflows, but their training remains limited by synthetic environments that look plausible while hiding broken links, inconsistent states, or infeasible tasks. We address the gap between scalable environment generation and trustworthy agent learning by constructing synthetic web environments that are executable, auditable, and grounded in backend state. Our framework represents each generated website as a structured scaffold of pages, navigation links, database records, state-change markers, and task constraints, then verifies and repairs structural, semantic, consistency, and feasibility defects before policy training. During interaction, ordinary UI transitions are executed deterministically, while persistent backend updates are invoked only through validated state-change markers, enabling dense rewards compiled from verified task-progress predicates. Across 500 synthetic environments spanning six domains, our method reduces task-blocking defects and improves feasible-task rate from 48.6% to 94.8%, while producing stronger PPO policies and improving transfer to WebArena, WebShop, and MiniWoB++ without LLM calls at evaluation time. These results show that verified synthetic environments can serve as a scalable and reliable training substrate for compact web agents, shifting synthetic webagent learning from surface-level plausibility toward executable, state-grounded supervision.
Figures & tables
Figure 1: Prior methods train on superficially plausible but inconsistent web interactions, while our approach verifies and repairs the environment to provide executable, state-grounded supervision for more effective agent policies.
Figure 2: Motivating diagnostics. (a) Raw LLM-generated environments contain frequent defects, with semantic and structural defects dominating the average total of 12.4 defects per environment. (b) Defect frequency and harmfulness differ: feasibility defects are less frequent but most likely to block task completion. (c) Marker-triggered updates are sparse across domains, indicating that most interactions are deterministic UI transitions. (d) The number of state-write calls remains small even for longer episodes.
Figure 3: Pipeline of verified synthetic web-environment construction and policy training. The offline stage generates a scaffold from a site specification, canonicalizes it into pages, data, tasks, and markers, and verifies/repairs navigation, data consistency, and workflow feasibility. The online stage trains a compact policy in the checked environment, where deterministic transitions and marker-validated state updates provide state-grounded rewards for PPO.
Figure 4: Verification–repair convergence. The repair loop rapidly removes defects and increases the fraction of tasks with bounded executable traces. The gains saturate after three to four iterations, supporting the use of targeted verification and repair rather than repeated full regeneration.
Method
Def. ↓
Block Def. ↓
Feasible ↑
Human SR ↑
State Viol. ↓
Time ↓
No Verification
12.4
4.9
48.6%
36.9%
18.7%
–
Rule-Based
9.8
4.1
53.2%
27.2%
14.6%
0.5m
Single-LLM
6.2
2.7
70.4%
54.3%
9.8%
12m
Self-Consistency
5.8
2.1
79.6%
69.1%
7.2%
35m
AutoGen
5.1
1.6
88.7%
89.2%
5.5%
28m
Ours
3.3
0.7
94.8%
91.4%
2.8%
18m
Table 1: Environment executability and fidelity. Verification should not merely reduce superficial defects; it should make generated tasks executable under consistent backend dynamics. Our method reduces both total defects and task-blocking defects.
Figure 5: Verified synthetic training improves compact policy learning. Raw environments provide noisy supervision, while verified environments make failures attributable to the policy. State-grounded dense rewards further accelerate learning and improve final success.
Figure 6: Component ablation matrix. Rows remove individual components and columns report metrics with 95% confidence intervals. Colors denote direction-corrected degradation. Feasibility verification, marker validation, and dense rewards drive executability, state fidelity, and policy learning.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Category
Typical failures
Detection criterion
Structural (ΔS)
Broken links, orphaned pages, missing navigation elements
The page graph is not connected from the entry page, a declared link has no target, or a required UI element is absent.
Semantic (ΔC)
Invalid prices, malformed dates, placeholder text, type mismatch
A rendered value violates the declared schema, expected field type, or page-level content semantics.
Consistency (ΔX)
Conflicting prices, inconsistent user attributes, mismatched entity states
The same entity attribute takes incompatible values across pages, views, or database bindings.
No bounded executable trace can satisfy the task completion constraint under the current scaffold.
Appendix
Table 2: Defect categories and detection criteria used in the preliminary diagnostics.
Type
Example
Detection
Structural
Link to /checkout returns 404
Graph + HTTP
Semantic
Price field shows “TBD” instead of $29.99
Schema + LLM
Consistency
Product 49onlist,59 on detail
Cross-page DB
Feasibility
“Add to cart” button not clickable
Task simulation
Appendix
Table 3: Defect Examples and Detection Complexity
Figure 7: Structural Verification Agent
Figure 8: Consistency Verification Agent
Figure 9: Cross-Page Verification Agent
Figure 10: Task Flow Verification Agent
Benchmark
# Tasks
Observation
Success Criterion
WebArena-compatible
180
Serialized DOM + task
Original programmatic evaluator
WebShop
500
Product-page DOM + task
Original purchase-match evaluator
MiniWoB++
560
DOM + task
Original environment reward
Appendix
Table 4: Transfer-evaluation protocol. We use a unified DOM-grounded interface for all methods. WebArena results are reported on the DOM-compatible subset described in the text; WebShop and MiniWoB++ use their original success evaluators.
Method
WebArena SR ↑
WebShop SR ↑
MiniWoB++ SR ↑
Eval LLM Calls ↓
WebArena Step ↓
GPT-4 Direct Prompting
10.6±2.1
32.5±2.7
41.2±2.9
17.5
17.5±0.6
GPT-4 ReAct
15.3±2.6
38.7±2.9
47.6±3.1
21.8
14.3±0.5
Small Policy, Synthetic BC
11.9±2.3
30.4±2.5
39.8±2.8
0.0
16.8±0.7
PPO on Raw Synthetic Env.
12.4±2.4
29.6±2.5
38.9±2.8
0.0
17.2±0.6
PPO on Verified Env. + Terminal
14.7±2.5
37.2±2.8
48.5±3.1
0.0
16.0±0.5
Ours
18.6±2.8
43.8±3.0
55.4±3.1
0.0
15.1±0.5
Appendix
Table 5: Transfer beyond synthetic environments. All methods are evaluated under the same DOM-grounded observation and action interface. The compact policy trained in verified environments transfers better than policies trained on raw synthetic environments, while requiring no LLM calls at evaluation time. The comparison to GPT-4 baselines should be interpreted under this constrained DOM-only interface, not as a claim of general superiority under native multimodal browser use.
Figure 11: Cost–fidelity tradeoff of simulators. Left: metric summary matrix. Each cell reports the mean and 95% confidence interval over rollout batches; color indicates metric-wise degradation after accounting for whether higher or lower is better. Right: rollout-level Pareto plot of token cost and state fidelity. Faint points denote individual rollout-batch observations; large markers denote simulator means with 95% confidence intervals, and marker size is proportional to rollout throughput. Our event-driven simulator lies near the practical Pareto frontier: it preserves fidelity close to step-wise LLM simulation while using far fewer tokens and achieving much higher throughput.
Figure 12: Defect impact on policy learning. Each point denotes one defect subtype measured across domains and random seeds. The x -axis reports how often the defect occurs in raw scaffolds, while the y -axis reports the PPO success-rate drop after injecting that defect into otherwise verified environments. Marker size indicates task-blocking rate and color denotes defect category. Feasibility and structural defects occupy the upper-impact region despite being less frequent than semantic defects. This shows that the most damaging defects are those that invalidate executable workflows or corrupt backend-grounded progress, rather than those that merely affect surface plausibility.
Figure 13: Alignment between dense reward and terminal success. Left: calibration curves between final progress score and empirical terminal success. A well-aligned dense reward should lie close to the diagonal. Middle: success rate by progress decile. Our state-grounded progress produces a monotonic success-lift pattern, whereas surface and LLM-based rewards assign high progress to many unsuccessful rollouts. Right: reward-alignment summary across domains, including AUROC, Spearman correlation, expected calibration error, and high-progress failure rate. State-grounded dense reward is both more predictive and better calibrated with terminal success, indicating that PPO receives intermediate supervision consistent with the true task objective.
Figure 14: Policy failure attribution. Left: outcome decomposition over all evaluated episodes. Raw synthetic environments contain many environment-induced failures, making PPO feedback noisy. Verification sharply reduces environment invalidity and state-update violations. Middle: failure attribution conditioned on failed episodes. After verification, most remaining failures become policy-attributable planning and grounding errors. Right: domain-level residual failures for our full method. Harder stateful domains such as banking and healthcare retain more planning and grounding failures, suggesting where future policy improvements are needed.
Method
Internal Feasible ↑
Held-out Feasible ↑
Author-Audited Exec. ↑
False Feasible ↓
Trace Agreement ↑
No Verification
48.6±1.4
46.1±1.6
44.8±2.4
12.4±1.5
83.2±2.1
Rule-Based
53.2±1.5
51.4±1.7
49.6±2.6
10.7±1.4
84.9±1.9
Single-LLM
70.4±1.3
67.8±1.5
65.9±2.2
6.8±1.1
89.4±1.6
Self-Consistency
79.6±1.1
77.2±1.3
75.4±2.0
5.2±0.9
91.1±1.4
AutoGen
88.7±0.9
86.9±1.0
84.8±1.9
3.9±0.7
93.8±1.1
Ours
94.8±0.7
92.9±0.9
91.7±1.8
2.3±0.6
96.4±0.8
Appendix
Table 6: Independent executability audit. The feasible-task improvement is confirmed by a held-out trace analyzer and author-audited replay. The low false-feasible rate indicates that our main feasible-task metric is not merely an artifact of the repair-time verifier.
Verifier
Precision ↑
Recall ↑
F1 ↑
Repair Success ↑
False Repair ↓
Rule-Based
96.1±1.2
39.4±2.1
56.0±2.0
71.8±2.8
1.2±0.4
Single-LLM
78.4±2.0
66.7±2.4
72.1±2.2
74.3±2.6
8.9±1.1
Self-Consistency
83.1±1.8
73.6±2.2
78.1±2.0
80.2±2.4
6.1±0.9
AutoGen
86.7±1.6
79.5±2.0
82.9±1.8
84.6±2.1
4.8±0.8
Ours
91.5±1.4
87.2±1.7
89.3±1.5
90.8±1.8
2.7±0.6
Appendix
Table 7: Verifier and repair reliability on an audited defect set. Rule-based checking is high precision but low recall. Single-pass LLM verification detects more defects but also produces more false repairs. Our coordinated verifier improves recall while keeping precision high and false repairs low.
Category
Precision ↑
Recall ↑
F1 ↑
Repair Success ↑
Symbolic
97.4
93.1
95.2
96.5
Structural
91.2
86.7
88.9
90.5
Semantic
88.6
84.2
86.3
87.1
Consistency
90.4
85.5
87.9
88.3
Feasibility
87.8
91.6
89.7
92.1
Marker
92.3
88.9
90.6
93.4
Appendix
Table 8: Category-level reliability of our verifier. Feasibility defects have slightly lower precision but higher recall, which is desirable because missed feasibility defects are especially harmful for policy learning.
Method
LLM Calls / Env. ↓
Tokens / Env. ↓
Time / Env. ↓
Def. ↓
Feasible ↑
False Feasible ↓
Rule-Based
0.0
0.0 K
0.5 m
9.8
53.2
10.7
Single-LLM
8.1
41.8 K
12.0 m
6.2
70.4
6.8
Self-Consistency
40.0
207.4 K
35.0 m
5.8
79.6
5.2
AutoGen
29.4
154.2 K
28.0 m
5.1
88.7
3.9
Ours-Token-Matched
8.3
42.5 K
11.5 m
4.7
87.4
3.4
Ours
22.6
116.4 K
18.0 m
3.3
94.8
2.3
Appendix
Table 9: LLM-budget comparison for environment verification. Our full method uses fewer tokens and less curation time than self-consistency and AutoGen while achieving higher executability. Even when matched to the Single-LLM token budget, our targeted coordination substantially improves feasible-task rate, suggesting that the gain is not merely due to spending more LLM calls.
Defect
Evidence in Raw Scaffold
Repair
Effect on Task
Broken navigation
Product page links to /cart , but the declared navigation graph has no reachable cart page.
Add the missing cart route and reconnect it to the product and checkout pages.
The bounded trace can reach the purchase workflow.
Inconsistent price
Product list shows the mouse as 24.99,whilethedetailpagebindsthesameentityto34.99.
Propagate the canonical database value $24.99 to all rendered views.
The price constraint in the task becomes evaluable and consistent.
Missing backend transition
The “Add to cart” button changes the visible page but does not update cart.items .
Attach an add-to-cart marker with typed arguments and write set {cart.items} .
Dense reward can credit verified cart insertion.
Unsatisfiable completion constraint
The completion predicate checks order.status=placed , but no checkout button writes this field.
Add a place-order marker and validate the state delta against order schema invariants.
Terminal success corresponds to an executable state transition.
Appendix
Table 10: Representative raw-to-verified repair example. The raw scaffold contains multiple locally plausible but globally task-blocking defects. Verification and repair convert the same task into an executable workflow with backend-grounded progress predicates.