Organizations: Tsinghua University · Huawei Noah’s Ark Lab · East China Normal University · Tongji University · Institute of Artificial Intelligence, Beihang University · National University of Singapore
Functional Web generation is increasingly optimized with executable rewards, yet existing methods largely focus on the quality of the final page and leave the process of diagnosing and repairing imperfect implementations underexplored. We identify a central challenge in this setting: the Generator and Refiner produce executable artifacts with direct environment rewards, whereas the intermediate Critic influences downstream behavior without a directly executable outcome. We introduce WebLoop, an execution-grounded framework that jointly learns generation, critique, and refinement within a shared policy. WebLoop trains an execution-free Critic with complementary signals for requirement-level discriminability and downstream helpfulness, first establishing reliable diagnosis and then introducing consequence-aware credit, while all three roles are jointly optimized with group-relative policy learning. With Qwen3.5-9B, WebLoop reaches 41.5 Overall on WebRise and 38.9% accuracy on WebGen-Bench, improving the base model by 11.3 and 15.4 points, respectively. The gains transfer to first-pass generation, persist at 27B scale, and generalize from text-only training to multimodal inputs. Controlled analyses further show that the improvement cannot be explained by an additional refinement pass alone, highlighting the importance of learning the Critic and the loop itself.
Figures & tables
Figure 1: Conceptual comparison of execution-based learning paradigms for functional Web generation. (a) Generation-only RL optimizes executable outputs directly. (b) Execution-guided RL feeds external feedback back into the code policy. (c) WebLoop jointly learns Generator, Critic, and Refiner, assigning execution-derived credit to the intermediate Critic.
Figure 2: Overview of our WebLoop framework. A shared policy sequentially acts as Generator, Critic, and Refiner. The Generator and Refiner produce executable Web pages evaluated in the same browser environment, while the execution-free Critic guides subsequent refinement and receives training credit from its diagnosis and downstream effects.
Table 1: Main results on WebRise and WebGen-Bench. General-purpose and Web-generation specialized models are included as capability references, while the Qwen3.5 block provides controlled comparisons. Bold denotes the best result within each reference group or controlled backbone size according to the metric direction. † denotes proprietary models.
Model / Stage
Text
Markdown
Sketch
Image
Video
Avg. ↑
T↑
Re↑
Ri↑
T↑
Re↑
Ri↑
T↑
Re↑
Ri↑
T↑
Re↑
Ri↑
T↑
Re↑
Ri↑
Base (Qwen3.5-9B)
23.2
45.3
22.1
28.2
51.1
26.0
25.2
48.8
21.9
23.0
46.2
22.8
24.7
45.7
23.9
31.9
WebLoop (Gen)
31.7
55.4
30.0
30.5
55.0
28.4
28.7
53.4
26.2
27.7
51.9
26.1
31.8
54.1
30.9
37.5
WebLoop (Refine)
34.6
57.9
32.0
34.4
57.5
32.1
31.1
55.0
28.1
30.6
53.1
29.0
35.9
57.3
33.6
40.1
Table 2: Cross-modal generalization on WebRise. WebLoop is trained only on Text-conditioned tasks and evaluated on all five input modalities using the same checkpoint. Avg. averages the modality-level (T+Re+Ri)/3 scores.
Training Scheme / Output
Optimized Roles
WebRise
WebGen-Bench
T↑
Re↑
Ri↑
Overall ↑
Yes (%) ↑
Partial (%) ↑
No (%) ↓
Start-Fail (%) ↓
Acc. ↑
Base
–
23.2
45.3
22.1
30.2
18.4
10.2
67.1
4.3
23.5
Generation Only (Gen)
Gen
29.5
54.1
27.3
36.9
17.8
10.5
67.9
3.9
23.0
Refine w/o Critique (Gen)
Gen + Refine
28.2
53.3
25.2
35.6
16.7
11.9
70.5
0.9
22.6
Refine w/o Critique (Refine)
32.9
57.3
29.7
40.0
21.6
10.0
68.3
0.0
26.7
Staged Training (Gen)
Gen → Critic + Refine
27.6
52.0
25.9
35.2
20.7
11.1
66.0
2.2
26.3
Table 3: Where does the gain come from? Controlled comparison of training schemes on WebRise and WebGen-Bench. All variants use the same Qwen3.5-9B backbone, training data, and rollout budget, and differ only in which roles are optimized and whether they are trained jointly or sequentially. Bold denotes the best result according to the metric direction.
Critic Training / Output
Rcrit
WebRise
WebGen-Bench
T↑
Re↑
Ri↑
Overall ↑
Yes (%) ↑
Partial (%) ↑
No (%) ↓
Start-Fail (%) ↓
Acc. ↑
Base
–
23.2
45.3
22.1
30.2
18.4
10.2
67.1
4.3
23.5
Discriminability Only (Gen)
D
30.0
53.6
27.2
36.9
22.3
8.7
65.8
3.2
26.6
Discriminability Only (Refine)
33.4
57.3
30.7
40.5
27.8
10.7
59.4
2.2
33.2
Helpfulness Only (Gen)
H
27.6
50.7
24.7
34.3
18.2
9.0
68.2
4.6
22.7
Helpfulness Only (Refine)
31.3
54.1
28.7
38.0
28.1
13.0
56.3
2.6
34.6
Table 4: Critic reward ablation. All variants use the same Qwen3.5-9B backbone, training data, and rollout structure and differ only in the Critic reward. D denotes discriminability and H denotes downstream helpfulness. Bold denotes the best result according to the metric direction.
Figure 3: Critic learning dynamics. (a) Discriminability D . (b) Downstream helpfulness H . (c) Generator and Refiner execution rewards. Thin and thick curves denote per-step values and moving averages; the dashed line marks the reward switch after 15 updates (around 1K samples).
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Comparison
#Samples
Acc. ↑
MAE ↓
Spearman ↑
Pearson ↑
κ↑
Human (A) vs. Human (B)
300
0.920
0.079
0.924
0.923
0.838
Evaluator vs. Human (A)
300
0.880
0.095
0.901
0.902
0.758
Evaluator vs. Human (B)
300
0.900
0.094
0.902
0.898
0.799
Appendix
Table 5: Human validation of functional evaluation. Agreement between the automatic evaluator and two independent human reviewers on the same 300 sampled interaction cases. Human–human agreement is included as a reference.
Figure 4: Refinement learning and Generator–Critic checkpoint interaction. (a) Refined-page quality over training, reported as improvement from the initial training point, with refinement without learned critique as a control. The dashed line marks the Critic reward switch. (b) WebRise Overall under cross-pairing Generator and Critic checkpoints from steps 20, 30, and 45. The Refiner uses the same checkpoint as the Generator in each pairing.
Model / Method
By Website Category
By Check Category
Content
User
Data
Functional
Data Display
Design
Presentation
Interaction
Management
Testing
Testing
Validation
General-Purpose Models
Qwen3.5-122B-A10B
21.3
9.6
10.6
5.2
18.0
27.0
Kimi-K2.6
42.5
26.4
28.4
20.9
37.9
49.6
Qwen3.5-397B-A17B
27.3
19.2
16.2
11.5
24.5
40.2
Appendix
Table 6: Fine-grained WebGen-Bench results. Accuracy under the same template-free setting as Table 1 . The 647 interaction checks are independently grouped by website category and check category; each grouping covers the full benchmark, so the six columns are not additive. † denotes proprietary models.
Figure 5: Critique length versus refinement quality. Each cell reports the mean WebRise Overall of refined pages within a critique-length bucket for a given Critic objective. WebLoop exhibits a stronger positive association, while alternative objectives show weaker or non-monotonic trends.
Figure 6: Successful critique-guided repair on Poll Vote (WebRise). The Critic localizes premature disclosure of voting results and a missing post identifier that blocks vote updates. The Refiner applies the corresponding fixes, changing both affected transition checks from FAIL to PASS.
Figure 7: Boundary-condition repair on File Search with Filters (WebRise). The Critic traces the empty result set to size bounds parsed as NaN after clearing filters. The Refiner adds guards for empty bounds, restoring all 12 files while preserving the already-correct control reset.
Figure 8: False-positive Critic judgment on Share Chat Link Interface (WebRise). The Critic attributes message-selection failure to an unsynchronized hidden checkbox, although the visible selection and preview are already controlled correctly. The Refiner synchronizes the checkbox without affecting either passing behavior.
Figure 9: Correct verdict but mislocalized repair guidance on Search Result Tabs (WebRise). The Critic correctly marks the per-category count requirement as unmet but targets a type-matching issue rather than the missing count labels. The Refiner applies the proposed edit, while the target predicate remains FAIL.
Figure 10: Divergent fault localization on Petition Progress Dashboard (WebRise). Both critiques correctly identify signing as broken, but only one localizes the state error that disables signing when the dialog opens. Although both suggested edits are implemented, only the correctly localized repair changes the transition from FAIL to PASS.
Figure 11: Incomplete repair guidance on Bulk Content Operations (WebRise). The Refiner adds the missing listener as recommended, but an unaddressed selector error still clears the header-checkbox state. The prescribed edit is implemented, yet the target requirement remains unsatisfied.
Figure 12: Divergent Refiner outcomes under equivalent diagnoses on Audit Timeline (WebRise). Both critiques identify the same missing event binding, but only one refinement implements the recommended listener and passes the transition check. The contrast isolates implementation fidelity from Critic diagnosis.
Figure 13: Refiner-induced regression on Topic Trending (WebRise). While addressing missing liked-state feedback, the Refiner conflates the user’s like state with the aggregate like count, causing the count-update predicate to regress from PASS to FAIL.
Figure 14: Generator prompt for first-pass Web generation.
Figure 15: Critic prompt for requirement-level diagnosis and repair guidance.
Figure 16: Refiner prompt for critique-conditioned Web repair.
State Key Lab of CAD&CG, Zhejiang University · Ant Group · State Key Lab of CAD&CG, Zhejiang University; Laboratory of Art and Archaeology Image (Zhejiang University), Ministry of Education, China