Learning the Loop, Not Just the Page: Execution-Grounded Loop Learning for Web Generation
Organizations: Tsinghua University · Huawei Noah’s Ark Lab · East China Normal University · Tongji University · Institute of Artificial Intelligence, Beihang University · National University of Singapore
Abstract
Functional Web generation is increasingly optimized with executable rewards, yet existing methods largely focus on the quality of the final page and leave the process of diagnosing and repairing imperfect implementations underexplored. We identify a central challenge in this setting: the Generator and Refiner produce executable artifacts with direct environment rewards, whereas the intermediate Critic influences downstream behavior without a directly executable outcome. We introduce WebLoop, an execution-grounded framework that jointly learns generation, critique, and refinement within a shared policy. WebLoop trains an execution-free Critic with complementary signals for requirement-level discriminability and downstream helpfulness, first establishing reliable diagnosis and then introducing consequence-aware credit, while all three roles are jointly optimized with group-relative policy learning. With Qwen3.5-9B, WebLoop reaches 41.5 Overall on WebRise and 38.9% accuracy on WebGen-Bench, improving the base model by 11.3 and 15.4 points, respectively. The gains transfer to first-pass generation, persist at 27B scale, and generalize from text-only training to multimodal inputs. Controlled analyses further show that the improvement cannot be explained by an additional refinement pass alone, highlighting the importance of learning the Critic and the loop itself.
Figures & tables
| Model / Method | WebRise | WebGen-Bench | |||||||
| Overall | Yes (%) | Partial (%) | No (%) | Start-Fail (%) | Acc. | ||||
| General-Purpose Models (Open-Source & Proprietary) | |||||||||
| Qwen3.5-122B-A10B | 38.0 | 48.9 | 35.2 | 40.7 | 9.9 | 6.2 | 78.1 | 5.9 | 13.0 |
| Kimi-K2.6 | 44.6 | 54.2 | 41.8 | 46.9 | 25.7 | 11.1 | 55.3 | 7.9 | 31.2 |
| Qwen3.5-397B-A17B | 45.7 | 57.2 | 42.8 | 48.6 | 16.1 | 9.1 | 70.0 | 4.8 | 20.6 |
| GLM-5.3 | 59.5 | 77.1 | 54.9 | 63.8 | 20.6 | 11.9 | 67.5 | 0.0 | 26.5 |
| Model / Stage | Text | Markdown | Sketch | Image | Video | Avg. | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Base (Qwen3.5-9B) | 23.2 | 45.3 | 22.1 | 28.2 | 51.1 | 26.0 | 25.2 | 48.8 | 21.9 | 23.0 | 46.2 | 22.8 | 24.7 | 45.7 | 23.9 | 31.9 |
| WebLoop (Gen) | 31.7 | 55.4 | 30.0 | 30.5 | 55.0 | 28.4 | 28.7 | 53.4 | 26.2 | 27.7 | 51.9 | 26.1 | 31.8 | 54.1 | 30.9 | 37.5 |
| WebLoop (Refine) | 34.6 | 57.9 | 32.0 | 34.4 | 57.5 | 32.1 | 31.1 | 55.0 | 28.1 | 30.6 | 53.1 | 29.0 | 35.9 | 57.3 | 33.6 | 40.1 |
| Training Scheme / Output | Optimized Roles | WebRise | WebGen-Bench | |||||||
| Overall | Yes (%) | Partial (%) | No (%) | Start-Fail (%) | Acc. | |||||
| Base | – | 23.2 | 45.3 | 22.1 | 30.2 | 18.4 | 10.2 | 67.1 | 4.3 | 23.5 |
| Generation Only (Gen) | Gen | 29.5 | 54.1 | 27.3 | 36.9 | 17.8 | 10.5 | 67.9 | 3.9 | 23.0 |
| Refine w/o Critique (Gen) | Gen + Refine | 28.2 | 53.3 | 25.2 | 35.6 | 16.7 | 11.9 | 70.5 | 0.9 | 22.6 |
| Refine w/o Critique (Refine) | 32.9 | 57.3 | 29.7 | 40.0 | 21.6 | 10.0 | 68.3 | 0.0 | 26.7 | |
| Staged Training (Gen) | Gen Critic + Refine | 27.6 | 52.0 | 25.9 | 35.2 | 20.7 | 11.1 | 66.0 | 2.2 | 26.3 |
| Critic Training / Output | WebRise | WebGen-Bench | ||||||||
| Overall | Yes (%) | Partial (%) | No (%) | Start-Fail (%) | Acc. | |||||
| Base | – | 23.2 | 45.3 | 22.1 | 30.2 | 18.4 | 10.2 | 67.1 | 4.3 | 23.5 |
| Discriminability Only (Gen) | 30.0 | 53.6 | 27.2 | 36.9 | 22.3 | 8.7 | 65.8 | 3.2 | 26.6 | |
| Discriminability Only (Refine) | 33.4 | 57.3 | 30.7 | 40.5 | 27.8 | 10.7 | 59.4 | 2.2 | 33.2 | |
| Helpfulness Only (Gen) | 27.6 | 50.7 | 24.7 | 34.3 | 18.2 | 9.0 | 68.2 | 4.6 | 22.7 | |
| Helpfulness Only (Refine) | 31.3 | 54.1 | 28.7 | 38.0 | 28.1 | 13.0 | 56.3 | 2.6 | 34.6 | |
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
| Comparison | #Samples | Acc. | MAE | Spearman | Pearson | |
|---|---|---|---|---|---|---|
| Human (A) vs. Human (B) | 300 | 0.920 | 0.079 | 0.924 | 0.923 | 0.838 |
| Evaluator vs. Human (A) | 300 | 0.880 | 0.095 | 0.901 | 0.902 | 0.758 |
| Evaluator vs. Human (B) | 300 | 0.900 | 0.094 | 0.902 | 0.898 | 0.799 |
| Model / Method | By Website Category | By Check Category | ||||
|---|---|---|---|---|---|---|
| Content | User | Data | Functional | Data Display | Design | |
| Presentation | Interaction | Management | Testing | Testing | Validation | |
| General-Purpose Models | ||||||
| Qwen3.5-122B-A10B | 21.3 | 9.6 | 10.6 | 5.2 | 18.0 | 27.0 |
| Kimi-K2.6 | 42.5 | 26.4 | 28.4 | 20.9 | 37.9 | 49.6 |
| Qwen3.5-397B-A17B | 27.3 | 19.2 | 16.2 | 11.5 | 24.5 | 40.2 |