cs.CLOct 8, 2026

Learning the Loop, Not Just the Page: Execution-Grounded Loop Learning for Web Generation

Authors: Yuxin Meng, Ruixu Zhang, Junjie Wang, Yuhan Suo, Yuhan Sun, Ruining Hu, Yiyao Yu, Yubin Wang, +5 more

Organizations: Tsinghua University · Huawei Noah’s Ark Lab · East China Normal University · Tongji University · Institute of Artificial Intelligence, Beihang University · National University of Singapore

Abstract

Functional Web generation is increasingly optimized with executable rewards, yet existing methods largely focus on the quality of the final page and leave the process of diagnosing and repairing imperfect implementations underexplored. We identify a central challenge in this setting: the Generator and Refiner produce executable artifacts with direct environment rewards, whereas the intermediate Critic influences downstream behavior without a directly executable outcome. We introduce WebLoop, an execution-grounded framework that jointly learns generation, critique, and refinement within a shared policy. WebLoop trains an execution-free Critic with complementary signals for requirement-level discriminability and downstream helpfulness, first establishing reliable diagnosis and then introducing consequence-aware credit, while all three roles are jointly optimized with group-relative policy learning. With Qwen3.5-9B, WebLoop reaches 41.5 Overall on WebRise and 38.9% accuracy on WebGen-Bench, improving the base model by 11.3 and 15.4 points, respectively. The gains transfer to first-pass generation, persist at 27B scale, and generalize from text-only training to multimodal inputs. Controlled analyses further show that the improvement cannot be explained by an additional refinement pass alone, highlighting the importance of learning the Critic and the loop itself.

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. WebGen-R1: Incentivizing Large Language Models to Generate Functional and Aesthetic Websites with Reinforcement Learning

    Apr 22, 2026Juyong Jiang, Chenglin Cai, Chansung Park +4RL for Code GenerationRL for Language Models

  2. WebGrader: Training LLMs for Web Development with Self-Evolving Programmatic Grader

    Aug 6, 2026Boshui Chen, Huiping Liu, Shaolei ZhangRL for Code GenerationAutomated Software Testing

  3. LiveEvalBench: Toward Open-World Evaluation for Web Generation

    Aug 4, 2026Yiyao Wang, Zhen Wen, Yinghao Tang +5Web Application GenerationAutomated Evaluation