cs.AISep 28, 2026

AUV-Bench: Aesthetic Understanding and Generation Evaluation for User Interfaces

Authors: Zhijie Deng, Ling Li, Junhao Ji, Siwei Lyu, Zhipeng Xu, Zulong Chen, Rongyao Fang, Shuai Bai, +2 more

Organizations: The Hong Kong University of Science and Technology (Guangzhou) · Alibaba Group · The Hong Kong University of Science and Technology

Abstract

Multimodal foundation models are increasingly used for evaluating and generating user interfaces (UIs), often producing seemingly reasonable aesthetic judgments and visually plausible pages. However, under professional design scrutiny, their behavior can differ substantially from that of human designers. In professional design practice, designers rely on a systematic set of aesthetic principles that consistently guide judgment, diagnosis, repair, and creation. A coherent aesthetic capability should therefore connect aesthetic judgment with design actions. Existing evaluations, however, typically assess these abilities in isolation, making it difficult to determine whether task-level success reflects a shared aesthetic understanding or merely fragmented task-specific competence. To address this gap, we introduce AUV-Bench, developed in collaboration with professional UI designers around 1,395 executable web interfaces and four tasks: aesthetic scoring, diagnosis, repair, and text-to-UI generation. The tasks share a pool of UIs and aesthetic principles, with diagnosis and repair further aligned on 660 controlled-degradation instances to enable instance-level analysis of judgment and action. Evaluation of 12 models reveals a capability imbalance: models show moderate agreement with professional designers in holistic aesthetic scoring, yet exact diagnosis-chain success peaks at only 24.7%. On the aligned diagnosis-repair cases, correct judgments and successful repairs do not consistently coincide, exposing a Judgment-Action Gap between identifying aesthetic problems and successfully acting on them. In open-ended generation, even leading models achieve only moderate aesthetic quality under human-calibrated evaluation. Overall, current models exhibit partial aesthetic competence, but still lack the fine-grained understanding and judgment-action coherence required for reliable UI design.

Figures & tables

Appendix figures & tables40 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. FlowEval: Reference-based Evaluation of Generated User Interfaces

    May 5, 2026Jason Wu, Priyan Vaithilingam, Eldon Schoop +2Generative User InterfaceInteraction Design

  2. Design Theater: Evaluating the Gap Between User-Facing Design Reasoning and Implementation in Generative UI Tools

    Jul 24, 2026Kashif Imteyaz, Kaif Imteyaz, Nakul Rajpal +3Generative User InterfaceInteraction Design

  3. Beyond a Single Judge: The Evidence-Grounded, Social-Weighted Persona Panel for Generative UI Evaluation

    Jul 30, 2026Zheng Wu, Yibo Luo, Pu Zhang +2Generative User InterfacePersona Consistency