AUV-Bench: Aesthetic Understanding and Generation Evaluation for User Interfaces
Organizations: The Hong Kong University of Science and Technology (Guangzhou) · Alibaba Group · The Hong Kong University of Science and Technology
Abstract
Multimodal foundation models are increasingly used for evaluating and generating user interfaces (UIs), often producing seemingly reasonable aesthetic judgments and visually plausible pages. However, under professional design scrutiny, their behavior can differ substantially from that of human designers. In professional design practice, designers rely on a systematic set of aesthetic principles that consistently guide judgment, diagnosis, repair, and creation. A coherent aesthetic capability should therefore connect aesthetic judgment with design actions. Existing evaluations, however, typically assess these abilities in isolation, making it difficult to determine whether task-level success reflects a shared aesthetic understanding or merely fragmented task-specific competence. To address this gap, we introduce AUV-Bench, developed in collaboration with professional UI designers around 1,395 executable web interfaces and four tasks: aesthetic scoring, diagnosis, repair, and text-to-UI generation. The tasks share a pool of UIs and aesthetic principles, with diagnosis and repair further aligned on 660 controlled-degradation instances to enable instance-level analysis of judgment and action. Evaluation of 12 models reveals a capability imbalance: models show moderate agreement with professional designers in holistic aesthetic scoring, yet exact diagnosis-chain success peaks at only 24.7%. On the aligned diagnosis-repair cases, correct judgments and successful repairs do not consistently coincide, exposing a Judgment-Action Gap between identifying aesthetic problems and successfully acting on them. In open-ended generation, even leading models achieve only moderate aesthetic quality under human-calibrated evaluation. Overall, current models exhibit partial aesthetic competence, but still lack the fine-grained understanding and judgment-action coherence required for reliable UI design.
Figures & tables
| Benchmark | Benchmark Setup | Principle Grounding | Core Capabilities | Judgment–Action Linkage | Statistics | ||||||
| Executable UI | Expert Aesthetic GT | Principle- Grounded | Controlled Violations | Aesthetic Scoring | Aesthetic Diagnosis | Aesthetic Repair | Text-to-UI Generation | Domain | Scale | ||
| Benchmarks for UI Generation | |||||||||||
| Design2Code ( Si et al., 2025 ) | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | Web UI | 484 pages |
| WebUIBench ( Lin et al., 2025 ) | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | Web UI | 21K QA / 0.7K+ sites |
| WebMMU ( Awal et al., 2025 ) | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | Web UI | 8.1K tasks / 2,059 pages |
| WebGen-Bench ( Lu et al., 2026 ) | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | Web App | 647 tests |
| Model | Aesthetic Scoring | Aesthetic Diagnosis | Aesthetic Repair | Text-to-UI Generation | Judgment–Action Association ( ) | |||||
| SRCC | MAE | Det. BAcc | Attr. Macro-J | Loc. F1@0.5 | ECS | Repair Pass | Cal. Aes. Judge | Pairwise WR | ||
| Frontier Models | ||||||||||
| Claude Opus 5 | 0.56 | 0.81 | 91.10 | 58.90 | 36.70 | 24.70 | 75.60 | 3.13 | 71.50 | 0.32 |
| Doubao-Seed-2.1-Pro | 0.50 | 1.10 | 79.60 | 36.20 | 7.60 | 4.10 | 45.00 | 2.78 | 62.90 | 0.31 |
| GPT-5.6 Sol | 0.55 | 0.90 | 87.60 | 55.40 | 28.20 | 19.40 | 65.50 | 3.17 | 78.00 | 0.38 |
| Grok 4.6 | 0.54 | 0.95 | 81.20 | 42.70 | 6.50 | 2.90 | 70.00 | 3.08 | 62.50 | 0.25 |
Appendix figures & tables40 assets
Supplementary material from the paper’s appendix.
Appendix
| Dimension | Rule | Implementation |
| color | Primary : secondary : accent | Measure the pixel proportion of each color role and compute its KL divergence from the target ratio |
| Normal text/background contrast ; large text ; graphics and UI components | Compute the WCAG contrast ratio from foreground and background relative luminances using Eq. ( 6 ) | |
| Spacing | Element sizes and spacing use , , or integer multiples of pixels | Compute the residual of element coordinates and dimensions modulo , while admitting and as additional values |
| Typography | At most one CJK and one Latin typeface per page | Extract typeface information from the design source and compare typeface usage |
| and at most five type sizes | Extract font-size and line-height parameters from the design source | |
| Font weights use and at most three weights per application | Extract font-weight values from the design source |
| Metric | Definition and properties |
| PSNR | Peak signal-to-noise ratio is a full-reference image-quality metric that compares the maximum possible signal power with the power of corrupting noise. It is measured in decibels, and larger values indicate less distortion. Because it is based on pixel-level error, it does not explicitly model the visual characteristics of human perception. |
| SSIM | Structural similarity evaluates image similarity from luminance, contrast, and structure ( Wang et al., 2004 ) . Its value lies in , with larger values indicating less distortion. In practical computation, an image can be divided into local windows. For windows, the average structural similarity is |
| IFC | The information fidelity criterion evaluates image quality using natural scene statistics and characteristics of the human visual system. It measures the mutual information between a test image and a reference image ( Sheikh et al., 2005 ) . |
| VIF | Visual information fidelity extends the information-based formulation of IFC and focuses on the amount of visual information lost between a test image and its reference ( Sheikh and Bovik, 2006 ) . |
| MSE / RMSE | Mean squared error and root mean squared error measure pixel-level error between corresponding images. They are simple objective measures but do not explicitly account for characteristics of human visual perception. |
| Level | Typical components | Offset | Blur | color |
| 0 | Buttons, inputs, search fields, tags, tables, links, pagination, steps, breadcrumbs, switches, radio buttons, checkboxes, and progress bars | rgba(0,0,0,0) | ||
| 1 | Navigation and card hover states | px | px | rgba(0,0,0,0.12) |
| 2 | Drop-down containers and drawers | px | px | rgba(0,0,0,0.12) |
| 3 | Dialogs, modal windows, and toasts | px | px | rgba(0,0,0,0.15) |
| Criterion | Definition |
| Color | Whether the page palette is visually harmonious, colors are used consistently, and sufficient contrast is maintained between text and background. |
| Typography | Whether font choices, sizes, weights, line heights, and textual hierarchy are clear, consistent, and readable. |
| Graphics & Imagery | Whether images, icons, illustrations, and other visual assets are clear, intact, appropriately proportioned, and visually well presented. |
| Layout | Whether page structure, information hierarchy, alignment, spacing, grouping, and space utilization are visually appropriate. |
| Component Consistency | Whether components with equivalent functions or hierarchy follow consistent rules in their size, structure, appearance, and state representation. |
| Visual Style Consistency | Whether corner radii, shadows, borders, line weights, icons, and other visual treatments form a coherent visual language across the page. |
| Dimension | Krippendorff’s | ICC(1,3) |
| Color | 0.328 | 0.596 |
| Typography | 0.334 | 0.609 |
| Graphics & Imagery | 0.346 | 0.614 |
| Layout | 0.431 | 0.694 |
| Component Consistency | 0.399 | 0.672 |
| Visual-style Consistency | 0.366 | 0.635 |
| Dimension | Violation Rule | Description |
| T | Peer Font Weight | Changes the weight of a peer text element, breaking consistency among texts at the same visual level. |
| T | Peer Font Size | Changes the size of a peer text element, disrupting the local typographic hierarchy. |
| T | Font Family Count | Introduces an additional font family, increasing unnecessary variation in the page typography. |
| T | Text Scale Count | Introduces an additional font-size level, increasing complexity in the existing typographic scale. |
| T | Table Text Alignment | Changes text or numeric alignment within a table, breaking its alignment consistency. |
| L | Vertical Alignment | Offsets a peer element from its original vertical alignment. |
| Dimension | MAE | Bias | SRCC | |||
| Raw | Cal. | Raw | Cal. | Raw | Cal. | |
| Color | 0.95 | 0.61 | +0.86 | 0.00 | 0.55 | 0.48 |
| Typography | 0.70 | 0.55 | +0.52 | 0.00 | 0.60 | 0.54 |
| Graphics & Imagery | 0.98 | 0.61 | +0.85 | 0.00 | 0.59 | 0.52 |
| Layout | 0.71 | 0.58 | +0.50 | 0.00 | 0.66 | 0.60 |
| Component Consistency | 1.03 | 0.58 | +0.95 | 0.00 | 0.62 | 0.55 |
| Model | Repair | ||||
| Claude Opus 5 | 0.83 | 0.97 | 0.97 | 0.82 | 0.76 |
| Doubao-Seed-2.1-Pro | 0.50 | 0.99 | 0.99 | 0.70 | 0.45 |
| GPT-5.6 Sol | 0.69 | 0.95 | 0.95 | 0.75 | 0.66 |
| Grok 4.6 | 0.74 | 0.94 | 0.94 | 0.77 | 0.70 |
| Kimi-K3 | 0.59 | 0.77 | 0.77 | 0.64 | 0.57 |
| Qwen3.7-Plus | 0.37 | 1.00 | 1.00 | 0.77 | 0.34 |