Multimodal foundation models are increasingly used for evaluating and generating user interfaces (UIs), often producing seemingly reasonable aesthetic judgments and visually plausible pages. However, under professional design scrutiny, their behavior can differ substantially from that of human designers. In professional design practice, designers rely on a systematic set of aesthetic principles that consistently guide judgment, diagnosis, repair, and creation. A coherent aesthetic capability should therefore connect aesthetic judgment with design actions. Existing evaluations, however, typically assess these abilities in isolation, making it difficult to determine whether task-level success reflects a shared aesthetic understanding or merely fragmented task-specific competence. To address this gap, we introduce AUV-Bench, developed in collaboration with professional UI designers around 1,395 executable web interfaces and four tasks: aesthetic scoring, diagnosis, repair, and text-to-UI generation. The tasks share a pool of UIs and aesthetic principles, with diagnosis and repair further aligned on 660 controlled-degradation instances to enable instance-level analysis of judgment and action. Evaluation of 12 models reveals a capability imbalance: models show moderate agreement with professional designers in holistic aesthetic scoring, yet exact diagnosis-chain success peaks at only 24.7%. On the aligned diagnosis-repair cases, correct judgments and successful repairs do not consistently coincide, exposing a Judgment-Action Gap between identifying aesthetic problems and successfully acting on them. In open-ended generation, even leading models achieve only moderate aesthetic quality under human-calibrated evaluation. Overall, current models exhibit partial aesthetic competence, but still lack the fine-grained understanding and judgment-action coherence required for reliable UI design.
Figures & tables
Benchmark
Benchmark Setup
Principle Grounding
Core Capabilities
Judgment–Action Linkage
Statistics
Executable UI
Expert Aesthetic GT
Principle- Grounded
Controlled Violations
Aesthetic Scoring
Aesthetic Diagnosis
Aesthetic Repair
Text-to-UI Generation
Domain
Scale
Benchmarks for UI Generation
Design2Code ( Si et al., 2025 )
✓
✗
✗
✗
✗
✗
✗
✗
✗
Web UI
484 pages
WebUIBench ( Lin et al., 2025 )
✓
✗
✗
✗
✗
✗
✗
✗
✗
Web UI
21K QA / 0.7K+ sites
WebMMU ( Awal et al., 2025 )
✓
✗
✗
✗
✗
✗
✗
✗
✗
Web UI
8.1K tasks / 2,059 pages
WebGen-Bench ( Lu et al., 2026 )
✓
✗
✗
✗
✗
✗
✗
✓
✗
Web App
647 tests
Table 1: Comparison of existing benchmarks along key dimensions for evaluating UI aesthetic judgment and design action. Expert Aesthetic GT indicates whether aesthetic ground truth is provided or validated by professional designers or relevant domain experts. Principle-Grounded denotes explicit evaluation against predefined aesthetic design principles, while Controlled Violations indicates controlled manipulation of specific aesthetic principles. Judgment–Action Linkage indicates whether aesthetic diagnosis and corrective action are directly paired on the same interface and controlled aesthetic violation, enabling instance-level analysis of judgment–action coherence. ✓ and ✗ denote support and non-support, respectively.
Figure 1: Overview of AUV-Bench . (a) Data Statistics and (b) representative Data Examples of four tasks spanning aesthetic understanding ( Aesthetic Scoring and Aesthetic Diagnosis ) and design action ( Aesthetic Repair and Text-to-UI Generation ).
Figure 2: Distribution of the 1,395 reference UIs. (a) Data sources. (b) Industry categories. (c) Page types.
Figure 3: Overview of AUV-Bench construction. Starting from 1,395 executable UIs and professional aesthetic guidelines, we build four complementary tasks through expert annotation, controlled degradation, and requirement reconstruction. The controlled-degradation pipeline draws inspiration from all seven guideline dimensions and retains five operational categories for source-level editing and verification.
Model
Aesthetic Scoring
Aesthetic Diagnosis
Aesthetic Repair
Text-to-UI Generation
Judgment–Action Association ( ϕ ) ↑
SRCC ↑
MAE ↓
Det. BAcc ↑
Attr. Macro-J ↑
Loc. F1@0.5 ↑
ECS ↑
Repair Pass ↑
Cal. Aes. Judge ↑
Pairwise WR ↑
Frontier Models
Claude Opus 5
0.56
0.81
91.10
58.90
36.70
24.70
75.60
3.13
71.50
0.32
Doubao-Seed-2.1-Pro
0.50
1.10
79.60
36.20
7.60
4.10
45.00
2.78
62.90
0.31
GPT-5.6 Sol
0.55
0.90
87.60
55.40
28.20
19.40
65.50
3.17
78.00
0.38
Grok 4.6
0.54
0.95
81.20
42.70
6.50
2.90
70.00
3.08
62.50
0.25
Table 2: Main results across the four evaluation tasks and cross-task analysis. ↑ ( ↓ ) indicates that higher (lower) is better. – indicates that JAA is undefined because no instance satisfies Ji=1 .
Figure 4: In-depth analyses of aesthetic sensitivity and judgment–action coherence. (a) Aesthetic score changes under controlled degradations. (b) Repair performance under input ablations. (c) Repair with versus without Task 2 diagnosis; points above (below) the diagonal indicate positive (negative) diagnosis effects.
Figure 7
Appendix figures & tables40 assets
Supplementary material from the paper’s appendix.
Appendix
Dimension
Rule
Implementation
color
Primary : secondary : accent =6:3:1
Measure the pixel proportion of each color role and compute its KL divergence from the target ratio
Normal text/background contrast ≥4.5:1 ; large text ≥3:1 ; graphics and UI components ≥3:1
Compute the WCAG contrast ratio from foreground and background relative luminances using Eq. ( 6 )
Spacing
Element sizes and spacing use 4 , 12 , or integer multiples of 8 pixels
Compute the residual of element coordinates and dimensions modulo 8 , while admitting 4 and 12 as additional values
Typography
At most one CJK and one Latin typeface per page
Extract typeface information from the design source and compare typeface usage
hline=s+8 and at most five type sizes
Extract font-size and line-height parameters from the design source
Font weights use {400,500,600} and at most three weights per application
Extract font-weight values from the design source
Appendix
Table 4: Summary of aesthetic rules with explicitly specified quantitative or programmatic checks in the co-developed professional UI design guideline.
Figure 6: color proportion. The violating interface applies the accent color to a large portion of the page, while the conforming interface follows a more restrained color allocation.
Figure 7: color contrast. A confirmation message with insufficient foreground-background contrast, compared with a version satisfying the prescribed contrast requirement.
Figure 8: Functional color semantics. The violating example uses a color whose conventional semantic meaning conflicts with the represented state, while the conforming example follows the expected color semantics.
Figure 9: Spacing scale. Representative spacing values derived from the 8 -pixel base grid, together with the additional 4 - and 12 -pixel increments.
Figure 10: Spacing grid. The violating layout contains spacing values that do not follow the prescribed grid, while the conforming version uses admissible spacing increments.
Figure 11: Fitts’ Law. A primary action placed far from the screen edge, compared with a placement that follows the recommended edge distance.
Figure 12: Miller’s Law. Six similar content blocks are presented as one group in the violating example, while the conforming version reorganizes them into smaller groups.
Figure 13: Gutenberg Diagram. The violating example places important information and actions against the expected reading flow, while the conforming layout follows the recommended upper-left to lower-right organization.
Figure 14: Typeface discipline. The violating example mixes several unrelated typefaces within one interface, while the conforming version uses a unified typeface system.
Figure 15: Type scale and line height. The violating example uses line spacing that weakens the continuity of the text block, while the conforming version follows the additive line-height rule in Eq. equation 8 .
Figure 16: Font weight. The violating example uses an inconsistent weight assignment, while the conforming example follows the prescribed font-weight system.
Figure 17: Text and numerical alignment. Numerical values are left-aligned in the violating table, while the conforming version right-aligns numerical content and left-aligns text.
Figure 18: Image aspect ratio. The violating example distorts the image proportion, while the conforming version preserves an admissible aspect ratio.
Figure 19: Image clarity. The image in the violating example is insufficiently clear, while the conforming version preserves recognizable visual detail.
Metric
Definition and properties
PSNR
Peak signal-to-noise ratio is a full-reference image-quality metric that compares the maximum possible signal power with the power of corrupting noise. It is measured in decibels, and larger values indicate less distortion. Because it is based on pixel-level error, it does not explicitly model the visual characteristics of human perception.
SSIM
Structural similarity evaluates image similarity from luminance, contrast, and structure ( Wang et al., 2004 ) . Its value lies in [−1,1] , with larger values indicating less distortion. In practical computation, an image can be divided into local windows. For N windows, the average structural similarity is MSSIM(X,Y)=N1∑i=1NSSIM(xi,yi).
IFC
The information fidelity criterion evaluates image quality using natural scene statistics and characteristics of the human visual system. It measures the mutual information between a test image and a reference image ( Sheikh et al., 2005 ) .
VIF
Visual information fidelity extends the information-based formulation of IFC and focuses on the amount of visual information lost between a test image and its reference ( Sheikh and Bovik, 2006 ) .
MSE / RMSE
Mean squared error and root mean squared error measure pixel-level error between corresponding images. They are simple objective measures but do not explicitly account for characteristics of human visual perception.
Appendix
Table 5: Image-quality metrics included in our professional UI design guideline as references for evaluating image clarity.
Level
Typical components
Offset d
Blur
color
0
Buttons, inputs, search fields, tags, tables, links, pagination, steps, breadcrumbs, switches, radio buttons, checkboxes, and progress bars
0
0
rgba(0,0,0,0)
1
Navigation and card hover states
1 px
3 px
rgba(0,0,0,0.12)
2
Drop-down containers and drawers
2 px
4 px
rgba(0,0,0,0.12)
3
Dialogs, modal windows, and toasts
20 px
30 px
rgba(0,0,0,0.15)
Appendix
Table 6: Shadow hierarchy specified in our professional design guideline, following the Fusion Design shadow scale.
Figure 20: Shadow hierarchy. The violating example uses a shadow treatment inconsistent with the component’s semantic level, while the conforming example follows the prescribed shadow hierarchy.
Figure 21: Container consistency. Equivalent containers use different corner radii, shadows, or spacing in the violating example, while the conforming version applies a consistent container specification.
Figure 22: Component consistency. Instances of the same component type use different specifications in the violating example, while the conforming version follows a shared component specification.
Figure 23: Visual rubric exemplars used to calibrate annotators before formal annotation. Rows correspond to the eight aesthetic dimensions, and columns correspond to rating levels from 1 to 5.
Criterion
Definition
Color
Whether the page palette is visually harmonious, colors are used consistently, and sufficient contrast is maintained between text and background.
Typography
Whether font choices, sizes, weights, line heights, and textual hierarchy are clear, consistent, and readable.
Graphics & Imagery
Whether images, icons, illustrations, and other visual assets are clear, intact, appropriately proportioned, and visually well presented.
Layout
Whether page structure, information hierarchy, alignment, spacing, grouping, and space utilization are visually appropriate.
Component Consistency
Whether components with equivalent functions or hierarchy follow consistent rules in their size, structure, appearance, and state representation.
Visual Style Consistency
Whether corner radii, shadows, borders, line weights, icons, and other visual treatments form a coherent visual language across the page.
Appendix
Table 7: Fine-grained criteria used for human UI evaluation.
Figure 24: Score distributions from three individual annotators and the aggregated human MOS across eight aesthetic dimensions.
Dimension
Krippendorff’s α
ICC(1,3)
Color
0.328
0.596
Typography
0.334
0.609
Graphics & Imagery
0.346
0.614
Layout
0.431
0.694
Component Consistency
0.399
0.672
Visual-style Consistency
0.366
0.635
Appendix
Table 8: Inter-annotator reliability across aesthetic dimensions. We report ordinal Krippendorff’s α for inter-annotator agreement and ICC(1,3) for the reliability of the average of three ratings.
Figure 25: Model evaluation prompt for aesthetic scoring, Part I.
Figure 26: Model evaluation prompt for aesthetic scoring, Part II.
Figure 27: Model evaluation prompt for aesthetic scoring, Part III.
Figure 28: Representative human annotation example.
Figure 29: Representative human annotation example.
Figure 30: Representative human annotation example.
Dimension
Violation Rule
Description
T
Peer Font Weight
Changes the weight of a peer text element, breaking consistency among texts at the same visual level.
T
Peer Font Size
Changes the size of a peer text element, disrupting the local typographic hierarchy.
T
Font Family Count
Introduces an additional font family, increasing unnecessary variation in the page typography.
T
Text Scale Count
Introduces an additional font-size level, increasing complexity in the existing typographic scale.
T
Table Text Alignment
Changes text or numeric alignment within a table, breaking its alignment consistency.
L
Vertical Alignment
Offsets a peer element from its original vertical alignment.
Appendix
Table 9: The 13 controlled aesthetic violation rules used in our benchmark. A rule is instantiated only when the reference UI satisfies its required structural or semantic context. T, L, S, V, and C denote Typography, Layout & Reading Flow, Spacing, Visual Style Consistency, and Color, respectively.
Figure 31: Distribution of the formal evaluation cohort shared by Tasks 2 and 3. (a) Complexity distribution of the 660 underlying UI instances. (b) Distribution of the 1,209 controlled aesthetic defects across five aesthetic dimensions. (c) Distribution of mild and severe defects.
Dimension
MAE ↓
Bias →0
SRCC ↑
Raw
Cal.
Raw
Cal.
Raw
Cal.
Color
0.95
0.61
+0.86
0.00
0.55
0.48
Typography
0.70
0.55
+0.52
0.00
0.60
0.54
Graphics & Imagery
0.98
0.61
+0.85
0.00
0.59
0.52
Layout
0.71
0.58
+0.50
0.00
0.66
0.60
Component Consistency
1.03
0.58
+0.95
0.00
0.62
0.55
Appendix
Table 10: Five-fold cross-validation of GPT-5.4 aesthetic-score calibration against professional-designer MOS on Task 1. Bias denotes the mean signed error ( prediction − MOS ); values closer to zero indicate better calibration. Macro denotes the equally weighted average across the eight aesthetic dimensions.
Figure 32: Performance under increasing defect complexity. We compare single-violation ( Diagnostic ), compositional ( Combo ), and high-density ( Stress ) settings across degradation detection, violation attribution, exact-chain diagnosis, and aesthetic repair. Detection remains comparatively stable as more defects are introduced, whereas fine-grained diagnosis and repair become substantially more difficult.
Figure 33: Dimension-wise aesthetic capability profiles. Left: SRCC with professional-designer MOS for Aesthetic Scoring. Right: human-calibrated aesthetic scores for successfully rendered Text-to-UI generations. The two panels use the same eight-dimensional perceptual rubric but characterize scoring and generation, respectively.
Model
F
C
P
S
Repair
Claude Opus 5
0.83
0.97
0.97
0.82
0.76
Doubao-Seed-2.1-Pro
0.50
0.99
0.99
0.70
0.45
GPT-5.6 Sol
0.69
0.95
0.95
0.75
0.66
Grok 4.6
0.74
0.94
0.94
0.77
0.70
Kimi-K3
0.59
0.77
0.77
0.64
0.57
Qwen3.7-Plus
0.37
1.00
1.00
0.77
0.34
Appendix
Table 11: Decomposition of Aesthetic Repair performance. F : target fix; C : collateral preservation; P : content preservation; S : visual-scope compliance.
Figure 34: Absolute performance across UI industries. Group-average performance across 11 industries for Aesthetic Scoring (SRCC), Aesthetic Diagnosis (ECS), Aesthetic Repair (Repair Pass), and Text-to-UI Generation (Pairwise Win Rate). Each panel uses the original scale of its corresponding metric and should be compared within task.
Figure 35: Fine-grained performance of individual frontier models across 11 UI industries. We report industry-wise results for (a) Aesthetic Scoring (SRCC), (b) Aesthetic Diagnosis (ECS), (c) Aesthetic Repair (Repair Pass), and (d) Text-to-UI Generation (Pairwise Win Rate). Cell values show the original metric scores, while color scales are independently normalized within each task for clearer comparison.
Figure 36: Model-normalized industry effects across the four tasks. Each cell reports the mean within-model deviation from overall performance: Δg,t=∣M∣1∑m∈M(Sm,g,t−Sm,overall,t) . Positive values indicate above-overall performance and negative values indicate below-overall performance.
Figure 37: Industry-wise performance from coarse diagnosis to corrective action. We report degradation detection (BAcc), violation attribution (Macro-J), exact diagnosis-chain success (ECS), and final Repair Pass for the three model groups across industries.
While large language models (LLMs) and coding agents are often applied to user interface (UI) development, developers find it difficult to reliably assess their proficiency in visual and interaction design. Existing evaluations either rely on human experts, who can accurately assess usability by testing critical flows but are slow and costly, or on automated judges, which are scalable but less accurate and opaque. We present FlowEval, a reference-based framework that measures whether a generated UI supports realistic interaction flows by comparing navigation traces from real websites to traces from generated analogs using reference-based similarity metrics (e.g., dynamic time warping). In a small-scale study with expert UI evaluators, we show that reference-based metrics strongly correlate with human judgments, suggesting that they can provide scalable yet trustworthy evaluation for UI generation systems.
Generative UI tools promise to democratize UI design by turning natural language descriptions into complete interfaces. Alongside the interface, these tools generate user-facing design rationales that explain their layout, accessibility, and design choices. However, it remains unclear whether these stated rationales are actually reflected in the interfaces they produce. We call this disconnect ``Design Theater'': plausible and confident design rationales that have little relationship to the actual implementation. To study this phenomenon, we introduce a benchmark and three metrics for measuring Design Theater. The benchmark includes 24 UI generation tasks spanning structural, styling, and functional design requirements. Using this benchmark, we evaluate 120 interfaces created by five generative UI tools. On average, over 25% of user-facing design rationales are not implemented in the generated interface, and the implementation failure increases to 34% for functional requirements. Tools recognize roughly half of the UX principles embedded in prompts (mean = 0.54), with four of five tools implementing 6% or fewer functional principles. We also measure interface similarity across tools and find convergence in visual appearance and layout organization, with greater variation in color choices. Overall, we contribute: 1) the concept of Design Theater; 2) a benchmark with metrics for assessing whether the stated reasoning of generative UI tools is reflected in their implementations; 3) and findings from a systematic evaluation of these tools. We discuss what these findings mean for the design and evaluation of generative UI tools.
Kashif Imteyaz, Kaif Imteyaz, Nakul Rajpal +3
Northeastern University, United States · Jamia Hamdard, India · Saarland University, Germany +2
Generative UI (GenUI) lets large language models synthesize a complete, renderable interface directly from a natural-language instruction, but evaluating the quality of what they generate remains an open problem. Human evaluation is costly and rater-variant, while LLM-as-a-judge is scalable but reflects only a single implicit viewpoint, unable to capture how different populations of real users actually perceive the same interface. We propose the Evidence-Grounded, Social-Weighted Persona Panel (ESPP), a three-stage GenUI evaluation method in which a panel of psychologically diverse, evidence-grounded personas independently rates a screenshot, exchanges opinions under a trait-derived, semantically-gated bounded-confidence mechanism, and is aggregated via Delphi-inspired social weighting into a single judgment. ESPP tracks human judgment substantially more closely than a naive single-pass judge, raising Pearson r from 0.716 to 0.922, and a prompt-ensemble control recovers only about a third of this gap, isolating genuine persona and evidence grounding as the dominant source of improvement. Beyond this fidelity gain, retaining each panelist's individual rating further reveals that user subgroups agree on overall model rankings yet diverge sharply on specific rating dimensions, a structural disagreement a single homogeneous judge would systematically erase. The codes are available at https://github.com/Wuzheng02/ESPP.
Zheng Wu, Yibo Luo, Pu Zhang +2
School of Computer Science, Shanghai Jiao Tong University · 2ByteDance Inc