A key challenge for multimodal large language models (MLLMs) is moving beyond visual recognition to constraint-aware cross-modal reasoning. This involves combining visual cues with information from other modalities to understand elements' relationships under domain-specific rules. This challenge is acutely evident in industrial design-to-code (D2C), which converts user interface (UI) designs into code and requires MLLMs to connect design images with disorganized layer metadata, infer component and layout implementation requirements, and realize them in code under target-library constraints. However, these capabilities remain insufficiently evaluated in realistic industrial settings. To fill this gap, we present TaoD2C-Bench, a benchmark for evaluating MLLMs' ability to generate UI code that satisfies implementation requirements in industrial applications. The TaoD2C dataset consists of 2,861 production designs from 17 commercial platforms with 97,652 expert annotations across four categories: Component, Group, Alignment, and Position. These annotations distinguish required constraints from permitted implementation choices. TaoD2C-Bench defines three tasks: end-to-end UI code generation, requirement inference, and requirement realization. Evaluating eight MLLMs reveals substantial gaps in generating UI code that satisfies implementation requirements, alongside distinct performance profiles in inference and realization. We further show that MLLMs' visual reconstruction ability does not necessarily imply an ability to generate code that meets these requirements. We release TaoD2C to support research on industrial UI code generation.
Figures & tables
Figure 1: Two capabilities in UI code generation: (1) inferring positioning requirements and (2) realizing them in code.
Figure 2: Overview of TaoD2C-Bench. We collect 2,861 production UI designs in TaoD2C and annotate their implementation requirements across Component, Group, Alignment, and Position. These annotations provide a shared evaluation reference for UI code generation (T1), requirement inference (T2), and realization with expert requirements supplied (T3).
Figure 3
Benchmark
Data source
Scale
Metadata
Implementation requirements
Granularity
Design2Code ( Si et al., 2025 )
Web pages
484
✗
✗
Page
Web2Code ( Yun et al., 2024 )
Synthetic+web
884.7K
✗
✗
Page/QA
WebUIBench ( Lin et al., 2025 )
Websites
21.8K QA
✗
✗
QA/element
IW-Bench ( Guo et al., 2025 )
Generated+web
1,200
✗
✗
DOM/layout
WebCode2M ( Gui et al., 2025a )
Common Crawl
2.56M
✗
✗
Element/layout
Figma2Code ( Gui et al., 2026 )
Figma Comm.
3,055 (213 eval)
✓
✗
Page
Table 1: Comparison with existing design-to-code benchmarks. Implementation requirements are annotated by experts.
Figure 5: Three tasks and their evaluation. Expert annotations provide a shared reference for requirement prediction (T2) and code implementation (T1/T3); visual fidelity compares rendered pages with design images.
Model
End-to-End UI Code Generation (T1)
Requirement Inference (T2)
Code Implementation
Visual Fidelity
Type
Props
Group
Align.
Pos.
CV ↑
VES ↑
MAE ↓
Comp.
Group
Align.
Pos.
GPT-5.6 Terra
62.47
44.25
56.84
30.53
36.43
85.78
0.9174
0.1163
59.15
60.92
37.11
47.68
GPT-5.4
60.19
40.29
56.48
27.48
33.44
85.09
0.9183
0.1078
57.32
60.93
31.88
48.85
Claude Sonnet 5
59.89
40.37
54.82
23.54
27.93
78.82
0.8946
0.1350
58.36
50.48
34.84
47.62
Qwen3.8-Max
61.77
38.83
59.46
26.42
45.35
85.53
0.9192
0.1134
58.86
71.40
29.45
49.15
Table 2: Main results for T1 and T2 on 421 designs. Code Implementation and Requirement Prediction scores are micro F1 (%). Higher CV/VES and lower MAE indicate better visual fidelity. Bold denotes the best result; shading indicates within-column rank.
Model
Code Implementation
Visual Fidelity
Type
Props
Group
Align.
Pos.
CV ↑
VES ↑
MAE ↓
GPT-5.6 Terra
81.72 [1pt]+19.25
66.19 [1pt]+21.94
61.11 [1pt]+4.27
47.22 [1pt]+16.69
51.03 [1pt]+14.60
86.52 [1pt]+0.74
0.9263 [1pt]+0.0089
0.1016 [1pt] − 0.0147
GPT-5.4
75.75 [1pt]+15.56
61.44 [1pt]+21.15
64.54 [1pt]+8.06
49.44 [1pt]+21.96
58.43 [1pt]+24.99
84.95 [1pt] − 0.14
0.9118 [1pt] − 0.0065
0.1088 [1pt]+0.0010
Claude Sonnet 5
78.08 [1pt]+18.19
63.99 [1pt]+23.62
61.05 [1pt]+6.23
35.16 [1pt]+11.62
49.35 [1pt]+21.42
79.40 [1pt]+0.58
0.9003 [1pt]+0.0057
0.1344 [1pt] − 0.0006
Qwen3.8-Max
79.99 [1pt]+18.22
63.08 [1pt]+24.25
64.79 [1pt]+5.33
45.79 [1pt]+19.37
64.82 [1pt]+19.47
84.73 [1pt] − 0.80
0.9126 [1pt] − 0.0066
0.1139 [1pt]+0.0005
Qwen3-VL-235B
42.83 [1pt]+11.23
29.48 [1pt]+12.14
30.73 [1pt]+0.20
9.81 [1pt]+3.09
13.45 [1pt]+5.84
58.69 [1pt]+0.32
0.7784 [1pt] − 0.0135
0.1782 [1pt] − 0.0026
Table 3: Results for Requirement Realization (T3) on 421 designs. Values below each score show Δ=T3−T1 . Purple and yellow denote gains and losses, respectively; bold marks the largest gain per column.
Figure 6: Visually similar interfaces can violate implementation requirements. From left to right: incorrect component type, incorrect positioning, missing grouping container, and fixed spacing instead of space-between alignment. Code excerpts reveal these errors; bottom panels explain their potential impact on component integration, interaction handling, and layout adaptation.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Statistic
Count
Commercial platforms
17
Designs
2,861
Visible layers
183,697
Expert annotation records
97,652
Appendix
Table 4: Overall scale of TaoD2C.
Figure 7: Application scenarios across all 2,861 designs. Each design is counted once in its primary scenario.
Quantity
Mean
Median
P90
Minimum
Maximum
Visible layers
64.21
43
146
1
776
Expert annotation records
34.13
23
81
0
270
Appendix
Table 5: Per-design statistics across all 2,861 designs. P90 denotes the 90th percentile.
Figure 8: The annotation platform. Experts inspect the design and its layers, overlay the four requirement categories, and review individual annotations in the adjacent panel.
Category
Required
Optional
Total
Component
7,976
21
7,997
Group
12,658
49
12,707
Alignment
747
170
917
Position
711
350
1,061
Appendix
Table 6: Required and optional reference entries in the 421-design evaluation subset.
Figure 9: End-to-End UI Code Generation (T1) prompt summary, including the shared inputs, code output format, and generation constraints.
Figure 10: Requirement Inference (T2) prompt summary. The four requirement categories correspond to the components, groups, arrangements, and stacks arrays.
Figure 11: Requirement Realization (T3) prompt summary. T3 supplements the T1 inputs with the expert annotations for the input design while retaining the same code generation instructions and output format.
Figure 12: Output formats for the three TaoD2C-Bench tasks. Left: ordered file blocks for T1 and T3. Right: the four arrays and their entry fields for T2. File contents and field notation are schematic, not an executable response or a complete JSON example.
Model
Reasoning setting
HTTP output cap
GPT-5.6 Terra
Medium
—
GPT-5.4
Medium
—
Claude Sonnet 5
Adaptive, medium effort
60,000
Qwen3.8-Max
—
60,000
Qwen3-VL-235B-A22B-Thinking
—
32,768
Gemini 3.5 Flash
Medium
60,000
Appendix
Table 7: Model-specific inference settings. A dash indicates that no setting is specified. HTTP caps are transport limits, not task-level budgets.
Model
Type
Props
Group
Align.
Pos.
CV ↑
VES ↑
MAE ↓
GPT-5.6 Terra
67.74
47.21
41.87
39.95
23.63
82.23
0.9225
0.1221
GPT-5.4
66.75
43.45
39.78
37.90
12.56
83.12
0.9408
0.0991
Claude Sonnet 5
65.04
41.73
39.19
29.63
16.70
76.03
0.9022
0.1462
Qwen3.8-Max
69.11
40.57
39.03
31.47
16.97
79.68
0.9240
0.1352
Qwen3-VL-235B
30.67
13.95
26.53
9.65
11.38
61.40
0.8706
0.1824
Gemini 3.5 Flash
66.47
38.51
33.08
25.98
24.70
77.82
0.9278
0.1130
Appendix
Table 8: B-end results for End-to-End UI Code Generation (T1) on the same 152 designs for all eight models. F1 is in percent. Bold marks the best value per column; Mean averages models. CV and VES are higher-better; MAE is lower-better.
Model
Comp.
Group
Align.
Pos.
GPT-5.6 Terra
59.70
56.94
38.17
25.19
GPT-5.4
59.60
57.59
29.88
26.48
Claude Sonnet 5
58.68
56.51
35.56
23.87
Qwen3.8-Max
59.97
61.16
24.15
21.57
Qwen3-VL-235B
53.54
37.00
21.91
16.05
Gemini 3.5 Flash
59.68
52.51
32.22
25.59
Appendix
Table 9: B-end results for Requirement Inference (T2) on the same 152 designs for all eight models. F1 is in percent. Bold marks the best value per column; Mean averages models.
Model
Type
Props
Group
Align.
Pos.
CV ↑
VES ↑
MAE ↓
GPT-5.6 Terra
85.05
67.94
48.32
75.44
60.36
85.36
0.9447
0.0897
GPT-5.4
80.78
66.72
43.69
66.27
62.90
84.26
0.9383
0.1013
Claude Sonnet 5
83.04
66.09
42.30
55.25
70.23
78.47
0.9189
0.1400
Qwen3.8-Max
81.32
60.91
39.18
34.47
36.69
78.87
0.9226
0.1294
Qwen3-VL-235B
39.43
23.97
29.23
8.24
12.33
61.83
0.8545
0.1787
Gemini 3.5 Flash
85.81
70.53
39.54
56.72
65.62
78.34
0.9333
0.0981
Appendix
Table 10: B-end results for Requirement Realization (T3) on the same 152 designs for all eight models. F1 is in percent. Bold marks the best value per column; Mean averages models. CV and VES are higher-better; MAE is lower-better.
Model
Type
Props
Group
Align.
Pos.
CV ↑
VES ↑
MAE ↓
GPT-5.6 Terra
53.45
39.19
62.44
28.08
38.34
87.78
0.9145
0.1129
GPT-5.4
48.93
34.88
63.52
25.24
38.41
86.20
0.9056
0.1127
Claude Sonnet 5
51.67
38.20
60.99
22.10
29.55
80.40
0.8903
0.1287
Qwen3.8-Max
49.85
36.01
67.89
25.19
48.66
88.83
0.9165
0.1010
Qwen3-VL-235B
33.10
22.81
32.11
6.24
6.73
56.66
0.7475
0.1799
Gemini 3.5 Flash
54.51
37.21
58.91
18.71
30.03
82.02
0.9019
0.1232
Appendix
Table 11: C-end results for End-to-End UI Code Generation (T1) on the same 269 designs for all eight models. F1 is in percent. Bold marks the best value per column; Mean averages models. CV and VES are higher-better; MAE is lower-better.
Model
Comp.
Group
Align.
Pos.
GPT-5.6 Terra
58.23
61.98
36.85
50.87
GPT-5.4
53.47
61.87
32.61
52.32
Claude Sonnet 5
57.81
48.90
34.81
50.59
Qwen3.8-Max
57.02
73.90
32.18
53.28
Qwen3-VL-235B
52.86
17.23
15.94
28.14
Gemini 3.5 Flash
57.27
51.18
35.39
50.11
Appendix
Table 12: C-end results for Requirement Inference (T2) on the same 269 designs for all eight models. F1 is in percent. Bold marks the best value per column; Mean averages models.
Model
Type
Props
Group
Align.
Pos.
CV ↑
VES ↑
MAE ↓
GPT-5.6 Terra
75.56
62.95
65.81
41.47
50.17
87.17
0.9158
0.1083
GPT-5.4
66.89
52.14
72.85
46.68
58.05
85.34
0.8968
0.1130
Claude Sonnet 5
69.56
60.39
68.06
31.94
47.48
79.93
0.8898
0.1312
Qwen3.8-Max
77.56
67.01
74.90
50.00
68.02
88.04
0.9070
0.1050
Qwen3-VL-235B
49.53
40.37
31.32
10.08
13.57
56.92
0.7354
0.1779
Gemini 3.5 Flash
75.18
61.89
66.70
33.90
53.38
81.84
0.8724
0.1238
Appendix
Table 13: C-end results for Requirement Realization (T3) on the same 269 designs for all eight models. F1 is in percent. Bold marks the best value per column; Mean averages models. CV and VES are higher-better; MAE is lower-better.
Design-to-code translates high-fidelity UI designs into executable front-end implementations, but progress remains hard to compare due to inconsistent datasets, toolchains, and evaluation protocols. We introduce 1D-Bench, a benchmark grounded in real e-commerce workflows, where each instance provides a reference rendering and an exported intermediate representation that may contain extraction errors. 1D is short for one day, representing the efficient completion of design-to-code tasks in less than one day. Models take both as input, using the intermediate representation as structural cues while being evaluated against the reference rendering, which tests robustness to intermediate representation defects rather than literal adherence. 1D-Bench requires generating an executable React codebase under a fixed toolchain with an explicit component hierarchy, and defines a multi-round setting in which models iteratively apply component-level edits using execution feedback. Experiments on commercial and open-weight multimodal models show that iterative editing generally improves final performance by increasing rendering success and often improving visual similarity. We further conduct a pilot study on post-training with synthetic repair trajectories and reinforcement learning based editing, and observe limited and unstable gains that may stem from sparse terminal rewards and high-variance file-level updates. The data and scripts used in this study are available in an anonymized repository at https://anonymous.4open.science/r/d2c-benchmark-A9C4/.
Multimodal large language models (MLLMs) are increasingly expected to automate visualization development by generating code directly from visual designs. However, existing evaluations mainly focus on single-chart generation and overlook coordinated multi-view interface construction, which requires joint reasoning about data semantics, view coordination, and interaction logic. Consequently, MLLM capabilities in this setting remain underexplored, and the field lacks a dedicated benchmark for systematic assessment. We introduce MV-Bench, a benchmark for evaluating MLLMs on coordinated multi-view interface construction. Instead of relying on incomplete or inconsistent open-source implementations, we use Tableau workbook files as ground truth because they explicitly encode data bindings, visual mappings, and interactions. We develop a multi-stage pipeline that converts these specifications into executable web interfaces through structured intermediate representations. The benchmark contains 92 base interfaces and 1,048 verified instances created by recombining chart types, datasets, and interaction patterns. Each instance includes executable code, a rendered interface, a dataset, and interaction annotations. We evaluate five state-of-the-art MLLMs in a single-pass setting using metrics for visual fidelity, data binding correctness, and interaction completeness. The strongest model achieves 75.45 percent accuracy in visual layout reproduction, but only 21.71 percent in data binding and 11.68 percent in interaction completeness. These results show that current MLLMs can reproduce visual appearance but remain limited in generating the data semantics and interactive logic required by coordinated multi-view interfaces. Iterative refinement improves code executability but does not substantially reduce the gap in data binding and interaction generation.
Yue Zhao, Hongxu Liu, Feiyu Wang +5
Shandong Second Medical University · Shandong University · Bairong Inc. +2
Existing screenshot-to-code systems face a trade-off between flexibility and controllability. Direct multimodal generation can hallucinate visible details, whereas structured pipelines reduce such errors through component-wise decomposition, predefined templates, and customized intermediate representations. These structures, however, introduce additional generative orchestration and restrict outputs to designs covered by the representation. We investigate whether selective tool grounding can improve the fidelity--efficiency trade-off of direct widget-to-code generation. We introduce \textbf{WidgetGen}, a lightweight tool-grounded framework that extracts observable text and color evidence, performs high-level layout and optional chart reasoning, and directly generates executable JavaScript XML (\emph{JSX}). This design reduces reliance on component-wise generation while avoiding a fixed UI schema. Across six multimodal models and 1,000 held-out widgets, WidgetGen outperforms direct prompting and the structured Widget2Code pipeline on most visual reconstruction metrics, with consistent gains in area, legibility, and style. Finally, reconstruction-derived image-code pairs improve six Qwen-family open-weight models across every reported metric through supervised fine-tuning. These results establish WidgetGen as a strong lightweight baseline and show that selective evidence grounding offers an effective alternative to extensive representation constraints.
Houston H. Zhang, Tao Zhang, Li Gu +5
McMaster University · University of Toronto · Concordia University