A key challenge for multimodal large language models (MLLMs) is moving beyond visual recognition to constraint-aware cross-modal reasoning. This involves combining visual cues with information from other modalities to understand elements' relationships under domain-specific rules. This challenge is acutely evident in industrial design-to-code (D2C), which converts user interface (UI) designs into code and requires MLLMs to connect design images with disorganized layer metadata, infer component and layout implementation requirements, and realize them in code under target-library constraints. However, these capabilities remain insufficiently evaluated in realistic industrial settings. To fill this gap, we present TaoD2C-Bench, a benchmark for evaluating MLLMs' ability to generate UI code that satisfies implementation requirements in industrial applications. The TaoD2C dataset consists of 2,861 production designs from 17 commercial platforms with 97,652 expert annotations across four categories: Component, Group, Alignment, and Position. These annotations distinguish required constraints from permitted implementation choices. TaoD2C-Bench defines three tasks: end-to-end UI code generation, requirement inference, and requirement realization. Evaluating eight MLLMs reveals substantial gaps in generating UI code that satisfies implementation requirements, alongside distinct performance profiles in inference and realization. We further show that MLLMs' visual reconstruction ability does not necessarily imply an ability to generate code that meets these requirements. We release TaoD2C to support research on industrial UI code generation.
Figures & tables
Figure 1: Two capabilities in UI code generation: (1) inferring positioning requirements and (2) realizing them in code.
Figure 2: Overview of TaoD2C-Bench. We collect 2,861 production UI designs in TaoD2C and annotate their implementation requirements across Component, Group, Alignment, and Position. These annotations provide a shared evaluation reference for UI code generation (T1), requirement inference (T2), and realization with expert requirements supplied (T3).
Figure 3
Benchmark
Data source
Scale
Metadata
Implementation requirements
Granularity
Design2Code ( Si et al., 2025 )
Web pages
484
✗
✗
Page
Web2Code ( Yun et al., 2024 )
Synthetic+web
884.7K
✗
✗
Page/QA
WebUIBench ( Lin et al., 2025 )
Websites
21.8K QA
✗
✗
QA/element
IW-Bench ( Guo et al., 2025 )
Generated+web
1,200
✗
✗
DOM/layout
WebCode2M ( Gui et al., 2025a )
Common Crawl
2.56M
✗
✗
Element/layout
Figma2Code ( Gui et al., 2026 )
Figma Comm.
3,055 (213 eval)
✓
✗
Page
Table 1: Comparison with existing design-to-code benchmarks. Implementation requirements are annotated by experts.
Figure 5: Three tasks and their evaluation. Expert annotations provide a shared reference for requirement prediction (T2) and code implementation (T1/T3); visual fidelity compares rendered pages with design images.
Model
End-to-End UI Code Generation (T1)
Requirement Inference (T2)
Code Implementation
Visual Fidelity
Type
Props
Group
Align.
Pos.
CV ↑
VES ↑
MAE ↓
Comp.
Group
Align.
Pos.
GPT-5.6 Terra
62.47
44.25
56.84
30.53
36.43
85.78
0.9174
0.1163
59.15
60.92
37.11
47.68
GPT-5.4
60.19
40.29
56.48
27.48
33.44
85.09
0.9183
0.1078
57.32
60.93
31.88
48.85
Claude Sonnet 5
59.89
40.37
54.82
23.54
27.93
78.82
0.8946
0.1350
58.36
50.48
34.84
47.62
Qwen3.8-Max
61.77
38.83
59.46
26.42
45.35
85.53
0.9192
0.1134
58.86
71.40
29.45
49.15
Table 2: Main results for T1 and T2 on 421 designs. Code Implementation and Requirement Prediction scores are micro F1 (%). Higher CV/VES and lower MAE indicate better visual fidelity. Bold denotes the best result; shading indicates within-column rank.
Model
Code Implementation
Visual Fidelity
Type
Props
Group
Align.
Pos.
CV ↑
VES ↑
MAE ↓
GPT-5.6 Terra
81.72 [1pt]+19.25
66.19 [1pt]+21.94
61.11 [1pt]+4.27
47.22 [1pt]+16.69
51.03 [1pt]+14.60
86.52 [1pt]+0.74
0.9263 [1pt]+0.0089
0.1016 [1pt] − 0.0147
GPT-5.4
75.75 [1pt]+15.56
61.44 [1pt]+21.15
64.54 [1pt]+8.06
49.44 [1pt]+21.96
58.43 [1pt]+24.99
84.95 [1pt] − 0.14
0.9118 [1pt] − 0.0065
0.1088 [1pt]+0.0010
Claude Sonnet 5
78.08 [1pt]+18.19
63.99 [1pt]+23.62
61.05 [1pt]+6.23
35.16 [1pt]+11.62
49.35 [1pt]+21.42
79.40 [1pt]+0.58
0.9003 [1pt]+0.0057
0.1344 [1pt] − 0.0006
Qwen3.8-Max
79.99 [1pt]+18.22
63.08 [1pt]+24.25
64.79 [1pt]+5.33
45.79 [1pt]+19.37
64.82 [1pt]+19.47
84.73 [1pt] − 0.80
0.9126 [1pt] − 0.0066
0.1139 [1pt]+0.0005
Qwen3-VL-235B
42.83 [1pt]+11.23
29.48 [1pt]+12.14
30.73 [1pt]+0.20
9.81 [1pt]+3.09
13.45 [1pt]+5.84
58.69 [1pt]+0.32
0.7784 [1pt] − 0.0135
0.1782 [1pt] − 0.0026
Table 3: Results for Requirement Realization (T3) on 421 designs. Values below each score show Δ=T3−T1 . Purple and yellow denote gains and losses, respectively; bold marks the largest gain per column.
Figure 6: Visually similar interfaces can violate implementation requirements. From left to right: incorrect component type, incorrect positioning, missing grouping container, and fixed spacing instead of space-between alignment. Code excerpts reveal these errors; bottom panels explain their potential impact on component integration, interaction handling, and layout adaptation.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Statistic
Count
Commercial platforms
17
Designs
2,861
Visible layers
183,697
Expert annotation records
97,652
Appendix
Table 4: Overall scale of TaoD2C.
Figure 7: Application scenarios across all 2,861 designs. Each design is counted once in its primary scenario.
Quantity
Mean
Median
P90
Minimum
Maximum
Visible layers
64.21
43
146
1
776
Expert annotation records
34.13
23
81
0
270
Appendix
Table 5: Per-design statistics across all 2,861 designs. P90 denotes the 90th percentile.
Figure 8: The annotation platform. Experts inspect the design and its layers, overlay the four requirement categories, and review individual annotations in the adjacent panel.
Category
Required
Optional
Total
Component
7,976
21
7,997
Group
12,658
49
12,707
Alignment
747
170
917
Position
711
350
1,061
Appendix
Table 6: Required and optional reference entries in the 421-design evaluation subset.
Figure 9: End-to-End UI Code Generation (T1) prompt summary, including the shared inputs, code output format, and generation constraints.
Figure 10: Requirement Inference (T2) prompt summary. The four requirement categories correspond to the components, groups, arrangements, and stacks arrays.
Figure 11: Requirement Realization (T3) prompt summary. T3 supplements the T1 inputs with the expert annotations for the input design while retaining the same code generation instructions and output format.
Figure 12: Output formats for the three TaoD2C-Bench tasks. Left: ordered file blocks for T1 and T3. Right: the four arrays and their entry fields for T2. File contents and field notation are schematic, not an executable response or a complete JSON example.
Model
Reasoning setting
HTTP output cap
GPT-5.6 Terra
Medium
—
GPT-5.4
Medium
—
Claude Sonnet 5
Adaptive, medium effort
60,000
Qwen3.8-Max
—
60,000
Qwen3-VL-235B-A22B-Thinking
—
32,768
Gemini 3.5 Flash
Medium
60,000
Appendix
Table 7: Model-specific inference settings. A dash indicates that no setting is specified. HTTP caps are transport limits, not task-level budgets.
Model
Type
Props
Group
Align.
Pos.
CV ↑
VES ↑
MAE ↓
GPT-5.6 Terra
67.74
47.21
41.87
39.95
23.63
82.23
0.9225
0.1221
GPT-5.4
66.75
43.45
39.78
37.90
12.56
83.12
0.9408
0.0991
Claude Sonnet 5
65.04
41.73
39.19
29.63
16.70
76.03
0.9022
0.1462
Qwen3.8-Max
69.11
40.57
39.03
31.47
16.97
79.68
0.9240
0.1352
Qwen3-VL-235B
30.67
13.95
26.53
9.65
11.38
61.40
0.8706
0.1824
Gemini 3.5 Flash
66.47
38.51
33.08
25.98
24.70
77.82
0.9278
0.1130
Appendix
Table 8: B-end results for End-to-End UI Code Generation (T1) on the same 152 designs for all eight models. F1 is in percent. Bold marks the best value per column; Mean averages models. CV and VES are higher-better; MAE is lower-better.
Model
Comp.
Group
Align.
Pos.
GPT-5.6 Terra
59.70
56.94
38.17
25.19
GPT-5.4
59.60
57.59
29.88
26.48
Claude Sonnet 5
58.68
56.51
35.56
23.87
Qwen3.8-Max
59.97
61.16
24.15
21.57
Qwen3-VL-235B
53.54
37.00
21.91
16.05
Gemini 3.5 Flash
59.68
52.51
32.22
25.59
Appendix
Table 9: B-end results for Requirement Inference (T2) on the same 152 designs for all eight models. F1 is in percent. Bold marks the best value per column; Mean averages models.
Model
Type
Props
Group
Align.
Pos.
CV ↑
VES ↑
MAE ↓
GPT-5.6 Terra
85.05
67.94
48.32
75.44
60.36
85.36
0.9447
0.0897
GPT-5.4
80.78
66.72
43.69
66.27
62.90
84.26
0.9383
0.1013
Claude Sonnet 5
83.04
66.09
42.30
55.25
70.23
78.47
0.9189
0.1400
Qwen3.8-Max
81.32
60.91
39.18
34.47
36.69
78.87
0.9226
0.1294
Qwen3-VL-235B
39.43
23.97
29.23
8.24
12.33
61.83
0.8545
0.1787
Gemini 3.5 Flash
85.81
70.53
39.54
56.72
65.62
78.34
0.9333
0.0981
Appendix
Table 10: B-end results for Requirement Realization (T3) on the same 152 designs for all eight models. F1 is in percent. Bold marks the best value per column; Mean averages models. CV and VES are higher-better; MAE is lower-better.
Model
Type
Props
Group
Align.
Pos.
CV ↑
VES ↑
MAE ↓
GPT-5.6 Terra
53.45
39.19
62.44
28.08
38.34
87.78
0.9145
0.1129
GPT-5.4
48.93
34.88
63.52
25.24
38.41
86.20
0.9056
0.1127
Claude Sonnet 5
51.67
38.20
60.99
22.10
29.55
80.40
0.8903
0.1287
Qwen3.8-Max
49.85
36.01
67.89
25.19
48.66
88.83
0.9165
0.1010
Qwen3-VL-235B
33.10
22.81
32.11
6.24
6.73
56.66
0.7475
0.1799
Gemini 3.5 Flash
54.51
37.21
58.91
18.71
30.03
82.02
0.9019
0.1232
Appendix
Table 11: C-end results for End-to-End UI Code Generation (T1) on the same 269 designs for all eight models. F1 is in percent. Bold marks the best value per column; Mean averages models. CV and VES are higher-better; MAE is lower-better.
Model
Comp.
Group
Align.
Pos.
GPT-5.6 Terra
58.23
61.98
36.85
50.87
GPT-5.4
53.47
61.87
32.61
52.32
Claude Sonnet 5
57.81
48.90
34.81
50.59
Qwen3.8-Max
57.02
73.90
32.18
53.28
Qwen3-VL-235B
52.86
17.23
15.94
28.14
Gemini 3.5 Flash
57.27
51.18
35.39
50.11
Appendix
Table 12: C-end results for Requirement Inference (T2) on the same 269 designs for all eight models. F1 is in percent. Bold marks the best value per column; Mean averages models.
Model
Type
Props
Group
Align.
Pos.
CV ↑
VES ↑
MAE ↓
GPT-5.6 Terra
75.56
62.95
65.81
41.47
50.17
87.17
0.9158
0.1083
GPT-5.4
66.89
52.14
72.85
46.68
58.05
85.34
0.8968
0.1130
Claude Sonnet 5
69.56
60.39
68.06
31.94
47.48
79.93
0.8898
0.1312
Qwen3.8-Max
77.56
67.01
74.90
50.00
68.02
88.04
0.9070
0.1050
Qwen3-VL-235B
49.53
40.37
31.32
10.08
13.57
56.92
0.7354
0.1779
Gemini 3.5 Flash
75.18
61.89
66.70
33.90
53.38
81.84
0.8724
0.1238
Appendix
Table 13: C-end results for Requirement Realization (T3) on the same 269 designs for all eight models. F1 is in percent. Bold marks the best value per column; Mean averages models. CV and VES are higher-better; MAE is lower-better.