Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in both visual understanding and code generation. However, existing benchmarks typically evaluate these two modalities in isolation, lacking a dedicated assessment of their unification, i.e., how a model can perceive complex visual structures and synthesize them into precise, executable code. Moreover, current visual code generation benchmarks often rely on simplified layouts within single programming environments, falling short of evaluating true unified multimodal reasoning. To bridge this gap, we propose FigCodeBench, a comprehensive framework for rigorously evaluating MLLMs on figure reproduction, integrating multimodal comprehension and generation. We first design a systematic dataset construction pipeline, resulting in a total of 6,194 instances that cover 7 functional categories and 4 types of programming languages. We further categorize figure reproduction into three tiers with visual and code complexity modeling, specifically targeting complex structural reasoning, varying aspect ratios, and dense geometric constraints. We introduce a multi-dimensional evaluation protocol, encompassing visual fidelity and syntactic isomorphism, that aligns highly with the Mean Machine Opinion Score (MMOS) and human preferences. Based on our framework, we conducted extensive experiments on 24 widely used proprietary and open-source MLLMs (e.g., Gemini 3.1 Pro, GPT-5.4, and Kimi-K2.5), where we observed a universal, non-linear performance cliff across different programming languages and difficulty scenarios for all models, and gained several insights, such as the significant metric decline in rigid declarative languages.
Figures & tables
Figure 1: FigCodeBench leaderboard . Left : Mean Machine Opinion Score (MMOS) versus average cost per problem for various models. Right : Performance comparison of six representative models on different image categories.
Language
Package
Tag
# Prototype
# Query Imgs
Pct.%
Code Len.
Difficulty
Python
Matplotlib &Seaborn
Statistical
56/51
168/148
9.72/3.31
379.7/772.4
[30.52,60.25]/[2.42,61.69]
Relational
29/44
87/131
5.03/2.93
337.3/625.2
[31.35,60.92]/[36.72,64.56]
Temporal
17/34
51/97
2.95/2.17
377.1/514.2
[38.09,52.42]/[16.76,59.65]
Compositional
9/8
27/24
1.56/0.54
305.9/485.1
[39.13,52.50]/[43.46,52.31]
Geospatial
8/1
24/3
1.39/0.07
338.6/566.3
[34.39,55.02]/[44.87,50.88]
Mathematical
37/39
111/116
6.42/2.60
376.2/496.1
[32.40,55.47]/[41.84,60.91]
Table 1: FigCodeBench statistics for annotated tags with corresponding functional-focus taxonomy. Breakdown by the number of figure plotting topics, query image counts, percentages, average code length ( tokens calculated by the LLama3 tokenizer ), and difficulty scores. The slash separates statistics of exemplary data ( left ) and user-generated data ( right ).
Figure 2: Overview of FigCodeBench . Upper : The construction pipeline of FigCodeBench includes source data collection, attribute-guided perturbation, auxiliary information extraction, and quality control. Bottom left : This gallery displays 16 representative cases, organized into seven functional categories, i.e., statistical, relational, temporal, compositional, geospatial, mathematical, and conceptual visualizations. We summarize the key features of this benchmark in the bottom right part.
Figure 3: Scatter plot and marginal distribution of the image-code complexity of FigCodeBench. The red dotted lines separate the easy ( [2.42,44.73] ), medium ( [44.73,50.05] ) and hard tiers [50.05,76.98] . We also visualize some image-code cases ( points that are neither zero-sum nor on the coordinate axes ) at the corresponding minimum and maximum difficulty tiers.
Dataset/Benchmark
Source
# Lang.
# Fig. Type
# Test Inst.
Reso.
Eval. Format
Metric
ChartLlama ( Han et al., 2023 )
Synthesized
1
10
458
704 × 516
I+T → T
GPT Score, EM
MMCode ( Li et al., 2024c )
Crawl
1
12
263
534 × 302
I+T → C
Pass rate
MatPlotBench ( Yang et al., 2024 )
Crawl
1
13
100
849 × 639
T → C → I
GPT Score
ChartX ( Xia et al., 2025 )
Synthesized
1
18
6,000
1176 × 819
I+T → T&C → I
Multiple
Plot2Code ( Wu et al., 2025a )
Crawl
2
6
368
903 × 664
I+T → C → I
Multiple
Design2Code ( Si et al., 2025 )
Crawl
1
HTML
484
1280 × 1482
I+T → C → I
Multiple
Table 2: Comparison of our FigCodeBench with other related datasets and benchmarks. “I”, “T”, and “C” denote the image, text, and code modalities, respectively. “EM” denotes the exact match.
Figure 4: A gallery of random samples from seven compared benchmarks (i.e., ChartLlama, MMCode, MatPlotBench, ChartX, Plot2Code, Design2Code, and ChartMimic) and our proposed FigCodeBench. FigCodeBench covers 4 different languages, 7 categories, and 2,548 prototypes with 6,194 query images for standardized evaluation of MLLMs, significantly surpassing the existing benchmarks in terms of scale and coverage. Zoom-in for better visualization.
Figure 5: Feature distribution comparisons among eight benchmarks: ChartLlama, MMCode, MatPlotBench, ChartX, Plot2Code, Design2Code, ChartMimic, and our FigCodeBench.
Figure 6: Query image (blue ‘x’) distribution in paired feature space with corresponding convex hulls (red boundaries). Left : Brightness (BR) × Contrast (CT), right : Colorfulness (CF) × Sharpness (SR).
Model Name
Organizations
Cut-off Date
Citation
gpt-5.4
OpenAI
August 2025
( OpenAI, 2026d )
gpt-5.4-mini
OpenAI
August 2025
( OpenAI, 2026d )
gpt-5.2
OpenAI
August 2025
( OpenAI, 2025b )
claude-opus-4.7
Anthropic
January 2026
( Anthropic, 2026 )
gemini-3.1-pro-preview
Google
January 2025
( Google, 2026a )
gemini-3-flash-preview-nothinking
Google
January 2025
( Google, 2025b )
Table 3: List of models evaluated and their respective organizations, cut-off date, and citations.
Model
Python (n=987)
Matlab (n=600)
R (n=1130)
Latex (n=3477)
AvgTok
AvgCost
AvgTok
AvgCost
AvgTok
AvgCost
AvgTok
AvgCost
Proprietary Models
GPT-5.4
814.98
$0.0117
946.19
$0.0122
898.51
$0.0119
1176.65
$0.0154
GPT-5.4-Mini
822.65
$0.0034
702.30
$0.0029
831.93
$0.0040
1075.91
$0.0043
GPT-5.2
792.12
$0.0088
951.72
$0.0110
830.21
$0.0103
1026.63
$0.0105
Claude Opus 4.7
597.10
$0.0223
590.41
$0.0204
341.74
$0.0128
537.90
$0.0192
Table 4: The evaluation resource consumption of FigCodeBench. AvgTok is the average number of tokens generated per query and AvgCost is the approximate $-cost per query.
Model
Python
Matlab
R
Latex
Exec.
PSNR/SSIM
LPIPS/ ΔE
Exec.
PSNR/SSIM
LPIPS/ ΔE
Exec.
PSNR/SSIM
LPIPS/ ΔE
Exec.
PSNR/SSIM
LPIPS/ ΔE
Proprietary Models
GPT-5.4
91.48
12.807 / 0.628
0.411 /16.607
44.67
5.697/0.307
0.724/59.897
62.48
8.080/0.460
0.664/44.177
52.43
8.151/0.416
0.668/50.986
GPT-5.4-Mini
90.87
12.545/0.614
0.467/16.912
75.00
9.103/0.500
0.592/33.133
78.85
10.340/0.581
0.571/28.265
48.12
7.764/0.388
0.702/54.294
GPT-5.2
91.38
12.452/0.614
0.487/ 16.459
68.17
8.239/0.446
0.649/39.135
41.33
5.248/0.297
0.786/62.414
39.69
6.218/0.315
0.760/62.473
Claude Opus 4.7
99.29
10.189/0.583
0.726/17.188
98.17
10.205/ 0.585
0.701/ 15.603
97.35
10.714/0.689
0.633/15.794
88.47
12.822 / 0.683
0.594/ 18.398
Table 5: Image-oriented reproduction performance on FigCodeBench . PSNR/SSIM ( ↑ ) and LPIPS/ ΔE ( ↓ ) represent the pixel-level and perceptual metrics, respectively. “TK” denotes the thinking version. The best and second-best results are in bold and underlined .
Model
Python
Matlab
R
Latex
CBS
CB
SAST
FAPI
CB
SAST
FAPI
CB
SAST
FAPI
CB
SAST
FAPI
Proprietary Models
GPT-5.4
0.818
0.150
0.649
0.537
0.090
0.538
0.342
0.216
0.719
0.506
0.162
0.596
0.630
GPT-5.4-Mini
0.812
0.138
0.630
0.518
0.111
0.590
0.415
0.215
0.725
0.494
0.152
0.599
0.617
GPT-5.2
0.819
0.143
0.645
0.547
0.081
0.581
0.403
0.221
0.727
0.525
0.159
0.595
0.630
Claude Opus 4.7
0.753
0.036
0.618
0.312
0.033
0.565
0.314
0.045
0.688
0.447
0.028
0.480
0.437
Table 6: Code-level fidelity of figure reproduction performance on FigCodeBench. We report CodeBERTScore (CBS; Python only), CrystalBLEU (CB), normalized AST similarity ( SAST ), and plotting-API F1 ( FAPI ) between the generated code and the corresponding ground-truth implementation. Higher values indicate greater similarity to the corresponding ground-truth code. Missing or unmatched code outputs are assigned a score of zero when computing the all-instance averages. The best and second-best results are in bold and underlined .
Model
Python
Matlab
R
Latex
J1
J2
J3
MMOS
J1
J2
J3
MMOS
J1
J2
J3
MMOS
J1
J2
J3
MMOS
Proprietary Models
GPT-5.4
85.59
75.87
77.74
79.73
41.58
36.83
35.52
37.98
53.65
45.70
41.57
46.97
46.40
39.04
35.86
40.43
GPT-5.4-Mini
84.31
74.60
66.76
75.22
65.70
56.62
56.64
59.65
67.46
57.88
51.12
58.82
41.27
33.86
30.57
35.23
GPT-5.2
85.43
75.52
68.04
76.33
61.77
53.87
52.06
55.90
34.97
29.59
26.43
30.33
35.86
29.77
26.97
30.87
Claude Opus 4.7
86.33
77.59
73.56
76.19
72.37
61.04
58.73
64.05
68.35
62.17
54.66
61.73
55.14
50.93
49.35
51.81
Table 7: Performance comparison on mean model opinion score (MMOS). We report the scores rated by three MLLM judges. J1 , J2 , and J3 represent GPT-5.1, GPT-5.6-Luna, and Gemini-3.5-Flash, respectively.
Figure 7: Statistical significance of cross-language differences in image-oriented performance. Pairwise differences among Python, Matlab, R, and LaTeX are examined for execution success rate (Exec), PSNR, SSIM, LPIPS, and CIEDE2000 color difference ( ΔE ). Bars and vertically displayed numbers show the mean performance across 24 evaluated models, and error bars denote the standard error of the mean. All P values were calculated using two-sided Wilcoxon signed-rank tests on paired model-level measurements, followed by Holm–Bonferroni correction. NS, not significant; * ≤0.05 , ** ≤0.01 , and *** ≤0.001 .
Figure 8: Fine-grained score distribution generated by MLLM judges. We report the average scores of three MLLM judges for each evaluated model across four programming languages.
Figure 9: Pearson correlation between different MLLM judges across four programming languages.
Figure 10: Qualitative visualization of MLLM judges. All judge models produce human-aligned scores and mutually consistent rankings. From top to bottom, the reproduced figures show progressively decreasing similarity to the reference.
Figure 11: Comparison of Fig-Code Fidelity (FCF) results across four programming languages.
Model
Statistical
Relational
Temporal
Compositional
Geospatial
Mathematical
Conceptual
Proprietary Models
GPT-5.4
61.15
55.01
54.30
55.16
41.08
56.56
40.43
GPT-5.4-Mini
69.28
61.66
65.81
59.63
36.47
66.02
35.23
GPT-5.2
52.19
51.09
55.25
43.77
21.67
60.06
30.87
Claude Opus 4.7
68.76
62.31
72.58
66.74
30.75
61.48
51.81
Gemini 3.1 Pro
70.84
66.92
72.42
64.00
65.57
72.50
50.72
Table 8: Performance comparison across different functional categories of figures . MMOS is used as the indicator.
Figure 12: Model size versus MMOS . Trend lines are added to show the performance trend of each model family across different parameter scales. Note that we only investigate open-source MLLMs with publicly available parameters.
Figure 13: Correlation analysis of image-oriented metrics.
Figure 14: Correlation analysis of code-oriented metrics.
Model
Python
Matlab
R
Latex
PSNR
SSIM
LPIPS
ΔE
PSNR
SSIM
LPIPS
ΔE
PSNR
SSIM
LPIPS
ΔE
PSNR
SSIM
LPIPS
ΔE
Spearman’s ρ
0.7687
0.7856
-0.7380
-0.7730
0.8035
0.7617
-0.8915
-0.5290
0.6783
0.7180
-0.7423
-0.7278
0.5313
0.5462
-0.6269
-0.5235
Kendall’s τ
0.5725
0.5621
-0.5699
-0.5435
0.6232
0.5725
-0.7441
-0.7104
0.4928
0.5554
-0.5554
-0.5870
0.3333
0.3376
-0.4335
-0.3406
Table 9: Correlation between image-oriented metrics and MMOS . Given that code-oriented metrics primarily quantify similarity and thus do not directly reflect final graphical fidelity, we exclude them in this experiment.
Figure 15: Performance correlation of FigCodeBench with benchmarks evaluating chart understanding and code generation capabilities.
Figure 16: Screenshot of the user interface (UI) for subjective evaluation.
Metric
Chart Type
Layout
Text Content
Data
Style
Overall
J1
0.863
0.804
0.852
0.752
0.738
0.793
J2
0.881
0.822
0.846
0.774
0.744
0.823
J3
0.897
0.815
0.863
0.807
0.761
0.837
MMOS
0.883
0.817
0.855
0.783
0.747
0.826
Table 10: Pearson correlation coefficient r between multi-level MMOS and human evaluation. J1 , J2 , and J3 represent GPT-5.1, GPT-5.6-Luna, and Gemini-3.5-Flash, respectively.
Figure 17: The MMOS of six representative models at different difficulty tiers.
Figure 18: Comparison of figures reproduced by different models. We visualize results from six representative MLLMs, including GPT-5.4, Gemini 3.1 Pro, Grok 4.3, GLM-5.1, Kimi-K2.5, and Qwen3-VL-235B-A22B. The rectangle with diagonal lines indicates a failure in generation.
We introduce Chart2Code, a new benchmark for evaluating the chart understanding and code generation capabilities of large multimodal models (LMMs). Chart2Code is explicitly designed from a user-driven perspective, capturing diverse real-world scenarios and progressively increasing task difficulty. It consists of three levels: Level 1 (Chart Reproduction) reproduces charts from a reference figure and user query; Level 2 (Chart Editing) involves complex modifications such as changing chart types or adding elements; and Level 3 (Long-Table to Chart Generation) requires models to transform long, information-dense tables into faithful charts following user instructions. To our knowledge, this is the first hierarchical benchmark that reflects practical chart2code usage while systematically scaling task complexity. In total, Chart2Code contains 2,023 tasks across 22 chart types, paired with multi-level evaluation metrics that assess both code correctness and the visual fidelity of rendered charts. We benchmark 25 state-of-the-art (SoTA) LMMs, including both proprietary and the latest open-source models such as GPT-5, Qwen2.5-VL, InternVL3/3.5, MiMo-VL, and Seed-1.6-VL. Experimental results demonstrate that even the SoTA model GPT-5 averages only 0.57 on code-based evaluation and 0.22 on chart-quality assessment across the editing tasks, underscoring the difficulty of Chart2Code. We anticipate this benchmark will drive advances in multimodal reasoning and foster the development of more robust and general-purpose LMMs. Our code and data are available on Chart2Code.
Jiahao Tang, Henry Hengyuan Zhao, Lijian Wu +8
CSU-JPG, Central South University · National University of Singapore · Nanyang Technological University
While Large Language Models (LLMs) have substantially advanced text-to-code synthesis, many real programming tasks specify intent through visual artifacts such as screenshots, charts, vector drawings, videos, and interactive states. These tasks require models to connect visual perception to executable programs, because correctness depends not only on syntax but also on layout, data semantics, interaction behavior, and domain-specific constraints that apply after execution. This survey examines Multimodal Code Intelligence, covering systems that generate, edit, refine, or reason with code under visually grounded inputs and outputs. We first formulate the field by the role that code plays in each task, distinguishing code as a rendered artifact, an editable symbolic structure, a scientific representation, an intermediate reasoning trace, or an executable policy or tool interface. We then organize benchmarks and methods into four domains: Graphical User Interface, Scientific Visualization, Structured Graphics, and Frontier Tasks and Frameworks. This taxonomy connects mature artifact-generation problems to emerging agentic and unified settings and allows us to compare how different tasks treat evidence of correctness. Looking ahead, we argue that future research may benefit from four verification-centered directions. Multi-signal validation can combine complementary evidence of correctness, multi-state verification can test behavior across execution trajectories, cross-task transfer testing can probe reusable visual-code skills, and verifiable agent traces can reveal whether agent actions are grounded in visual evidence. Together, these directions may move this field from single-output imitation toward evidence-grounded executable systems. An ongoing project and resources are available on GitHub.
Xuanle Zhao, Qiushi Sun, Jingyu Xiao +16
Meituan · The University of Hong Kong · The Chinese University of Hong Kong +7
Image-to-code generation tests whether a vision-language model (VLM) can recover the structure of an image enough to express it as executable code. Existing benchmarks either focus on narrow visual domains, depend on paired executable reference code, or rely on generic rubrics that miss domain-specific reconstruction errors. We introduce Vision2Code, a reference-code-free benchmark and evaluation framework for multi-domain image-to-code generation. Vision2Code contains 2,169 test examples from 15 source datasets that span charts and plots, geometry, graphs, scientific imagery, documents, and 3D spatial scenes. Models generate executable programs, which we render and score against the source image using a VLM rater with dataset-specific rubrics and deterministic guardrails for severe semantic failures. We report render-success diagnostics that separate code execution failures from reconstruction quality. Human validation shows that this evaluation protocol aligns better with human judgments than either a generic visual rubric or embedding-similarity baselines. Across nine open-weight and proprietary models, we find that image-to-code performance is domain-dependent: leading models perform well on regular chart- and graph-like visuals but remain weak on spatial scenes, chemistry, documents, and circuit-style diagrams. Finally, we show that evaluator-filtered model outputs can serve as training data to improve image-to-code capability, with Qwen3.5-9B improving from 1.60 to 1.86 on the benchmark without paired source programs. Vision2Code provides a reproducible testbed for measuring, diagnosing, and improving image-to-code generation. Our code and data are publicly available at https://image2code.github.io/vision2code/.