We introduce Chart2Code, a new benchmark for evaluating the chart understanding and code generation capabilities of large multimodal models (LMMs). Chart2Code is explicitly designed from a user-driven perspective, capturing diverse real-world scenarios and progressively increasing task difficulty. It consists of three levels: Level 1 (Chart Reproduction) reproduces charts from a reference figure and user query; Level 2 (Chart Editing) involves complex modifications such as changing chart types or adding elements; and Level 3 (Long-Table to Chart Generation) requires models to transform long, information-dense tables into faithful charts following user instructions. To our knowledge, this is the first hierarchical benchmark that reflects practical chart2code usage while systematically scaling task complexity. In total, Chart2Code contains 2,023 tasks across 22 chart types, paired with multi-level evaluation metrics that assess both code correctness and the visual fidelity of rendered charts. We benchmark 25 state-of-the-art (SoTA) LMMs, including both proprietary and the latest open-source models such as GPT-5, Qwen2.5-VL, InternVL3/3.5, MiMo-VL, and Seed-1.6-VL. Experimental results demonstrate that even the SoTA model GPT-5 averages only 0.57 on code-based evaluation and 0.22 on chart-quality assessment across the editing tasks, underscoring the difficulty of Chart2Code. We anticipate this benchmark will drive advances in multimodal reasoning and foster the development of more robust and general-purpose LMMs. Our code and data are available on Chart2Code.
Figures & tables
Figure 1 : Chart2Code covers three progressively challenging levels : reproduction, editing, and long-table to chart generation. It provides a user-driven and diverse benchmark that better reflects real-world chart2code demands.
Table 2: Evaluation results on Chart Reproduction (Level 1) with various LMMs. Each task includes a reference chart as input. DR : input without the any data. CRD : input with customized text-format table data. CFD : input with customized figure-format table data. Exec.Rate : execution rate(%); Base : the multi-dimensional weighted Base-Score. The Base score uses the same dimension-wise weighting scheme as the LLM-score( LLM ). More detail Metrics in Appendix 12 . We use GPT-5-mini as the base model for both LLM-score and LMM-score( LMM ); Base , LLM , LMM : scores on a 0–100 scale (higher is better).
Model
Exec. Rate
Code-Level
Chart-Level
Color
Grid
Layout
Legend
Visual
Data
Text
Type
Base
LLM
LMM
Proprietary
Gemini-3-Pro
97.23
52.32
75.31
86.45
63.33
81.58
62.75
77.16
93.86
70.78
72.21
33.41
Claude-Sonnet-4
90.20
47.17
65.29
55.32
56.51
81.50
54.88
80.52
93.29
63.65
66.45
25.40
GPT-5.2
96.04
58.44
80.83
61.51
58.16
84.65
64.66
83.77
94.52
70.93
75.66
33.03
Seed-1.5-VL
60.20
44.39
69.37
52.06
52.80
79.85
52.02
77.21
92.66
61.67
65.45
18.30
Table 3: Evaluation results on Chart Editing (Level 2) with various LMMs.
Model
Exec. Rate
Code-Level
Chart-Level
Color
Grid
Layout
Legend
Visual
Data
Text
Type
Base
LLM
LMM
Proprietary
Gemini-3-Pro
30.03
56.29
83.71
92.55
68.39
84.00
62.26
80.45
96.60
74.28
77.90
35.97
Claude-Sonnet-4
46.65
35.43
68.97
84.08
61.77
74.44
46.73
67.77
89.95
61.13
62.04
16.60
GPT-5.2
20.13
34.46
82.14
80.87
52.64
78.34
43.56
76.09
93.81
61.99
68.40
16.29
Seed-1.5-VL
10.86
44.55
76.89
93.94
69.30
73.52
51.75
73.57
95.96
67.58
75.59
22.85
Table 4: Evaluation results on Long-Table to Chart task (Level 3) with various LMMs.
Figure 4 : Left : Most proprietary and open-source models generalize well on Level 1 and Level 2 tasks when calculating the LLM-score for predicted code assessment. Right : Proprietary models tend to obtain higher LMM-scores on the Level 1 task rather than the Level 2, while open-source models perform poorly on both tasks (scores are lower than 0.5).
Figure 5 : Timestamps distribution of chart sources from arxiv preprint.
Annotator
Raw Data
Chart Data
Annotator 1
20 diverse data tables and Excels.
50 charts.
Annotator 2
20 diverse data tables and Excels.
50 charts.
Annotator 3
50 diverse data figures and Excels.
50 charts.
Annotator 4
50 diverse data figures and Excels.
50 charts.
Annotator 5
50 diverse data tables and Excels.
150 charts.
Annotator 6
50 diverse data tables and Excels.
150 charts.
Table 5 : Statistics of Annotation Tasks per Annotator
Model
Exec. Rate
Code-Level
Chart-Level
Color
Grid
Layout
Legend
Visual
Data
Text
Type
Base
LLM
LMM
Proprietary
Gemini-3-Pro
97.50
61.92
83.49
84.14
87.17
83.78
68.47
93.59
93.60
78.65
75.86
45.42
Claude-Sonnet-4
96.52
43.69
51.02
77.15
80.39
74.97
55.45
85.66
88.75
65.60
56.65
32.36
GPT-5.2
97.08
65.77
84.94
83.96
87.93
85.21
69.25
93.35
93.85
79.91
77.88
43.73
Seed1.5-VL
87.34
37.85
65.79
76.16
75.22
71.77
53.27
81.05
86.40
63.85
53.84
26.40
Table 6: Details Evaluation results on level-1 Direct mimic result(Base-Score Details)
Model
Exec. Rate
Code-Level
Chart-Level
Color
Grid
Layout
Legend
Visual
Data
Text
Type
Base
LLM
LMM
Proprietary
Gemini-3-Pro
97.50
75.54
87.78
77.15
88.02
67.22
58.05
84.86
86.36
78.65
75.86
45.42
Claude-Sonnet-4
96.52
53.41
69.56
61.63
76.64
50.57
34.23
70.69
62.12
65.60
56.65
32.36
GPT-5.2
97.08
76.95
88.58
80.08
91.06
70.60
59.80
85.75
89.20
79.91
77.88
43.73
Seed1.5-VL
87.34
48.15
71.25
56.46
72.23
44.90
32.40
67.67
64.78
63.85
53.84
26.40
Table 7: Details Evaluation results on level-1 Direct mimic result(LLM-Score Details)
Model
Exec. Rate
Code-Level
Chart-Level
Color
Grid
Layout
Legend
Visual
Data
Text
Type
Base
LLM
LMM
Proprietary
Gemini-3-Pro
100.00
47.55
81.25
80.56
54.34
78.54
66.00
76.10
94.44
69.23
69.76
40.72
Claude-Sonnet-4
100.00
39.06
50.52
85.19
32.70
74.51
65.22
67.70
95.37
61.46
55.03
40.63
GPT-5.2
97.22
45.29
78.93
82.86
38.79
78.83
64.71
72.25
91.43
66.31
65.51
39.26
Seed1.5-VL
97.22
34.89
67.86
87.14
45.53
74.32
72.14
71.55
97.14
65.76
58.73
34.74
Table 8: Details Evaluation results on level-1 Customize Raw Data result(Base-Score Details)
Model
Exec. Rate
Code-Level
Chart-Level
Color
Grid
Layout
Legend
Visual
Data
Text
Type
Base
LLM
LMM
Proprietary
Gemini-3-Pro
100.00
65.56
85.06
64.58
72.50
61.81
76.06
56.67
73.75
69.23
69.76
40.72
Claude-Sonnet-4
100.00
44.03
64.58
52.08
54.86
48.19
64.44
46.67
66.94
61.46
55.03
40.63
GPT-5.2
97.22
60.71
82.14
62.29
68.00
61.43
65.71
62.00
66.43
66.31
65.51
39.26
Seed1.5-VL
97.22
37.00
72.14
55.43
61.43
45.29
76.00
52.57
74.43
65.76
58.73
34.74
Table 9: Details Evaluation results on level-1 Customize Raw Data result(LLM-Score Details)
Model
Exec. Rate
Code-Level
Chart-Level
Color
Grid
Layout
Legend
Visual
Data
Text
Type
Base
LLM
LMM
Proprietary
Gemini-3-Pro
99.07
52.32
75.31
86.45
63.33
81.58
62.75
77.16
93.86
70.78
71.12
32.85
Claude-Sonnet-4
93.52
48.84
52.64
85.48
58.65
79.04
57.41
71.10
93.30
65.27
65.99
26.44
GPT-5.2
99.07
61.93
78.66
85.51
61.32
83.12
62.66
75.46
96.95
73.02
71.42
35.40
Seed1.5-VL
79.63
48.02
59.30
87.21
57.02
74.28
57.13
75.32
92.36
65.58
64.19
19.53
Table 10: Details Evaluation results on level-1 Customize Figure Data result(Base-Score Details)
Model
Exec. Rate
Code-Level
Chart-Level
Color
Grid
Layout
Legend
Visual
Data
Text
Type
Base
LLM
LMM
Proprietary
Gemini-3-Pro
99.07
68.90
87.67
73.10
79.67
66.95
65.81
62.38
85.55
70.78
71.12
32.85
Claude-Sonnet-4
93.52
61.55
76.18
70.60
71.90
60.85
58.95
60.25
85.70
65.27
65.99
26.44
GPT-5.2
99.07
68.91
89.04
71.47
79.98
67.74
64.29
62.38
83.96
73.02
71.42
35.40
Seed1.5-VL
79.63
58.59
78.27
65.15
72.35
55.35
61.29
60.47
78.06
65.58
64.19
19.53
Table 11: Details Evaluation results on level-1 Customize Figure Data result(LLM-Score Details)
Model
Exec. Rate
Code-Level
Chart-Level
Color
Grid
Layout
Legend
Visual
Data
Text
Type
Base
LLM
LMM
Proprietary
Gemini-3-Pro
97.23
52.32
75.31
86.45
63.33
81.58
62.75
77.16
93.86
70.78
72.21
33.41
Claude-Sonnet-4
90.20
47.17
65.29
55.32
56.51
81.50
54.88
80.52
93.29
63.65
66.45
25.40
GPT-5.2
96.04
58.44
80.83
61.51
58.16
84.65
64.66
83.77
94.52
70.93
75.66
33.03
Seed1.5-VL
60.20
44.39
69.37
52.06
52.80
79.85
52.02
77.21
92.66
61.67
65.45
18.30
Table 12: Details Evaluation results on level-2 result(Base-Score Details)
Model
Exec. Rate
Code-Level
Chart-Level
Color
Grid
Layout
Legend
Visual
Data
Text
Type
Base
LLM
LMM
Proprietary
Gemini-3-Pro
97.23
85.21
89.26
86.91
78.32
69.72
48.24
73.06
87.00
70.78
72.21
33.41
Claude-Sonnet-4
90.20
79.15
81.38
81.27
75.23
61.89
31.96
68.13
75.81
63.65
66.45
25.40
GPT-5.2
96.04
86.34
90.05
87.25
80.06
71.65
52.00
74.57
89.89
70.93
75.66
33.03
Seed1.5-VL
60.20
78.37
80.30
79.64
72.17
60.46
31.42
67.87
75.55
61.67
65.45
18.30
Table 13: Details Evaluation results on level-2 result(LLM-Scores Details)
Model
Exec. Rate
Code-Level
Chart-Level
Color
Grid
Layout
Legend
Visual
Data
Text
Type
Base
LLM
LMM
Proprietary
Gemini-3-Pro
30.03
56.29
83.71
92.55
68.39
84.00
62.26
80.45
96.60
74.28
77.90
35.97
Claude-Sonnet-4
46.65
35.43
68.97
84.08
61.77
74.44
46.73
67.77
89.95
61.13
62.04
16.60
GPT-5.2
20.13
34.46
82.14
80.87
52.64
78.34
43.56
76.09
93.81
61.99
68.40
16.29
Seed1.5-VL
10.86
44.55
76.89
93.94
69.30
73.52
51.75
73.57
95.96
67.58
75.59
22.85
Table 14: Details Evaluation results on level-3 details result(Base-Score Details)
Model
Exec. Rate
Code-Level
Chart-Level
Color
Grid
Layout
Legend
Visual
Data
Text
Type
Base
LLM
LMM
Proprietary
Gemini-3-Pro
30.03
82.91
91.70
83.33
86.65
78.69
54.68
74.70
88.72
74.28
77.90
35.97
Claude-Sonnet-4
46.65
65.68
78.97
71.95
78.97
62.71
32.16
57.78
74.32
61.13
62.04
16.60
GPT-5.2
20.13
70.00
84.60
78.10
81.35
68.10
42.51
69.10
77.78
61.99
68.40
16.29
Seed1.5-VL
10.86
77.65
87.06
81.74
90.88
72.94
57.06
69.79
84.12
67.58
75.59
22.85
Table 15: Details Evaluation results on level-3 detail result(LLM-Score Details)
Figure 6 : Correlation Analysis on Level-1 DR: Base-Score vs. LLM-Score (Code-Level Metrics).
Figure 7 : Correlation between Code-level and Chart-level on Level-1 DR
Figure 8 : Correlation Analysis on Level-1 CRD: Base-Score vs. LLM-Score (Code-Level Metrics).
Figure 9 : Correlation between Code-level and Chart-level on Level-1 CRD
Figure 10 : Correlation Analysis on Level-1 CFD: Base-Score vs. LLM-Score (Code-Level Metrics).
Figure 11 : Correlation between Code-level and Chart-level on Level-1 CFD
Figure 36 : Human vs model evaluate system.
Figure 37 : Human vs. model evaluation across all tasks: correlation between human evaluation and LMM Score.
Figure 38 : Analysis of model performance on different task cases with LLM-score and LMM-score.
Figure 39 : Correlation of the model performance (i.e, LMM-score) on different manually annotated difficulty levels (i.e., Easy, Medium, Hard) on Level 1, 2, 3, respectively.
Figure 40 : Human vs model performance: LLM-score and LMM-score across level 1 direct tasks.
Figure 41 : Human vs model performance: LLM-score and LMM-score across level 1 customize tasks.
Figure 42 : Human vs model performance: LLM-score and LMM-score across level 1 figure tasks.
Figure 43 : Human vs model performance: LLM-score and LMM-score across level 2 tasks.
Figure 44 : Human vs model performance: LLM-score and LMM-score across level 3 tasks.
Model
Version/HF Checkpoint
Do Sample
level 1 2 Max
level 3 Max
Temp.
Top-P
Proprietary Multimodal Large Language Models
GPT-5.2 [ 22 ]
gpt-5.2-2025-12-11
default
55000
0.1
1
Claude 4 Sonnet [ 1 ]
claude-4-sonnet-20250523
default
55000
0.1
1
Gemini-3-Pro [ 5 ]
gemini-3-pro-20251118
default
55000
0.1
1
doubao-seed-1-5 [ 9 ]
seed1.5-VL-20250513
default
16000
0.1
1
doubao-seed-1-6 [ 26 ]
seed1.6-VL-20250625
default
32768
0.1
1
Table 19 : Run configurations for all models. Unset values indicate that their default values are being used. For Proprietary models, we are unable to use a Top-P of exactly 1 due to their API settings, and we end up using a value of 0.99999 . Temp. denotes temperature. We use model pages’ code to set up the run configurations whenever possible.
Model
Vision
Language
Resolu-
Encoder
Model
tion
Qwen2-VL-7B
Qwen2-VL ViT-14-224
Qwen2-VL-LLM-7B
origianl
Qwen2-VL-72B
Qwen2-VL ViT-14-224
Qwen2-VL-LLM-72B
origianl
Qwen2.5-VL-7B
Qwen2.5-VL ViT-14-224
Qwen2.5-VL-LLM-7B
origianl
Qwen2.5-VL-72B
Qwen2.5-VL ViT-14-224
Qwen2.5-VL-LLM-72B
origianl
Deepseek-VL-7B
SigLIP-384-SO400M &
DeepSeek-LLM-7B
1152×1152 *
Table 20 : We summarize the visual and language components of the open-source models evaluated in our benchmark, along with the input resolutions used in our evaluation. Here, original denotes that we use the default image size, as the corresponding models support dynamic resolution inputs. Note that for DeepSeekVL-7B and GLM-4-9B , we apply a maximum input size constraint to accommodate their requirements.
Table 21 : The release time and model source of LMMs used in our benchmark.
Name
Model License
Code License
GPT-5.2
Proprietary
Proprietary
Claude 4 Sonnet
Proprietary
Proprietary
Gemini-3-Pro
Proprietary
Proprietary
doubao-seed-1.6
Proprietary
Proprietary
doubao-seed-1.5
Proprietary
Proprietary
Qwen2-VL-7B
qwen
Apache 2.0
Table 22 : Summary of licenses in models that are evaluated in Chart2Code. Entries marked with “Not Applicable” indicate that authors do not have an explicit code license displayed within the codebase or model checkpoint page.
Building strong chart-to-code systems increasingly relies on reinforcement learning, whose effectiveness depends critically on the quality of the reward signal. Large Multimodal Models (LMMs) play a natural critical role in jointly assessing chart visual appearance and task requirements. They are therefore increasingly used as visual critics and reward models, yet their reliability as judges remains largely unexplored. To this end, we introduce ChartJudgeBench, a diagnostic vision-language benchmark for assessing LMM judges in chart-to-code workflows. It includes 1,003 Chart Perception Alignment (CPA) instances for pairwise chart comparison and 650 Chart Reasoning Judgment (CRJ) instances for binary Accept/Reject verification in Chart Reproduction and Chart Editing. Together, these tasks emulate the core judging decisions required in agentic refinement and RL-based chart optimization. Our evaluation of strong LMMs reveals four systematic limitations: (i) positional bias in pairwise comparison, (ii) a strong tendency to overpredict Accept, (iii) difficulty in matching visual styles and aesthetics, and (iv) an unexpected leniency bias in RL-trained models. These findings show that current LMM judges require explicit reliability validation before being used as critics or reward models in chart-to-code optimization. The code and data are available on ChartJudgeBench.
Lijian Wu, Henry Hengyuan Zhao, Zijian Zhang +3
CSU-JPG, Central South University · National University of Singapore
Annotations are essential to communicative visualization, helping explain data, emphasize key findings, and guide attention. While multimodal large language models (MLLMs) offer new opportunities for automatic chart annotation authoring, their capabilities in this task remain underexplored. To address this gap, we introduce ChartAnno, a comprehensive benchmark for evaluating MLLMs on chart annotation generation. ChartAnno contains 1,200 real-world charts with paired annotated and unannotated executable code, along with 3,600 annotation instructions spanning three levels of specificity. We also develop a multidimensional evaluation framework combining rule-based and LLM-judged metrics to assess execution, structural compliance, semantic consistency, and design effectiveness. We evaluate 10 representative MLLMs under two primary chart input settings: (1) chart code alone and (2) both code and chart image. Results reveal that proprietary models lead overall, though open-source models narrow the gap. While higher instruction specificity improves annotation quality, inferring abstract communicative intent remains difficult across all models. Providing chart images yields marginal benefit when code is available. We also examine the effect of chart code through an image-only ablation and analyze the effects of multiple task complexity indicators and instruction-level transitions. Further analyses characterize common failure modes and validate the reliability of the LLM-based judge. Experiments with D3 and SVG demonstrate the generalizability of ChartAnno beyond its primary Python setting.
Zhenghan Chen, Zekai Shao, Lidan Tan +10
Fudan University · Sun Yat-sen University · The Hong Kong University of Science and Technology (Guangzhou) +1
Multimodal large language models (MLLMs) are rapidly evolving with expanded context windows and stronger reasoning capabilities, enabling multi-chart understanding and multi-step inference. These abilities are increasingly important as MLLMs are adopted in complex agentic tasks. However, existing benchmarks largely emphasize single-chart perception, while simple chart-to-chart connections are insufficient to evaluate these capabilities. To capture multi-chart complexity while ensuring consistency and validity, we design a synthesis pipeline supported by latent graphs. Building on this pipeline, we introduce LongChart, a benchmark whose VQA sets contain an average of 6.5 images and 31.2 questions. We evaluate 10 state-of-the-art MLLMs and examine three factors that influence performance: reasoning patterns, auxiliary tools, and robustness to image perturbations. Our results show that MLLM accuracy decreases and varies substantially as computational complexity increases, highlighting directions for future research in multi-chart reasoning.