ChartRevise: A Dataset and Evaluation Protocol for Exact Chart Editing via Code
Organizations: University of Minnesota, Twin-Cities · IBM Research · Horizon School of Digital Technologies
Abstract
Chart editing requires cross-modal edit grounding, realizing a requested visual change in the code that draws it, with necessary related updates and without altering unrelated content. Existing benchmarks emphasize either code executability or chart quality, but their metrics do not clearly distinguish request completion from missed coupled updates and gratuitous changes. We introduce ChartRevise, a structured dataset and evaluation protocol for exact program-grounded chart editing. For dataset construction, we build on the grammar of graphics to systematically cover chart-editing operations, using source-program checks to verify their applicability across chart types and libraries. To improve edit exactness, our pipeline checks individual requirements and guides repair or exclusion when they are unmet. The resulting dataset contains 92,438 records covering 344 edit types across 20 chart types and three plotting libraries. For evaluation, our reference-free protocol separately measures atomic requirement completion, identifies gratuitous changes, and detects missed coupled updates. These checks are combined with successful execution and rendering to determine exact-edit success. Across five models and four external benchmarks, fine-tuning yields relative gains of 16% in mean requirement recall and 22% in mean exact-edit rate.
Figures & tables
| Resource | Dataset scope | Construction checks | Evaluation checks | ||||
|---|---|---|---|---|---|---|---|
| Tasks | Chart types | Instruction types | Atomic requirements | Preservation | Coupled updates | ||
| ChartEdit | 1,405 | 19 | 6 | Human (request, code, chart) | — | — | — |
| ChartM 3 | 1K eval. 24K train | 10 | — | Model (charts) | — | ✓ (code-space) | — |
| ChartEditVista | 7,964 | 31 | 6 | Model (charts) Human (code, charts) | — | — | — |
| ChartEditBench | 4,142 | 37 | 35 | Model (code, charts) Program (assertions) | ✓ (assertions) | — | — |
| Chart2Code L2 | 1,010 | 19 | — | — | — | — | — |
| Benchmark | Family | Requirement recall (CR μ ) | Full completion | Gratuitous per task | Coupled-update recall | Exact-edit rate |
|---|---|---|---|---|---|---|
| ChartEdit 1,405 tasks | Phi-3.5-V | .773 +.138 | .674 +.155 | .346 -.205 | .540 +.076 | .425 +.119 |
| InternVL2.5 | .707 +.234 | .593 +.215 | .373 -.007 | .500 +.188 | .378 +.162 | |
| Qwen3.5-4B | .867 +.047 | .804 +.050 | .165 -.091 | .688 +.004 | .572 +.055 | |
| Qwen3.5-9B | .895 +.068 | .846 +.085 | .135 +.021 | .731 +.064 | .616 +.032 | |
| Granite | .812 +.114 | .733 +.128 | .459 -.482 | .622 +.077 | .439 +.168 | |
| rel. of mean | +17.4% | +21.0% | -34.1% | +15.3% | +28.3% |
| Family | ChartEdit | Chart2Code-L2 | ChartM 3 | ChartSync | ||||
|---|---|---|---|---|---|---|---|---|
| Exec. | Code. | Exec. | Code. | Exec. | Code. | Exec. | Code. | |
| Phi-3.5-V | 0.860 [-1pt] +0.000 | 75.4 [-1pt] +11.6 | 0.334 [-1pt] -0.066 | 38.0 [-1pt] +10.4 | 0.650 [-1pt] -0.010 | 46.3 [-1pt] +11.2 | 0.930 [-1pt] -0.010 | 85.5 [-1pt] +1.4 |
| InternVL2.5 | 0.860 [-1pt] -0.040 | 71.6 [-1pt] +20.4 | 0.380 [-1pt] -0.214 | 32.6 [-1pt] +11.9 | 0.740 [-1pt] +0.050 | 43.3 [-1pt] +10.6 | 0.930 [-1pt] -0.030 | 88.2 [-1pt] +8.4 |
| Qwen3.5-4B | 0.910 [-1pt] +0.010 | 86.7 [-1pt] +3.6 | 0.444 [-1pt] +0.030 | 62.1 [-1pt] +1.6 | 0.700 [-1pt] +0.040 | 66.8 [-1pt] +13.1 | 0.980 [-1pt] -0.010 | 92.6 [-1pt] +0.3 |
| Qwen3.5-9B | 0.920 [-1pt] -0.020 | 89.8 [-1pt] +6.0 | 0.540 [-1pt] -0.027 | 73.4 [-1pt] +6.7 | 0.720 [-1pt] -0.010 | 74.8 [-1pt] +13.6 | 0.980 [-1pt] -0.010 | 94.8 [-1pt] -0.9 |
| Granite | 0.860 [-1pt] -0.010 | 79.7 [-1pt] +11.2 | 0.373 [-1pt] -0.223 | 43.1 [-1pt] +13.4 | 0.630 [-1pt] -0.070 | 48.4 [-1pt] +4.5 | 0.900 [-1pt] +0.030 | 83.4 [-1pt] +3.7 |
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Category | Semantic objects |
|---|---|
| data | dataset, field, record, series, category |
| transform | aggregation, bin, stack, derived metric, sort, window |
| encoding | position, colour, size, shape, opacity, and text channels |
| layer | primary and overlay layers, error bars, reference line or band, secondary axis |
| mark | line, point, bar, area, arc, box, violin, cell |
| position | margins, anchor, z-order, aspect ratio, legend and annotation placement |
| Edit | Chart | Template realized instruction |
|---|---|---|
| E026 | Bubble (seaborn) | Add subgroup by {subgroup_field} within {primary_group_field} “For the subgroup, add subgroup by ‘Crop‘ within ‘Region‘ .” |
| E311 | Bubble (seaborn) | Map {facet_row_field} to rows and {facet_column_field} to columns “For the facet layout, map ‘Region‘ to rows and ‘Crop‘ to columns.” |
| E308 | Pie (matplotlib) | Set the multi-facet grid to {facet_rows} rows and {facet_columns} columns “For the subplot grid, set the multi-facet grid to 2 rows and 2 columns.” |
| E100 | Bubble (plotly) | Remove facet mapping for {facet_field} “For the facet channel, remove facet mapping for ‘Gender‘ .” |
| Category | Train | Test | Total |
|---|---|---|---|
| guide | 14,356 | 253 | 14,609 |
| data | 11,223 | 218 | 11,441 |
| style | 10,994 | 252 | 11,246 |
| position | 8,521 | 223 | 8,744 |
| annotation | 8,385 | 187 | 8,572 |
| transform | 7,171 | 177 | 7,348 |
| Subset | Audited units | Correct | Correctness | 95% CI |
|---|---|---|---|---|
| (a) Task level: records | ||||
| Single-action | 300 | 255 | 85.0% | [80.5, 88.6] |
| Compound | 200 | 139 | 69.5% | [62.8, 75.5] |
| Weighted | 500 | — | 84.2% | [80.4, 88.1] |
| (b) Requirement level: compound requirements | ||||
| Compound | 1,071 | — | 95.6% | [94.2, 96.7] |
| Phi-3.5-V | InternVL2.5-8B | Qwen3.5-4B | Qwen3.5-9B | Granite-4.1-4B | |
|---|---|---|---|---|---|
| Effective batch | |||||
| Optimizer steps | 2,830 | 2,805 | 1,416 | 1,416 | 1,416 |
| Warmup steps | 20 | 20 | 60 | 60 | 60 |
| Optimizer | AdamW (8-bit) | AdamW (8-bit) | AdamW (8-bit) | AdamW | AdamW (8-bit) |
| Max sequence length | 8,192 | 6,144 | 8,192 | 8,192 | 8,192 |
| Image input | 1 crop | 6 tiles + thumbnail | 1,280 visual tokens | 1,280 visual tokens |
| Metric | Source | Meaning / use |
|---|---|---|
| ExecRate | Mechanical | Fraction of all tasks whose generated program executes and renders. |
| CodeScore | LLM judge | Continuous editing-quality score over the generated program; can score non-running programs. |
| ChartScore | LLM judge | Overall rendered-chart quality; not target-edit success. |
| AppliedRate | LLM judge | Whether the requested edit reached the named target. |
| EditFidelity | LLM judge | Applied edit with no unrequested change, unintended reposition, or destroyed object. |
| Requirement recall (CR μ ) | GT-free judge | Micro-aggregated fraction of explicit requirements satisfied across all tasks. |
| Benchmark | Family | ExecRate | CodeScore | ChartScore | 8D F1 | LLM score | LMM score |
|---|---|---|---|---|---|---|---|
| ChartEdit | Phi-3.5-V | ||||||
| InternVL2.5 | |||||||
| Qwen3.5-4B | |||||||
| Qwen3.5-9B | |||||||
| Granite | |||||||
| Chart2Code-L2 | Phi-3.5-V |
| Benchmark | Family | AppliedRate | EditFidelity | CR μ | CR M |
|---|---|---|---|---|---|
| ChartEdit | Phi-3.5-V | .625 +.127 | .389 +.121 | .773 +.138 | .758 +.145 |
| InternVL2.5 | .573 +.224 | .345 +.143 | .707 +.234 | .703 +.241 | |
| Qwen3.5-4B | .711 +.024 | .470 +.052 | .867 +.047 | .871 +.049 | |
| Qwen3.5-9B | .743 +.059 | .498 +.022 | .895 +.068 | .899 +.073 | |
| Granite | .666 +.131 | .384 +.110 | .812 +.114 | .811 +.117 | |
| Chart2Code-L2 | Phi-3.5-V | .234 +.136 | .032 +.015 | .691 +.120 | .699 +.156 |
| Metric | Phi-3.5-V | InternVL2.5 | Qwen3.5-4B | Qwen3.5-9B | Granite |
|---|---|---|---|---|---|
| TESR | |||||
| VLCS | |||||
| BFS | |||||
| OCR F1 | |||||
| SSIM |
| Benchmark | ExecRate | CodeScore | ChartScore | AppliedRate | EditFidelity |
|---|---|---|---|---|---|
| ChartEdit | 1.00 | 95.41 | 99.88 | 0.796 | 0.552 |
| Chart2Code L2 | 0.97 | 98.04 | 95.92 | 0.757 | 0.455 |
| ChartSync | 0.99 | 99.79 | 98.58 | 0.956 | 0.901 |
| ChartM 3 | 0.91 | 89.52 | 89.26 | 0.345 | 0.297 |
| Benchmark | Pairs | Agreement | ||
|---|---|---|---|---|
| ChartEdit | 69 | 91% | ||
| ChartM 3 | 50 | 91% | ||
| ChartSync | 64 | 98% | ||
| Chart2Code-L2 | 17 | 91% | ||
| All | 200 | 93% |
| Benchmark | Model | Training data | Requirement recall | Full completion | Gratuitous per task | Coupled-update recall | Exact-edit rate |
|---|---|---|---|---|---|---|---|
| ChartEdit | Phi-3.5-V | Before filtering | .764 | .670 | .348 | .527 | .418 |
| Released | .773 | .674 | .346 | .540 | .425 | ||
| Qwen3.5-4B | Before filtering | .873 | .804 | .140 | .666 | .566 | |
| Released | .867 | .804 | .165 | .688 | .572 | ||
| Chart2Code-L2 | Phi-3.5-V | Before filtering | .689 | .034 | 1.404 | .489 | .001 |
| Released | .691 | .032 | 1.433 | .493 | .002 |