Scientific figures often encode quantitative results that are not readily available in machine-readable form, making accurate plot digitization important for verifying and reusing published findings. Yet it remains unclear how accurately current models recover plotted values from real scientific figures, as existing benchmarks rely largely on synthetic charts or cover only a limited range of chart types. We introduce PlotGround, an automated pipeline for building plot digitization benchmarks from real scientific figures and their author-released source data. PlotGround maps figures to source tables, identifies reconstructable panels, and generates quantitative questions with source-grounded reference values. We use PlotGround to construct PlotGround-1k, a human-verified benchmark of 1,119 questions from 1,066 bioRxiv preprints. Across sixteen multimodal models, the best reaches 87.5% accuracy at a ±5% relative-error tolerance. Tightening the tolerance to ±2% lowers every model's accuracy by 11-24 percentage points, revealing a gap between approximate visual reading and precise quantitative recovery. PlotGround's paired figure-source structure lets us compare how accurately the same values are recovered from figures and from source tables. Providing source tables instead of figures raises a coding agent's accuracy from 90.0% to 97.4% while cutting cost by 72%.
Figures & tables
Figure 1: Overview of the PlotGround pipeline. Given a bioRxiv paper organized as a filesystem (manuscript, figures, and supplementary source tables), Mapping proposes candidate figure–table bindings from textual evidence such as captions and in-text references; Grounding retains only panels whose values can be reliably reconstructed and records a reconstruction recipe, rejecting ambiguous ones; Synthesis follows the recipe, validates the recomputed values, and writes questions using only figure-visible labels. For clarity, the example shows a single panel with question text abbreviated. Real figures are typically multi-panel and are given to models in full (Appendix A.2 ).
Stage
Step
Retained
Collection
Index 2025 preprints with supplements
37,635 papers
Filtering
Require a tabular supplement
10,540 papers
Mapping
Bind figures to source tables (Section 3.2 )
28,536 mappings (8,197 papers)
Downsampling
Randomly sample papers, one mapping per paper
7,487 mappings
Grounding
Admit reconstructable panels (Section 3.3 )
3,894 panels (1,568 figures)
Synthesis
Generate and filter questions (Section 3.4 )
1,205 questions
Table 1: PlotGround pipeline funnel. Counts are given in the unit of each stage (papers, mappings, panels, or questions), since one paper can yield multiple figure–source mappings and one figure multiple panels.
Model
Acc. ( ± 2%)
Acc. ( ± 5%)
Avg. Cost (¢)
Avg. Output tokens
GPT-5.6 Sol
73.8
87.5
1.013
419
GPT-5.6 Luna
69.7
83.8
0.113
447
GPT-5.6 Terra
66.6
82.1
0.868
231
Gemini 3.7 Flash
65.7
81.2
0.428
909
Gemini 3.8 Flash
65.5
79.8
0.933
2,257
Qwen3.8 Max
57.7
77.9
2.945
4,269
Table 2: Model accuracy on PlotGround-1k (1,119 questions). A prediction counts as correct if its relative error with respect to the gold value is within the tolerance. All models were run with their default settings, including reasoning configuration. Cost and output tokens are averaged over all questions; output tokens include reasoning tokens where the provider reports them.
Figure 2: Accuracy versus inference cost on PlotGround-1k. Each point represents a model, with colors and marker shapes indicating developers. Accuracy is measured at a ±5% tolerance, and mean inference cost per question is shown in cents on a logarithmic scale. The dashed line connects Pareto-optimal models, whose labels are shown in bold. Higher cost does not necessarily yield higher accuracy: GPT-5.6 Luna achieves 83.8% accuracy at 0.11 cents per question, compared with 62.5% at 8.37 cents for Kimi K3.
Figure 3: Representative qualitative failure cases. (a) On a log axis, a visually small discrepancy in position can correspond to a substantial numerical error; the model predicts 300 for a gold value of 252.59. (b) Models often snap predictions to labeled tick marks, failing to account for the target’s position between or beyond ticks. For a bar with gold value 4.99, eight of sixteen models predict the final labeled tick at 4.
Condition
Accuracy ( ± 5%)
Avg. Cost ($)
Avg. Latency (s)
Figure (full text + figures)
90.0
0.408
139
Table (full text + source tables)
97.4
0.113
51
Both (full text + figures + source tables)
98.2
0.115
51
Table 3: Effect of evidence representation on agent performance. Claude Code (Claude Sonnet 5) answers the same 500 rewritten questions under three evidence conditions: figures, source tables, or both. Accuracy is measured at a ±5% relative-error tolerance; cost and latency are averaged over all questions.
Figure 4: Example of a discrepancy between a published figure and its deposited source data. The published GO enrichment plot differs from the values reconstructed from the deposited source data. The discrepancy is clearest for the most significant term. The deposited adjusted p -value for vesicle-mediated transport ( 5.27×10−9 ) corresponds to −log10=8.28 , but the bar shown for this term ends at approximately 7.7, short of the axis limit at 8. Axis clipping cannot explain the gap, since a clipped bar would extend to the limit.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Chart type
Questions
Share (%)
Bar
423
37.8
Volcano
268
23.9
Scatter
174
15.5
Box
112
10.0
Line
86
7.7
Violin
27
2.4
Appendix
Table 4: Chart-type distribution of PlotGround-1k. The benchmark contains 1,119 questions across 10 chart types. Bar, volcano, and scatter plots together account for 77.3% of the questions. The remaining types include line plots, distribution plots, and survival curves.
Figure 5: Accuracy by chart type. Per-model accuracy at a ± 5% relative tolerance for the models in Table 2 . Columns show the five most common chart types, ordered by mean accuracy across models. Among these five types, either scatter or volcano plots yield the lowest accuracy for each of the 16 models. This may reflect the difficulty of locating a labelled target point among many others.
Figure 6: Additional example of a discrepancy between a published figure and its deposited source data. The published survival curves differ from the values reconstructed from the deposited individual-level records. Although the discrepancy is visible throughout the trajectory, it is clearest at the endpoint: the published figure shows the blue curve ( S/S,+/+ ) reaching zero survival first, whereas the deposited records imply zero survival for S/S,+/DEL . Survival at day 28 is the Kaplan–Meier estimate; two of the three censored S/P,+/+ animals were censored on day 21, so the estimate (0.015) differs from the simple proportion.
Systematic reviews and meta-analyses often need numerical data reported only in figures, and extracting them with interactive digitisers usually requires a person to select and calibrate each figure. We present PlotPick, an open-source tool that uses vision-language models (VLMs) to extract tabular data from batches of scientific figures, and we benchmark the kind of model it calls: nine VLMs from four providers on ChartX and six of them on PlotQA, against DePlot, a dedicated chart-to-table model, with every system scored on the same items by numeric F1 (F1 over unlabelled numbers at 5% relative tolerance). On six ChartX chart types (n=300) all nine VLMs outperform DePlot in aggregate, at 79.1-96.0% against 74.3%. The lead comes mainly from box plots, where DePlot scores 24.8% against 64.2-97.3%; pooled over the other five types, seven VLMs keep a lead of 4.7 to 11.6 points and the two weakest do not. On a subset of the PlotQA test split (n=529; 427 horizontal bar charts), scored by a lenient best-series variant of the metric, DePlot reaches 87.0%; the two strongest VLMs are level with it or slightly above it, and four fall 3.8 to 30.3 points below. DePlot was trained on PlotQA's training split. Both benchmarks use synthetic charts, and the metric ignores which series a value belongs to. The application itself was not evaluated: its figure detection, structured output, prompt and default model were not tested, and the one benchmarked model it offers, Claude Haiku 4.5, is one of the four below DePlot on PlotQA. Accuracy on biomedical figures has not been established, and every extracted value must be checked against its source figure. This version corrects version 1, which scored most PlotQA replies against category labels instead of plotted values; its claim that every VLM outperformed DePlot on both benchmarks is withdrawn. PlotPick is available at https://plotpick.streamlit.app/.
Tommy Carstensen
Copenhagen Research Centre for Biological and Precision Psychiatry, Mental Health Centre Copenhagen, Copenhagen University Hospital, Copenhagen, Denmark
Scientific figures are the interface through which research claims are inspected and reused, but final published panels rarely expose the data or plotting code that produced them. Recovering this hidden provenance from pixels is therefore underdetermined. We introduce SciFigure2Code, an AI-reconstructed benchmark that instead evaluates presentation recovery: generating editable Python programs that preserve how a scientific panel is arranged and read. Role-specialized Codex agents generate, execute, visually refine, and audit silver-standard presentation programs that capture geometry, visual hierarchy, encodings, annotations, and typography without claiming to recover original measurements or author source code. This reconstruction-and-audit protocol turns final published panels into auditable reference packages; the resulting resource contains 6,740 reviewed panels and SciFigureBench, a balanced 337-panel test set across 31 chart subtypes, five domains, and three complexity levels. Across 14 zero-shot models in image-only and caption-assisted settings, execution, multi-component layouts, axes, legends, and scientific labels remain weak. Claude Opus 4.7 achieves the highest image-only Overall score, Claude Opus 4.6 leads caption-assisted reconstruction, and two-stage plan-then-code prompting improves Overall for all four tested models. SciFigure2Code provides an auditable testbed for agents that construct editable, visually faithful scientific figure presentations.
Scientific figures often encode the visual evidence behind scientific findings, yet figure plagiarism remains underexplored as a benchmarked multimodal evaluation problem. We present SciFigPlag-Bench, a benchmark for provenance-aware reasoning over scientific figures in scholarly documents. Unlike general image-similarity or image-forensics benchmarks, SciFigPlag-Bench evaluates whether a suspicious figure reuses evidence from a specific source figure, how the reused content has been transformed, and where the reused evidence appears. We introduce a factorized taxonomy that separates what is reused from how it is transformed, covering material-preserving reuse, such as full-figure and subfigure reuse, as well as abstract-content reuse, such as data re-expression and structural redraw. Guided by this taxonomy, we construct a hybrid benchmark with 2,582 positive pairs and 2,541 negative pairs, combining documented real-world cases, taxonomy-guided synthetic examples, and visually similar negatives. The benchmark supports four diagnostic tasks: pairwise detection, source attribution, hierarchical reuse-type classification, and reuse correspondence localization. Experiments with diverse vision-language models establish initial baselines and reveal persistent challenges in fine-grained provenance reasoning, reuse-type understanding, and spatial evidence grounding.
Zhiying Cui, Minghao Yang, Linlin Gao +2
1Ningbo University · 2Hokkaido University · 3IBM Research