SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
Organizations: The Chinese University of Hong Kong · Shanghai Artificial Intelligence Laboratory · New York University · Fudan University · Shanghai Jiao Tong University · Harbin Institute of Technology · JD Explore Academy · Centre for Perceptual and Interactive Intelligence (CPII) Limited
Abstract
Scientific papers require models to integrate evidence across text, equations, figures, tables, code, and datasets while preserving its provenance. Beyond answer correctness, scientific reading requires verifiable outputs from operations such as evidence localization, definition extraction, and consistency checking. We introduce SciDocBench, a workflow-centered benchmark targeting these operations through 124 expert-authored and difficulty-screened questions across seven research-assistant capability groups, 19 subtasks, and five scientific domains. Each question is instantiated in four matched settings formed by pairing its bilingual variants with the All Images First and Markdown Interleaved document representations, yielding 496 evaluation instances. The strongest evaluated model, Claude-Opus-5, scores 62.6 out of 100, with remaining gaps in evidence localization, structured information extraction, cross-document synthesis, and robustness to document representation. To convert these diagnostics into scalable training signals, we introduce SciDocIR, a structured representation of scientific document objects, layout and cross-reference relations, and provenance. Using SciDocIR, we construct SciDocDataset, which contains 4K supervised fine-tuning instances and 10K reinforcement-learning instances, built on 14 verifiable training subtasks. Post-training Qwen3.6-27B on task-aligned data improves its SciDocBench score from 40.03 to 45.33 with supervised fine-tuning and to 45.74 with subsequent reinforcement learning. Both adapted models preserve DocVQA and InfoVQA performance and improve ChartQA accuracy over the original model by 0.80 and 3.40 points, respectively. Together, SciDocBench, SciDocIR, and SciDocDataset connect capability diagnosis with verifiable training-data construction for scientific-document assistants.
Figures & tables
| Benchmark | Primary focus | Full paper or PDF | Scientific objects | Cross- doc. | Code/data prov. |
| DocVQA [ 2021 ] | Document-image question answering | ✗ | ✗ | ✗ | ✗ |
| QASPER [ 2021 ] | Evidence-grounded full-paper QA | ✓ | – | ✗ | ✗ |
| ChartQA [ 2022 ] | Chart question answering | ✗ | – | ✗ | ✗ |
| SCITAB [ 2023 ] | Scientific-table claim verification | ✗ | ✓ | ✗ | ✗ |
| MMLongBench-Doc [ 2024 ] | Long-context document multimodal QA | ✓ | – | ✗ | ✗ |
| DocGenome [ 2024b ] | Scientific-document structure and QA | ✓ | ✓ | ✗ | ✗ |
| Model | Overall | English | Chinese | Capability scores | ||||||||
| Images | Inter- | Images | Inter- | Document | Information | Evidence | Cross- | Reconstr. | Paper– | Dataset | ||
| leaved | leaved | Perception | Extraction | Verification | Document | & Exec. | Code | Underst. | ||||
| Claude-Opus-5 [ 5 ] | 62.60 | 63.19 | 61.34 | 65.21 | 60.68 | 61.37 | 55.52 | 74.47 | 62.81 | 59.65 | 51.75 | 43.75 |
| Claude-Opus-4.8 [ 4 ] | 53.41 | 53.63 | 52.43 | 54.13 | 53.46 | 46.59 | 57.12 | 58.63 | 56.65 | 57.09 | 44.00 | 53.12 |
| Claude-Sonnet-4.6 [ 6 ] | 49.57 | 52.25 | 51.18 | 53.62 | 41.21 | 44.80 | 51.51 | 54.25 | 47.43 | 52.64 | 52.17 | 59.38 |
| GPT-5.6-Sol [ 86 ] | 61.00 | 61.07 | 60.32 | 61.32 | 61.31 | 54.08 | 56.78 | 78.16 | 54.62 | 69.70 | 46.69 | 59.38 |
| Model | SciDocBench | General document understanding | |||||||||||||
| Overall | Capability scores | English | Chinese | DocVQA | InfoVQA | ChartQA | |||||||||
| Document | Information | Evidence | Cross- | Reconstr. | Paper– | Dataset | Images | Inter- | Images | Inter- | |||||
| Perception | Extraction | Verification | Document | & Exec. | Code | Underst. | leaved | leaved | |||||||
| Qwen3.6-27B | 40.03 | 38.14 | 38.55 | 44.71 | 38.46 | 42.28 | 40.31 | 34.38 | 44.61 | 38.61 | 35.38 | 41.53 | 95.81 | 90.67 | 74.60 |
| + SFT | 45.33 | 40.11 | 43.55 | 51.40 | 43.85 | 52.48 | 58.94 | 40.62 | 46.48 | 42.76 | 45.89 | 46.20 | 95.80 | 91.02 | 75.40 |
| Improvement | +5.30 | +1.97 | +5.01 | +6.69 | +5.39 | +10.20 | +18.62 | +6.25 | +1.87 | +4.15 | +10.51 | +4.67 | -0.01 | +0.36 | +0.80 |
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
| Statistic | Value |
| Data Source | |
| Source channels | arXiv, bioRxiv, OpenReview, PubMed |
| Source papers | 116 |
| Benchmark Scale | |
| Underlying questions | 124 |
| Evaluation instances | 496 |
| ID | Task | Evaluation focus |
| A. Document Perception & Structure | ||
| A1 | Layout Parsing & Reading Order Recovery | Recover local reading flow in two-column PDFs with floating figures, footnotes, and cross-page continuation. |
| A2 | Element Association & Evidence Localization | Given a claim, locate its supporting figure, table, equation, appendix, page, or block. |
| A3 | Citation Role Analysis | Identify whether cited work acts as background, method basis, data resource, baseline, result support, critique, or extension. |
| A4 | Dataset Lineage Analysis | Recover the source, composition, filtering, derivation, and transformation chain of datasets. |
| A5 | Figure Analysis | Interpret scientific figures, subfigures, legends, axes, annotations, and visual evidence supporting figure-level claims. |
| Rule | Questions | Credit and failure conditions |
| Reading-order ranking | 5 | Fraction of correctly placed labels; incorrect sequence length scores zero. |
| Binary decisions | 6 | Fraction of correct bits in a final bitstring of the required length. |
| Closed scalar | 1 | Normalized label match, with the task’s answer-field/label extraction. |
| Grouped choices | 1 | Four groups weighted 0.30, 0.10, 0.30, 0.30; disallowed choices invalidate the affected group. |
| Unordered exact records | 1 | Fraction of reference dictionaries recovered exactly, ignoring key order; one-to-one matching prevents duplicate credit. |
| Indexed exact records | 1 | Exact row count and keys are prerequisites; schema credit plus per-row exact-match credit. |
| Subtask | ID | Subject | Source Document(s) | Eval. |
| A1 | 1 | Computer Science | MM-IFEngine: Towards Multimodal Instruction Following [ 24 ] | J |
| 7 | Nuclear Theory | Revisiting p- 11 B Fusion: Updated Cross-sections, Reactivity, and Energy Balance [ 114 ] | R | |
| 13 | Computer Science | Depth Anything 3: Recovering the Visual Space from Any Views [ 60 ] | R | |
| 51 | Computer Science | MM-IFEngine: Towards Multimodal Instruction Following [ 24 ] | R | |
| 52 | Computer Science | Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning [ 65 ] | R | |
| 53 | Medicine | QuarkMedSearch: A Long-Horizon Deep Search Agent for Exploring Medical Intelligence [ 61 ] | R |
| Subtask | ID | Subject | Source Document(s) | Eval. |
| A3 | 3 | Computer Science | Proximal Policy Optimization Algorithms [ 101 ] | R |
| 9 | Biology | ProtFlow: Flow Matching-based Protein Sequence Design with Comprehensive Protein Semantic Distribution Learning and High-quality Generation [ 50 ] | R | |
| 15 | Computer Science | Spa3R: Predictive Spatial Field Modeling for 3D Visual Reasoning [ 44 ] | R | |
| 77 | Physics | Quantum computing: A taxonomy, systematic review and future directions [ 31 ] | R | |
| 101 | Physics | Quantum computing: A taxonomy, systematic review and future directions [ 31 ] | J | |
| 103 | Computer Science | Quantum computing: A taxonomy, systematic review and future directions [ 31 ] | R |
| Subtask | ID | Subject | Source Document(s) | Eval. |
| A5 | 58 | Biology | Estimation of the biological affinities of seven species of Sulawesi macaques based on multivariate analysis of dermatoglyphic pattern types [ 109 ] | J |
| 78 | Chemistry | A Delocalized Mixed-Valence Dinuclear Ytterbium Complex That Displays Intervalence Charge Transfer [ 81 ] | J | |
| 79 | Computer Science | Biologically Plausible Learning via Bidirectional Spike-Based Distillation [ 70 ] | R | |
| 80 | Physics | Design of a variable stiffness quasi-direct drive cable-actuated tensegrity robot [ 78 ] | R | |
| 88 | Biology | Deconstruction of a spino-brain–spinal cord circuit that drives chronic pain [ 115 ] | J | |
| 89 | Physics | New limits on the Pauli forbidden transitions in 12C nuclei obtained with the complete Borexino dataset [ 8 ] | R |
| Subtask | ID | Subject | Source Document(s) | Eval. |
| B2 | 5 | Computer Science | Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning [ 65 ] | R |
| 19 | Medicine | Glycosylation of anti-dsDNA IgG correlates with organ involvement in treatment-naive patients with systemic lupus erythematosus [ 139 ] | R | |
| 20 | Biology | LLM-assisted systematic review of large language models in clinical medicine [ 18 ] | J | |
| 21 | Biology | In vivo site-specific engineering to reprogram T cells [ 80 ] | J | |
| 59 | Maths | Quantum linear system solver based on time-optimal adiabatic quantum computing and quantum approximate optimization algorithm [ 2 ] | J | |
| 69 | Chemistry | Axial chirality-induced rigidification in aminoboranes enhances persistent room-temperature phosphorescence and circularly polarized luminescence [ 28 ] | J |
| Subtask | ID | Subject | Source Document(s) | Eval. |
| B3 | 6 | Computer Science | Proximal Policy Optimization Algorithms [ 101 ] | J |
| 22 | Physics | Observation of non-Abelian topological acoustic semimetals and their phase transitions [ 43 ] | J | |
| 23 | Biology | Poet: A generative model of protein families as sequences-of-sequences [ 112 ] | J | |
| 26 | Maths | Weighted theory for the Euclidean Dirac operator in higher dimensions [ 96 ] | J | |
| 70 | Maths | Weighted theory for the Euclidean Dirac operator in higher dimensions [ 96 ] | J | |
| C1 | 24 | Computer Science | Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning [ 65 ] | J |
| Subtask | ID | Subject | Source Document(s) | Eval. |
| C3 | 30 | Medicine | Glycosylation of anti-dsDNA IgG correlates with organ involvement in treatment-naive patients with systemic lupus erythematosus [ 139 ] | R |
| 31 | Environment | Simultaneous removal of heavy metals from aqueous solutions by pineapple crown and avocado peel hydrogel composites [ 119 ] | J | |
| 32 | Maths | The Euler Stratification for [ 38 ] | J | |
| 33 | Maths | Computation and sampling for Schubert specializations [ 3 ] | R | |
| 34 | Maths | Learning compositional functions with transformers from easy-to-hard data [ 118 ] | J | |
| 66 | Biology | Neural sequences underlying directed turning in Caenorhabditis elegans [ 52 ] | R |
| Subtask | ID | Subject | Source Document(s) | Eval. |
| D2 | 36 | Computer Science | Eagle 2.5: Boosting long-context post-training for frontier vision-language models [ 17 ] , Mammoth-vl: Eliciting multimodal reasoning with instruction tuning at scale [ 35 ] | J |
| 37 | Computer Science | Cambrian-1: A fully open, vision-centric exploration of multimodal llms [ 110 ] , Llava-onevision: Easy visual task transfer [ 56 ] , Mammoth-vl: Eliciting multimodal reasoning with instruction tuning at scale [ 35 ] | R | |
| 95 | Chemistry | Strategies for pre-training graph neural networks [ 40 ] , Geometry-enhanced molecular representation learning for property prediction [ 29 ] , Molecular contrastive learning of representations via graph neural networks [ 117 ] | R | |
| 96 | Biology | ProtFlow: Flow Matching-based Protein Sequence Design with Comprehensive Protein Semantic Distribution Learning and High-quality Generation [ 50 ] , PTM-Mamba: a PTM-aware protein language model with bidirectional gated Mamba blocks [ 89 ] , Compressing the collective knowledge of ESM into a single protein language model [ 25 ] | R | |
| 109 | Chemistry | Molecular Contrastive Learning of Representations via Graph Neural Networks [ 117 ] , ChemRL-GEM: Geometry Enhanced Molecular Representation Learning for Property Prediction [ 29 ] , Strategies for Pre-Training Graph Neural Networks [ 40 ] | J | |
| 120 | Computer Science | Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs [ 110 ] , LLaVA-OneVision: Easy Visual Task Transfer [ 56 ] , MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale [ 35 ] | R |
| Subtask | ID | Subject | Source Document(s) | Eval. |
| D3 | 38 | Computer Science | SpatialDreamer: Incentivizing Spatial Reasoning via Active Mental Imagery [ 11 ] , Spatialladder: Progressive training for spatial reasoning in vision-language models [ 58 ] | J |
| 39 | Computer Science | Ts-llava: Constructing visual tokens through thumbnail-and- sampling for training-free video large language models [ 90 ] , Pllava: Parameter-free llava extension from images to videos for video dense captioning [ 130 ] | J | |
| 47 | Medicine | Glycosylation of anti-dsDNA IgG correlates with organ involvement in treatment-naive patients with systemic lupus erythematosus [ 139 ] , IgG glycans in health and disease: Prediction, intervention, prognosis, and therapy [ 104 ] | R | |
| 97 | Computer Science | Caprl: Stimulating dense image caption capabilities via reinforcement learning [ 126 ] , Scalecap: Inference-time scalable image captioning via dual-modality debiasing [ 127 ] | R | |
| 98 | Chemistry | High voltage cycling stability of LiF-coated NMC811 electrode [ 67 ] , Long-term cyclability of NCM-811 at high voltages in lithium-ion batteries: an in-depth diagnostic study [ 59 ] , Enhanced Cycling Stability of NCM811 Cathodes at High C-Rates and Voltages via LiMTFSI-Based Polymer Coating [ 48 ] | R | |
| 110 | Computer Science | ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing [ 127 ] , CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning [ 126 ] | R |
| Subtask | ID | Subject | Source Document(s) | Eval. |
| E1 | 40 | Computer Science | Group Sequence Policy Optimization [ 136 ] | R |
| 41 | Computer Science | LLaVA-CoT: Let Vision Language Models Reason Step-by-Step [ 129 ] | J | |
| 48 | Physics | Stairway Codes: Floquetifying Bivariate Bicycle Codes and Beyond [ 42 ] | J | |
| 49 | Physics | Imaginary-time evolution of interacting spin systems in the truncated Wigner approximation [ 100 ] | R | |
| 68 | Biology | Neural sequences underlying directed turning in Caenorhabditis elegans [ 52 ] | J | |
| 99 | Biology | Causal modelling of gene effects from regulators to programs to traits [ 88 ] | R |
| Rater pair | Three-level exact | QWK | Binary exact | Binary | MAE |
| Human expert vs. GPT-5.4-mini | 92.0% | 0.866 | 70.0% | 0.355 | 0.169 |
| Codex vs. GPT-5.4-mini | 89.0% | 0.818 | 80.0% | 0.593 | 0.128 |
| Human expert vs. Codex | 93.0% | 0.872 | 74.0% | 0.405 | 0.145 |
| Direction | Training subtasks |
| Layout and reading-flow recovery | Block role classification; parent–child linking; reading-order prediction; next-hop reading-target prediction. |
| Table logic consistency | Table logic consistency check. |
| Cross-document table integration | Real cross-document table merge; single-document simulated table merge. |
| Chart readout and recovery | Single-point chart readout; multi-point or series readout. |
| Chart visual consistency | Chart visual consistency. |
| Cross-document dataset comparison | Cross-document dataset intersection. |
| Subtask | SciDocIR or document records | External or synthetic source | GPT-5.4 role and GT authority |
| Block role classification | Page image; target page index, bounding box, type, and layout annotation | None | Light question rewriting. The layout label remains the locked GT. |
| Parent–child linking | Block text, page and bounding box; heading level; caption pairs; footnote and continuation relations | None | Question rewriting and optional plausibility check. The verified relation remains GT. |
| Reading-order prediction | Order and reading annotations; block pages and boxes; section and discourse roles | None | Question rewriting only. GT follows the stored order and layout rules. |
| Next-hop reading target | Current and successor blocks; caption pairs; figure/table references; cross-page and appendix links | None | Question rewriting and optional next-hop plausibility check. The stored successor remains GT. |
| Table logic consistency | Table LaTeX, normalized JSON, caption, headers, cells, citing context, page, and box | None | Generates a natural corrupted statement, repair, and distractors, then self-checks them. Rules lock the perturbed field and answer. |
| Real cross-document table merge | Table schemas, rows, columns, values, captions, and nearby text from different papers | None | Polishes context and question. Rules align compatible schemas and compute the merged GT. |
| Subtask | SciDocIR or document records | External or synthetic source | GPT-5.4 role and GT authority |
| Single-point chart readout | Figure and caption blocks; page, box, and nearby text used for page replacement | PlotQA chart, axes, series labels, and source values | Naturalizes the question and performs repeated visual solvability checks. PlotQA values remain GT. |
| Multi-point or series readout | Same figure placement and context records as the single-point task | Complete PlotQA series data | Rewrites the question and checks the returned list. The source series remains GT. |
| Chart visual consistency | Figure position, caption, and page context used to reinsert a controlled chart | PlotQA values; Matplotlib radar, line, scatter, and bar renderers | Rewrites the question and checks solvability. Rules move one plotted value and retain the original annotation; the corruption log defines GT. |
| Dataset intersection | Dataset-like entities mined from text, tables, and captions when available; simulated two-paper text | Internal lexicon of approximately 50 CS, AI, mathematics, and natural-science resources; alias and composition pools | Checks name plausibility and polishes page prose. Exact set intersection defines GT. |
| Notation extraction | Final batch does not require real SciDocIR records | Mathematics, machine-learning, and physics formula templates; XeLaTeX page rendering | May polish surrounding prose while formulas and definitions stay locked. The saved symbol table defines GT. |
| Symbol ambiguity disambiguation | Final batch does not require real SciDocIR records | Formula and context templates; one-column, two-column, and spanning layouts | May naturalize local contexts without changing definitions. The rule log for each symbol occurrence defines GT. |
| Setting | SFT | GRPO |
| Model adaptation and optimization | ||
| Initialization | Original Qwen3.6-27B | Merged SFT step 50 |
| Adaptation | LoRA | LoRA |
| LoRA rank / alpha | 32 / 64 | 16 / 32 |
| Learning rate | ||
| Learning-rate schedule | Cosine | Cosine |
| Task | Reference and predicted output | Failure interpretation |
| ID 13: ordered metric extraction (A1); score 0 | Reference: [Chamfer Distance, Precision, Recall, F1-score]. Prediction: [accuracy, completeness, Chamfer Distance (CD), precision, recall, F1-score]. | The model retrieves the target metrics but adds two earlier entries. Since the question restricts the section and order, every aligned position is wrong. This is a scope/ordering failure, not failure to recognize all four metrics. |
| ID 59: logarithmic chart readout (B2); score 0.30 | Reference: AQC(exp) range G ( to ), QAOA range C. Prediction: AQC(exp) range A ( to ), QAOA range C. | The wrong AQC lower endpoint is also asserted in the explanation. Correct QAOA selection earns partial credit; the rubric does not credit an explanation grounded in the incorrect joint range selection. |
| ID 71: table consistency audit (C1); score 0.67 | Reference errors: 48.4, 47.8, 32.4. Prediction: 68.0, 32.4, 48.4. The missing 47.8 should be 47.3, the rounded mean of eight entries. | Two errors are found, including a correctly recomputed mean, but another mean is not checked correctly. The additional highlighting complaint about 68.0 does not recover the missing reference error. |