cs.AISep 4, 2026

SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding

Authors: Shenxi Wu, Yuhong Liu, Haosong Zhang, Tongjin Zou, Yanxun Zhang, Gaochang Chen, Dun Liang, Jiaqi Wang, +3 more

Organizations: The Chinese University of Hong Kong · Shanghai Artificial Intelligence Laboratory · New York University · Fudan University · Shanghai Jiao Tong University · Harbin Institute of Technology · JD Explore Academy · Centre for Perceptual and Interactive Intelligence (CPII) Limited

Abstract

Scientific papers require models to integrate evidence across text, equations, figures, tables, code, and datasets while preserving its provenance. Beyond answer correctness, scientific reading requires verifiable outputs from operations such as evidence localization, definition extraction, and consistency checking. We introduce SciDocBench, a workflow-centered benchmark targeting these operations through 124 expert-authored and difficulty-screened questions across seven research-assistant capability groups, 19 subtasks, and five scientific domains. Each question is instantiated in four matched settings formed by pairing its bilingual variants with the All Images First and Markdown Interleaved document representations, yielding 496 evaluation instances. The strongest evaluated model, Claude-Opus-5, scores 62.6 out of 100, with remaining gaps in evidence localization, structured information extraction, cross-document synthesis, and robustness to document representation. To convert these diagnostics into scalable training signals, we introduce SciDocIR, a structured representation of scientific document objects, layout and cross-reference relations, and provenance. Using SciDocIR, we construct SciDocDataset, which contains 4K supervised fine-tuning instances and 10K reinforcement-learning instances, built on 14 verifiable training subtasks. Post-training Qwen3.6-27B on task-aligned data improves its SciDocBench score from 40.03 to 45.33 with supervised fine-tuning and to 45.74 with subsequent reinforcement learning. Both adapted models preserve DocVQA and InfoVQA performance and improve ChartQA accuracy over the original model by 0.80 and 3.40 points, respectively. Together, SciDocBench, SciDocIR, and SciDocDataset connect capability diagnosis with verifiable training-data construction for scientific-document assistants.

Figures & tables

Appendix figures & tables22 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. LitTraceQA: A Benchmark for Multi-Stage Grounding and Verification in Scientific Question Answering

    Aug 7, 2026Xuye Liu, Yimu Wang, Peng Shi +7Literature Review

  2. DocScope: Benchmarking Verifiable Reasoning for Trustworthy Long-Document Understanding

    May 9, 2026Xiang Feng, Jiawei Zhou, Zhangfeng Huang +6Multimodal DocumentsQuestion-Answering Benchmarks