Scientific agents support a range of literature-based research tasks, such as retrieval, question answering, evidence-grounded generation, and claim assessment. Most existing systems, however, are organized around individual tasks: the same papers are repeatedly retrieved, segmented, and interpreted, and the understanding built in one task is difficult to reuse in the next. We present ScholarStack, a layered research asset framework that compiles a paper collection into reusable, versioned, and provenance-preserving assets at three complementary levels: source-grounded paper-level statements, domain-level organization, and evidence-grounded cross-paper syntheses. A common access interface returns task-specific views at the evidence granularity each task requires, preserving study conditions, source traceability, and verification status. We instantiate the framework on four task families spanning ten task settings, comparing agents that use the compiled assets with task-specific baselines under matched base models. Quality gains concentrate on tasks that require cross-paper evidence, such as multi-paper question answering and literature review generation, and query-time token cost falls on every task where it is measured, with assets compiled once and reused across tasks. These results suggest that layered research assets can serve as shared infrastructure for scientific agents, shifting literature-based assistance from isolated document processing toward cumulative, evidence-grounded workflows.
Figures & tables
Research need
Reusable assets
Reuse benefits
Grounded reading and verification
L1: Source-grounded facts with study conditions, source locations, and checking states.
Reduce repeated source processing while preserving the context of reported findings.
Organization and comparison
L1 + L2: Canonical entities, categories, and relations linked to result-specific assumptions, measurements, and settings.
Reuse organizational judgments while avoiding comparisons across incompatible settings.
Cross-paper synthesis
L3 (grounded in L1 + L2): Scoped syntheses built over organized research objects and linked to source-grounded evidence.
Reuse scoped interpretations without extending conclusions beyond the examined evidence.
Inspection and revision
Across layers: Evidence dependencies, identifiers, versions, and checking states.
Preserve justification and identify what requires rechecking when evidence changes.
Table 1: Research requirements and corresponding reusable assets.
Figure 1: Overview of ScholarStack. Knowledge construction compiles documents into source-grounded facts (L1), domain organization (L2), and scoped cross-paper syntheses (L3), linked by shared identifiers and evidence references. During task execution, agents select and combine the required objects into a Knowledge View, retaining their conditions, evidence references, and recorded verification states. Different tasks use different subsets of the layers and operations. The feedback path denotes optional registration of task-derived candidates after the applicable checks; registration does not itself establish semantic verification.
Operation
Output and semantics
Retrieve(K,q)→V
Locate recorded facts, domain objects, or existing syntheses relevant to the request, retaining their sources and status.
Select(V,c)→V′
Filter the view by explicit conditions, distinguishing confirmed matches from mismatches and missing or inapplicable information.
Expand(K,V,r)→V′
Supplement the view with objects connected through the specified relations or evidence links, retaining the connections to the selected objects.
Compare(K,V,d)→C
Organize findings by the requested comparison dimensions, exposing shared and differing conditions and identifying unresolved comparability.
Synthesize(K,V,q)→s
Generate a scoped interpretation from the selected evidence, recording its supporting, opposing, and limiting basis as a candidate synthesis.
Table 2: Core operations over the knowledge layer. K=(F,O,S) comprises source-grounded facts, domain organization, and scoped syntheses; q is a retrieval or synthesis request; V is a selected view of typed knowledge objects, and V′ is the view returned by selection or expansion; c specifies selection constraints, r the relations or evidence links to follow, and d comparison dimensions; C is a Comparison Result, and s a candidate Scoped Synthesis. The signatures specify semantic inputs and outputs rather than a fixed execution policy.
Figure 2: Knowledge-guided agentic literature retrieval. A compiled Knowledge View informs search planning and evidence selection. Candidate papers are checked against the request using retrieved source evidence, which also supports the final ranking. Fixed-corpus and open-world settings differ in the permitted search scope.
Method
Corpus
Asset Size
R@1
R@5
R@10
MRR@40
Tokens
BM25
64,183
–
0.220
0.470
0.537
0.331
–
Dense
64,183
–
0.257
0.503
0.533
0.379
–
Hybrid
64,183
–
0.347
0.537
0.583
0.452
–
Full-text
64,183
900
0.650
0.767
0.787
0.707
49.3k
ScholarStack
64,183
900
0.670
0.773
0.773
0.717
47.9k
Table 3: Fixed-corpus literature search results on LitSearch ( Ajith et al., 2024 ) . Tokens are per-query means. Bold marks the best retrieval value among methods evaluated on LitSearch. Corpus and asset size refer to the number of papers.
In-source
Out-of-source
Method
R@1
R@5
R@10
MRR@40
R@1
R@5
R@10
MRR@40
Full-text
0.711
0.807
0.807
0.781
0.613
0.742
0.774
0.662
ScholarStack
0.763
0.825
0.825
0.818
0.613
0.742
0.742
0.654
Table 4: Fixed-corpus literature search results by knowledge-source coverage: 19 queries with at least one gold paper among the 900 asset sources (in-source), and 31 with none (out-of-source). All gold papers are evaluated against the same retrieval corpus. Bold marks the best value within each model and coverage group, including ties.
Method
Recall
Precision
F1-Score
Macro W-Recall
Micro W-Recall
Full-text
0.124
0.138
0.128
0.117
0.113
ScholarStack
0.134
0.158
0.141
0.124
0.126
Table 5: Pooled open-world retrieval results on 89 SAGE queries. All metrics use the explicit Tier 1 set, which contains approximately five papers. Recall, Precision, F1, and Macro Weighted Recall are macro-averaged over queries; Micro Weighted Recall pools weighted labels.
Method
Weighted F1
Recall
Precision
Tokens
Full-text
0.645
0.761
0.560
35.6k
ScholarStack
0.662
0.779
0.576
21.4k
Table 6: Novelty-assessment results for n=40 overlap-present targets across 12 fields and 135 gold overlaps. Weighted identification F1 uses weights of 1.0 for core axes and 0.3 for peripheral axes; token counts are per-query means.
Answer-F1
Method
Overall
Extr.
Abst.
Bool.
Unans.
Evidence-F1
Tokens
Full-text
0.6146
0.6596
0.2697
0.6950
0.9198
0.6015
5.4k
Human
0.6090
–
–
–
–
0.7160
–
RAG
0.5613
0.5792
0.2527
0.6571
0.9091
0.5610
2.1k
ScholarStack
0.5708
0.5988
0.2631
0.6495
0.9114
0.5887
1.9k
Table 7: Single-paper QA on QASPER (416 papers / 1,428 questions). Overall denotes Answer-F1 across all questions; Extr., Abst., Bool., and Unans. denote extractive, abstractive, boolean, and unanswerable questions, respectively. Answer-F1 measures token-level answer overlap; Evidence-F1 measures paragraph-level evidence overlap. Human is the reported human-answer reference from QASPER ( Dasigi et al., 2021 ) . Tokens denotes mean input tokens per question; bold marks improvements over RAG.
Answerability
Generation (Rouge-L)
Method
Ans-F1
Unans-F1
Acc.
W-F1
Macro-F1
FF
AE
Tokens
Full-text
0.8056
0.5071
0.7212
0.7381
0.6564
0.1548
0.2048
12.6k
Human
0.7318
–
–
–
–
0.1501
0.1605
0.3k
RAG
0.7356
0.4759
0.6485
0.6768
0.6057
0.1416
0.1564
1.9k
ScholarStack
0.7485
0.4783
0.6606
0.6874
0.6134
0.1447
0.1558
1.4k
Table 8: Single-paper QA on PeerQA (208 papers / 579 questions). Answerability is evaluated on 495 questions: Ans-F1 and Unans-F1 denote F1 for answerable and unanswerable questions; Acc., W-F1, and Macro-F1 denote accuracy, class-frequency-weighted F1, and unweighted mean class F1. Generation ROUGE-L is evaluated on 245 questions against free-form reference answers (FF) and annotated evidence (AE). Human denotes model answers using human-annotated evidence, with answerability scored only on answerable questions ( Baumgärtner et al., 2025 ) . Tokens denotes mean input tokens per question; bold marks improvements over RAG.
Method
Corr.
Comp.
Synth.
Cal.
Dir.
Total
Tokens
Sec.
Full-text
1.11
2.01
1.41
0.92
3.01
8.46
20.8k
13.13
ScholarStack
3.89
3.35
3.48
3.90
3.80
18.41
5.5k
7.54
Table 9: Multi-paper QA on 797 MDAQA questions. Corr., Comp., Synth., Cal., and Dir. denote correctness, completeness, cross-paper synthesis, calibration, and directness. Scores are from 0 to 4, and Total sums the five scores (0 to 20). Tokens and Sec. are mean tokens and generation seconds per question over available usage records, excluding knowledge construction and judging.
Method
Rel.
Cov.
Spec.
Supp.
Overall
Full-text
4.688
4.219
4.359
4.234
4.156
ScholarStack
4.578
4.250
4.453
4.438
4.203
Table 10: Quality-dimension performance on the OARelatedWork benchmark (each dimension scored out of 5). Rel., Cov., Spec., and Supp. denote Relevance, Coverage, Specificity, and Support.
Method
True ↑
Wrong Cit. ↓
Unverif. ↓
False ↓
Citation F1 ↑
Full-text
91.82
2.14
4.23
1.81
0.95
ScholarStack
92.43
1.10
5.31
1.15
0.97
Table 11: Factuality and citation performance on the OARelatedWork benchmark. Claim-label values are sample-mean percentages; citation F1 uses a 0–1 scale.
System
F1
Prec.
Rec.
Hits
GPT-Researcher
0.039
0.290
0.022
72
STORM
0.015
0.250
0.008
21
LangChain ODR
0.007
0.055
0.004
13
Lacuna Deep Research
0.052
0.339
0.028
99
Claude Code DR
0.090
0.553
0.049
171
ScholarStack
0.214
0.260
0.182
637
Table 12: ReportBench-ML citation overlap against expert-survey references.
System
Overall
Comp.
Insight
Instr.
Read.
GPT-Researcher
5.24
5.31
4.87
4.41
7.62
STORM
2.90
3.45
2.67
1.66
4.20
LangChain ODR
7.42
7.48
7.23
7.22
8.16
Lacuna Deep Research
7.82
8.01
7.61
7.57
8.34
Claude Code DR
7.54
7.04
7.43
8.28
7.88
ScholarStack
9.18
9.23
9.30
9.16
8.81
Table 13: RACE report-quality scores on ReportBench-ML. Scores are out of 10.
Method
Coverage F1
Recall
Precision
Tokens
Base LLM
0.339
0.665
0.255
18.7k
Full-text
0.398
0.665
0.329
20.7k
ScholarStack
0.404
0.642
0.331
12.4k
Table 14: Experimental-design generation results. Coverage F1 against the paper’s actual experiments (weighted by experiment kind). Tokens indicate per query.
Table 16: Comparison of idea-generation quality across five dimensions on the benchmark.
Method
Correct
Accuracy (%)
Tokens/query
Full-text
37/80
46.25
12.0k
ScholarStack
41/80
51.25
2.8k
Table 17: Scientific claim verification on 80 NLPCC instances ( 20 per label). Tokens/query includes input, output, and retries, excluding knowledge construction. Bold marks the better observed value in each metric.
Table 18: Object fields used in the example. These are presentation fields for the representation in Section 3.2 , not a complete storage schema. All objects have an identity; verification and lifecycle states are annotations. The final column gives illustrative types and values, not exhaustive enumerations.
saga.convex-rate : component assumptions at SAGA:9
Guarantee type
Sublinear rate
saga.convex-rate : convex result at SAGA:29
Appendix
Table 19: Distinct organizational views of the same method. The assumption and guarantee rows share the convex result as their displayed basis. These labels describe that result, not every guarantee available for SAGA. The family assignment records three agreeing votes; all three displayed assignments have stored confirmed certainty, current validity, and passed evidence verification.
Role
Source fact
Contribution and source locator
Supporting
Prox-SVRG method [ Xiao and Zhang, 2014 ]
Periodic full-gradient computation; Prox-SVRG:83 states when the full gradient is computed.
Supporting
saga.method [ Defazio et al., 2014 ]
Historical gradient storage; SAGA:20 . Same fact as the L2 family-assignment basis.
Supporting
Saddle-point method [ Palaniappan and Bach, 2016 ]
Extension to saddle-point problems: Saddle:19 ; convex-concave setting: Saddle:14 .
Limiting
saga.convex-rate [ Defazio et al., 2014 ]
Specifies the guarantee in the convex setting ( SAGA:9,29 ); this upper bound does not establish that faster convergence is impossible.
Limiting
Saddle-point comparison [ Palaniappan and Bach, 2016 ]
Uniform-sampling SAGA/SVRG may be inferior to an accelerated batch method; Saddle:118 .
Appendix
Table 20: Evidence roles in the displayed synthesis. A supporting entry contributes to the account; a limiting entry qualifies its application. Locators identify sentences in each paper’s saved source map. The convex-rate entry uses the corrected basis at SAGA:29 .
Stage
Corpus searches
Title lookups
Asset lookups
Abstract inspections
Full-text reads
0: Initial planning
3
2
1
0
0
1: Expansion and checks
3
2
1
6
4
2: Final verification
2
2
0
4
2
3: Ranking
0
0
0
0
0
Appendix
Table 21: Maximum numbers of tool calls per stage. Asset lookups are available to Full-text and ScholarStack but not to the no-asset control. All allocations are subject to the cumulative limits in Table 22 .
Resource
Limit
Model calls / candidate-context size
4 / 180 papers
Corpus searches / title lookups
8 (20 papers per search) / 6
Metadata resolutions / abstract inspections
10 / 10
Asset lookups and returned content
2 lookups; at most 3 records per lookup; 8,000 characters in total
Full-text reads
6 calls covering at most 6 distinct papers
Full-text response length
4,000 characters per call; 6,000 per paper; 24,000 in total
Appendix
Table 22: Execution and response limits per query. Byte limits are measured in UTF-8; character and token limits are enforced separately.
Method
Corpus
Asset Size
R1
R5
R10
MRR40
Tokens
Full-text
64,183
900
0.670
0.773
0.773
0.721
41.1k
ScholarStack
64,183
900
0.710
0.803
0.803
0.763
39.5k
Appendix
Table 23: Fixed-corpus retrieval results with GPT-5.6 Terra. Source papers denotes the number of papers used to construct the auxiliary assets. Runtime and token usage are averaged per query; GPT token counts are reported-usage lower bounds. Bold indicates the best value in each retrieval-metric column.
In-source
Out-of-source
Method
R1
R5
R10
MRR40
R1
R5
R10
MRR40
Codex raw (w/o full-text)
0.605
0.719
0.719
0.684
0.677
0.774
0.774
0.720
Full-text
0.711
0.772
0.772
0.763
0.645
0.774
0.774
0.695
ScholarStack
0.711
0.798
0.798
0.781
0.710
0.807
0.807
0.753
Appendix
Table 24: Fixed-corpus retrieval results by knowledge-source coverage with GPT-5.6 Terra. Bold indicates the best value in each column, including ties.
Domain
SAGE queries
Source papers
Computer science
29
1,754
Natural science
28
1,749
Healthcare
32
1,748
Appendix
Table 25: Frozen source collections for open-world retrieval. "Source papers" denotes the number of full-text documents underlying both the Full-text and ScholarStack conditions.
Field
Count
Field
Count
Field
Count
Linguistics
5
Physics
4
Education
2
Medicine
5
Psychology
4
Art
1
Engineering
5
Business
4
Chemistry
1
Mathematics
4
Biology
4
Materials Sci.
1
Appendix
Table 26: Per-field breakdown of the 40 overlap-present novelty targets (12 fields).
Answerability
Generation (Rouge-L)
Method
Ans-F1
Unans-F1
Acc.
W-F1
Macro-F1
FF
AE
Tokens
Full-text
0.806
0.507
0.721
0.738
0.656
0.155
0.205
12,577
Human (reference)
0.732
–
–
–
–
0.150
0.161
301
RAG10
0.736
0.476
0.649
0.677
0.606
0.142
0.156
1,864
RAG20
0.766
0.489
0.679
0.703
0.627
0.149
0.184
3,481
ScholarStack
0.749
0.478
0.661
0.687
0.613
0.145
0.156
1,350
Appendix
Table 27: All evaluated conditions on PeerQA (208 papers / 579 questions). RAG@ k retrieves the top- k passages. Answerability is evaluated on 495 questions: Ans-F1 and Unans-F1 denote F1 for answerable and unanswerable questions; Acc., W-F1, and Macro-F1 denote accuracy, class-frequency-weighted F1, and unweighted mean class F1. Generation ROUGE-L is evaluated on 245 questions against free-form reference answers (FF) and annotated evidence (AE). Human denotes model answers using human-annotated evidence, with answerability scored only on answerable questions. Tokens denotes mean input tokens per question; bold marks ScholarStack’s improvements over RAG10.
Table 28: Per-domain results on PeerQA. Parentheses give the numbers of answerability / generation questions. Ans-F1 and Unans-F1 denote F1 for answerable and unanswerable questions; Acc., W-F1, and Macro-F1 denote accuracy, class-frequency-weighted F1, and unweighted mean class F1. RgL ∗FF and RgL ∗AE denote generation ROUGE-L against free-form reference answers and annotated evidence, respectively. Bold marks the best Macro-F1 in each domain.
Question
Gold
Full-text
RAG
ScholarStack
QASPER
What inference models are used?
XGBoost, AdaBoost, Random Forest
XGBoost, AdaBoost and Random Forest
Gaussian Processes
XGBoost, AdaBoost, Random Forest
How many actions are present in the dataset?
14,769
14,769
No Answer
14,769
PeerQA
Does the optimization happen only over map locations observed so far?
No; it also covers unobserved (hallucinated) regions
No; any region observed or hallucinated
No Answer
No; goals from any observed or hallucinated region
What values are represented in Table 2?
Zero-shot classification performance across 5 datasets and 14 language models
Zero-shot classification performance for all models and datasets
No Answer
Performance statistics for all models and datasets
Appendix
Table 29: Selected qualitative examples on QASPER and PeerQA. Gold denotes the reference answer. Reference answers and model responses are abridged for presentation; answer completeness varies. Aggregate quality and input usage are reported in Tables 7 and 8 .
Method
Corr.
Comp.
Synth.
Cal.
Dir.
Total
Tokens
Sec.
Full-text + Bare
1.11
2.01
1.41
0.92
3.01
8.46
20.8k
13.13
ScholarStack (L1) + Bare
3.90
3.56
2.99
3.78
3.82
18.04
4.1k
9.83
ScholarStack (L1) + Skill
3.91
3.60
3.03
3.82
3.83
18.18
4.2k
10.45
ScholarStack (L1+L2) + Bare
3.82
3.25
2.80
3.26
3.88
17.01
4.8k
7.20
ScholarStack (L1+L2) + Skill
3.89
3.53
3.02
3.78
3.84
18.05
4.9k
10.67
ScholarStack (L1+L2+L3) + Bare
3.89
3.35
3.48
3.90
3.80
18.41
5.5k
7.54
Appendix
Table 30: Full ablation and generation-cost results on the 797-question MDAQA bank. Bare denotes the answering model using the ordinary prompt and supplied evidence, without the additional skill instructions. Skill denotes Cross-Paper Evidence Synthesis skill, which guides comparison of study conditions, synthesis of findings, and attribution of claims to evidence. Corr., Comp., Synth., Cal., and Dir. denote correctness, completeness, cross-paper synthesis, calibration, and directness. Scores are weighted means on a 0–4 scale, except Total (0–20). Generation-cost columns use the available per-question usage audit. Construction and judging are excluded. Bold marks the best value in each column, where higher is better for the six score columns and lower is better for the two cost columns.
System
Input(k)
Output(k)
Cache hit(%)
Cost(USD)
Claude Code DR
3,128
106
78.94
7.18
ScholarStack
10,533
117
96.02
3.49
Appendix
Table 31: Per-report resource use and standardized cost for the two locally evaluated review-generation workflows. k denotes one thousand tokens.
Field
Count
Field
Count
Field
Count
Medicine
5
Physics
4
Education
2
Engineering
5
Business
4
Art
1
Linguistics
4
Biology
4
Chemistry
1
Mathematics
4
Psychology
3
Materials Sci.
1
Appendix
Table 32: Per-field breakdown of the 38 experimental-design targets (12 fields).
Domain
Count
AI and engineering
14
Chemistry and materials
12
Life sciences and medicine
4
Public health and social sciences
3
Environmental science
2
Total
35
Appendix
Table 33: Domain composition of the idea-generation benchmark.
Dimension
Definition
Novelty
Are the problems or approaches new? Is this a novel combination of familiar techniques? Is it clear how this work differs from previous contributions? Is related work adequately referenced?
Significance
Is the idea important? Are other people (practitioners or researchers) likely to use these ideas or build on them? Does the idea address a difficult problem in a better way than previous research? Does it provide a unique theoretical or pragmatic approach?
Clarity
Is the paper clearly written? Is it well-organized? Does it adequately inform the reader?
Feasibility
Can the idea be realized with existing technology or methods? Are there any technical difficulties or bottlenecks? Is the idea clear and logical? Are there any obvious errors or unreasonable parts in the idea, and can the experiments be designed normally according to this idea.
Effectiveness
How likely the proposed idea is going to work well (e.g., better than existing baselines).
Appendix
Table 34: Definitions of evaluation dimensions from Li et al. [2024] .
Current LLM-based research agents have advanced through agent orchestration, yet largely overlook scientific knowledge orchestration. Existing works often reduce papers to abstracts, surface mentions, and flat \texttt{cites} edges, omitting key entities, claims, evidence, mechanisms, and method lineages essential for scientific reasoning. To this end, we introduce \textbf{Agents-K1}, an end-to-end knowledge orchestration pipeline that converts raw documents into agent-native scientific knowledge graphs. Agents-K1 integrates three components under a unifying theoretical foundation: a multimodal parser whose five-module schema captures entities, multimodal evidence, citations, and typed inter-entity relations across the full paper rather than abstracts alone; a 4B information-extraction backbone trained with GRPO under a rule-based reward; and a graphanything CLI, a tri-source agent interface that unifies web search, multimodal graph retrieval, and cross-document traversal. On top of this, we process 2.46 million scientific papers across six subjects to produce \textbf{Scholar-KG}, of which we release a one-million-paper subset, and the full Scholar-KG is accessible via the SCP link below. The same pipeline can be extended to general-domain corpora and to schema-conformant data synthesis. Extensive experiments demonstrate that Agents-K1 achieves superior performance in scientific information extraction, knowledge graph construction, and multi-hop scientific reasoning.
Zongsheng Cao, Bihao Zhan, Jinxin Shi +23
Shanghai Artificial Intelligence Laboratory · East China Normal University · Fudan University
Scientific publication compresses a branching, iterative research process into a linear narrative, discarding the majority of what was discovered along the way. This compilation imposes two structural costs: a Storytelling Tax, where failed experiments, rejected hypotheses, and the branching exploration process are discarded to fit a linear narrative; and an Engineering Tax, where the gap between reviewer-sufficient prose and agent-sufficient specification leaves critical implementation details unwritten. Tolerable for human readers, these costs become critical when AI agents must understand, reproduce, and extend published work. We introduce the Agent-Native Research Artifact (ARA), a protocol that replaces the narrative paper with a machine-executable research package structured around four layers: scientific logic, executable code with full specifications, an exploration graph that preserves the failures compilation discards, and evidence grounding every claim in raw outputs. Three mechanisms support the ecosystem: a Live Research Manager that captures decisions and dead ends during ordinary development; an ARA Compiler that translates legacy PDFs and repos into ARAs; and an ARA-native review system that automates objective checks so human reviewers can focus on significance, novelty, and taste. On PaperBench and RE-Bench, ARA raises question-answering accuracy from 72.4% to 93.7% and reproduction success from 57.4% to 64.4%. On RE-Bench's five open-ended extension tasks, preserved failure traces in ARA accelerate progress, but can also constrain a capable agent from stepping outside the prior-run box depending on the agent's capabilities. Our code is open-sourced at https://github.com/Orchestra-Research/Agent-Native-Research-Artifact.
Jiachen Liu, Jiaxin Pei, Jintao Huang +34
University of Michigan · Stanford University · Ohio State University +22
Autonomous scientific research is significantly advanced thanks to the development of AI agents. One key step in this process is finding the right scientific literature, whether to explore existing knowledge for a research problem, or to acquire evidence for verifying assumptions and supporting claims. To assess AI agents' capability in driving this process, we present AutoResearchBench, a dedicated benchmark for autonomous scientific literature discovery. AutoResearchBench consists of two complementary task types: (1) Deep Research, which requires tracking down a specific target paper through a progressive, multi-step probing process, and (2) Wide Research, which requires comprehensively collecting a set of papers satisfying given conditions. Compared to previous benchmarks on agentic web browsing, AutoResearchBench is distinguished along three dimensions: it is research-oriented, calling for in-depth comprehension of scientific concepts; literature-focused, demanding fine-grained utilization of detailed information; and open-ended, involving an unknown number of qualified papers and thus requiring deliberate reasoning and search throughout. These properties make AutoResearchBench uniquely suited for evaluating autonomous research capabilities, and extraordinarily challenging. Even the most powerful LLMs, despite having largely conquered general agentic web-browsing benchmarks such as BrowseComp, achieve only 9.39% accuracy on Deep Research and 9.31% IoU on Wide Research, while many other strong baselines fall below 5%. We publicly release the dataset and evaluation pipeline to facilitate future research in this direction. We publicly release the dataset, evaluation pipeline, and code at https://github.com/CherYou/AutoResearchBench.