ScholarStack: Layered Research Asset Orchestration and Cross-Task Reuse for Scientific Agents
Organizations: DAMO Academy, Alibaba Group
Abstract
Scientific agents support a range of literature-based research tasks, such as retrieval, question answering, evidence-grounded generation, and claim assessment. Most existing systems, however, are organized around individual tasks: the same papers are repeatedly retrieved, segmented, and interpreted, and the understanding built in one task is difficult to reuse in the next. We present ScholarStack, a layered research asset framework that compiles a paper collection into reusable, versioned, and provenance-preserving assets at three complementary levels: source-grounded paper-level statements, domain-level organization, and evidence-grounded cross-paper syntheses. A common access interface returns task-specific views at the evidence granularity each task requires, preserving study conditions, source traceability, and verification status. We instantiate the framework on four task families spanning ten task settings, comparing agents that use the compiled assets with task-specific baselines under matched base models. Quality gains concentrate on tasks that require cross-paper evidence, such as multi-paper question answering and literature review generation, and query-time token cost falls on every task where it is measured, with assets compiled once and reused across tasks. These results suggest that layered research assets can serve as shared infrastructure for scientific agents, shifting literature-based assistance from isolated document processing toward cumulative, evidence-grounded workflows.
Figures & tables
| Research need | Reusable assets | Reuse benefits |
|---|---|---|
| Grounded reading and verification | L1: Source-grounded facts with study conditions, source locations, and checking states. | Reduce repeated source processing while preserving the context of reported findings. |
| Organization and comparison | L1 + L2: Canonical entities, categories, and relations linked to result-specific assumptions, measurements, and settings. | Reuse organizational judgments while avoiding comparisons across incompatible settings. |
| Cross-paper synthesis | L3 (grounded in L1 + L2): Scoped syntheses built over organized research objects and linked to source-grounded evidence. | Reuse scoped interpretations without extending conclusions beyond the examined evidence. |
| Inspection and revision | Across layers: Evidence dependencies, identifiers, versions, and checking states. | Preserve justification and identify what requires rechecking when evidence changes. |
| Operation | Output and semantics |
|---|---|
| Locate recorded facts, domain objects, or existing syntheses relevant to the request, retaining their sources and status. | |
| Filter the view by explicit conditions, distinguishing confirmed matches from mismatches and missing or inapplicable information. | |
| Supplement the view with objects connected through the specified relations or evidence links, retaining the connections to the selected objects. | |
| Organize findings by the requested comparison dimensions, exposing shared and differing conditions and identifying unresolved comparability. | |
| Generate a scoped interpretation from the selected evidence, recording its supporting, opposing, and limiting basis as a candidate synthesis. |
| Method | Corpus | Asset Size | R@1 | R@5 | R@10 | MRR@40 | Tokens |
|---|---|---|---|---|---|---|---|
| BM25 | 64,183 | – | 0.220 | 0.470 | 0.537 | 0.331 | – |
| Dense | 64,183 | – | 0.257 | 0.503 | 0.533 | 0.379 | – |
| Hybrid | 64,183 | – | 0.347 | 0.537 | 0.583 | 0.452 | – |
| Full-text | 64,183 | 900 | 0.650 | 0.767 | 0.787 | 0.707 | 49.3k |
| ScholarStack | 64,183 | 900 | 0.670 | 0.773 | 0.773 | 0.717 | 47.9k |
| In-source | Out-of-source | |||||||
|---|---|---|---|---|---|---|---|---|
| Method | R@1 | R@5 | R@10 | MRR@40 | R@1 | R@5 | R@10 | MRR@40 |
| Full-text | 0.711 | 0.807 | 0.807 | 0.781 | 0.613 | 0.742 | 0.774 | 0.662 |
| ScholarStack | 0.763 | 0.825 | 0.825 | 0.818 | 0.613 | 0.742 | 0.742 | 0.654 |
| Method | Recall | Precision | F1-Score | Macro W-Recall | Micro W-Recall |
|---|---|---|---|---|---|
| Full-text | 0.124 | 0.138 | 0.128 | 0.117 | 0.113 |
| ScholarStack | 0.134 | 0.158 | 0.141 | 0.124 | 0.126 |
| Method | Weighted F1 | Recall | Precision | Tokens |
|---|---|---|---|---|
| Full-text | 0.645 | 0.761 | 0.560 | 35.6k |
| ScholarStack | 0.662 | 0.779 | 0.576 | 21.4k |
| Answer-F1 | |||||||
|---|---|---|---|---|---|---|---|
| Method | Overall | Extr. | Abst. | Bool. | Unans. | Evidence-F1 | Tokens |
| Full-text | 0.6146 | 0.6596 | 0.2697 | 0.6950 | 0.9198 | 0.6015 | 5.4k |
| Human | 0.6090 | – | – | – | – | 0.7160 | – |
| RAG | 0.5613 | 0.5792 | 0.2527 | 0.6571 | 0.9091 | 0.5610 | 2.1k |
| ScholarStack | 0.5708 | 0.5988 | 0.2631 | 0.6495 | 0.9114 | 0.5887 | 1.9k |
| Answerability | Generation (Rouge-L) | |||||||
|---|---|---|---|---|---|---|---|---|
| Method | Ans-F1 | Unans-F1 | Acc. | W-F1 | Macro-F1 | FF | AE | Tokens |
| Full-text | 0.8056 | 0.5071 | 0.7212 | 0.7381 | 0.6564 | 0.1548 | 0.2048 | 12.6k |
| Human | 0.7318 | – | – | – | – | 0.1501 | 0.1605 | 0.3k |
| RAG | 0.7356 | 0.4759 | 0.6485 | 0.6768 | 0.6057 | 0.1416 | 0.1564 | 1.9k |
| ScholarStack | 0.7485 | 0.4783 | 0.6606 | 0.6874 | 0.6134 | 0.1447 | 0.1558 | 1.4k |
| Method | Corr. | Comp. | Synth. | Cal. | Dir. | Total | Tokens | Sec. |
|---|---|---|---|---|---|---|---|---|
| Full-text | 1.11 | 2.01 | 1.41 | 0.92 | 3.01 | 8.46 | 20.8k | 13.13 |
| ScholarStack | 3.89 | 3.35 | 3.48 | 3.90 | 3.80 | 18.41 | 5.5k | 7.54 |
| Method | Rel. | Cov. | Spec. | Supp. | Overall |
|---|---|---|---|---|---|
| Full-text | 4.688 | 4.219 | 4.359 | 4.234 | 4.156 |
| ScholarStack | 4.578 | 4.250 | 4.453 | 4.438 | 4.203 |
| Method | True | Wrong Cit. | Unverif. | False | Citation F1 |
|---|---|---|---|---|---|
| Full-text | 91.82 | 2.14 | 4.23 | 1.81 | 0.95 |
| ScholarStack | 92.43 | 1.10 | 5.31 | 1.15 | 0.97 |
| System | F1 | Prec. | Rec. | Hits |
|---|---|---|---|---|
| GPT-Researcher | 0.039 | 0.290 | 0.022 | 72 |
| STORM | 0.015 | 0.250 | 0.008 | 21 |
| LangChain ODR | 0.007 | 0.055 | 0.004 | 13 |
| Lacuna Deep Research | 0.052 | 0.339 | 0.028 | 99 |
| Claude Code DR | 0.090 | 0.553 | 0.049 | 171 |
| ScholarStack | 0.214 | 0.260 | 0.182 | 637 |
| System | Overall | Comp. | Insight | Instr. | Read. |
|---|---|---|---|---|---|
| GPT-Researcher | 5.24 | 5.31 | 4.87 | 4.41 | 7.62 |
| STORM | 2.90 | 3.45 | 2.67 | 1.66 | 4.20 |
| LangChain ODR | 7.42 | 7.48 | 7.23 | 7.22 | 8.16 |
| Lacuna Deep Research | 7.82 | 8.01 | 7.61 | 7.57 | 8.34 |
| Claude Code DR | 7.54 | 7.04 | 7.43 | 8.28 | 7.88 |
| ScholarStack | 9.18 | 9.23 | 9.30 | 9.16 | 8.81 |
| Method | Coverage F1 | Recall | Precision | Tokens |
|---|---|---|---|---|
| Base LLM | 0.339 | 0.665 | 0.255 | 18.7k |
| Full-text | 0.398 | 0.665 | 0.329 | 20.7k |
| ScholarStack | 0.404 | 0.642 | 0.331 | 12.4k |
| Comparison | Wins | Losses | Ties |
|---|---|---|---|
| Full-text vs. Base LLM | 32 | 6 | 0 |
| ScholarStack vs. Base LLM | 34 | 4 | 0 |
| ScholarStack vs. Full-text | 16 | 22 | 0 |
| Method | Novelty | Significance | Clarity | Feasibility | Effectiveness | Mean |
|---|---|---|---|---|---|---|
| Full-text (w/o Skill) | 0.486 | 0.457 | 0.343 | 0.257 | 0.443 | 0.397 |
| Full-text (w/ Skill) | 0.500 | 0.514 | 0.571 | 0.543 | 0.500 | 0.526 |
| ScholarStack | 0.514 | 0.529 | 0.586 | 0.700 | 0.557 | 0.577 |
| Method | Correct | Accuracy (%) | Tokens/query |
|---|---|---|---|
| Full-text | 37/80 | 46.25 | 12.0k |
| ScholarStack | 41/80 | 51.25 | 2.8k |
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
| Object | Core fields | Example types or values |
|---|---|---|
| Scientific Fact | statement, context, evidence | Method description; theoretical result; limitation |
| Domain Taxonomy | dimensions, categories, membership criteria | Method family; assumption regime; guarantee type |
| Canonical Entity | name, type, aliases, source anchors | Method: SAGA; dataset: MNIST |
| Category Assignment | subject, dimension, category, basis, inference type, certainty | SAGA variance reduction; system judgment |
| Typed Relation | source, relation, target, evidence, payload | Paper evaluated on dataset; reported metric |
| Scoped Synthesis | account, scope, supporting / opposing / limiting basis | Method-family overview; selected studies |
| Dimension | Assigned category | Basis |
|---|---|---|
| Method family | Variance reduction | saga.method : historical gradient-table correction |
| Assumption regime | Convex | saga.convex-rate : component assumptions at SAGA:9 |
| Guarantee type | Sublinear rate | saga.convex-rate : convex result at SAGA:29 |
| Role | Source fact | Contribution and source locator |
|---|---|---|
| Supporting | Prox-SVRG method [ Xiao and Zhang, 2014 ] | Periodic full-gradient computation; Prox-SVRG:83 states when the full gradient is computed. |
| Supporting | saga.method [ Defazio et al., 2014 ] | Historical gradient storage; SAGA:20 . Same fact as the L2 family-assignment basis. |
| Supporting | Saddle-point method [ Palaniappan and Bach, 2016 ] | Extension to saddle-point problems: Saddle:19 ; convex-concave setting: Saddle:14 . |
| Limiting | saga.convex-rate [ Defazio et al., 2014 ] | Specifies the guarantee in the convex setting ( SAGA:9,29 ); this upper bound does not establish that faster convergence is impossible. |
| Limiting | Saddle-point comparison [ Palaniappan and Bach, 2016 ] | Uniform-sampling SAGA/SVRG may be inferior to an accelerated batch method; Saddle:118 . |
| Stage | Corpus searches | Title lookups | Asset lookups | Abstract inspections | Full-text reads |
|---|---|---|---|---|---|
| 0: Initial planning | 3 | 2 | 1 | 0 | 0 |
| 1: Expansion and checks | 3 | 2 | 1 | 6 | 4 |
| 2: Final verification | 2 | 2 | 0 | 4 | 2 |
| 3: Ranking | 0 | 0 | 0 | 0 | 0 |
| Resource | Limit |
|---|---|
| Model calls / candidate-context size | 4 / 180 papers |
| Corpus searches / title lookups | 8 (20 papers per search) / 6 |
| Metadata resolutions / abstract inspections | 10 / 10 |
| Asset lookups and returned content | 2 lookups; at most 3 records per lookup; 8,000 characters in total |
| Full-text reads | 6 calls covering at most 6 distinct papers |
| Full-text response length | 4,000 characters per call; 6,000 per paper; 24,000 in total |
| Method | Corpus | Asset Size | R1 | R5 | R10 | MRR40 | Tokens |
|---|---|---|---|---|---|---|---|
| Full-text | 64,183 | 900 | 0.670 | 0.773 | 0.773 | 0.721 | 41.1k |
| ScholarStack | 64,183 | 900 | 0.710 | 0.803 | 0.803 | 0.763 | 39.5k |
| In-source | Out-of-source | |||||||
|---|---|---|---|---|---|---|---|---|
| Method | R1 | R5 | R10 | MRR40 | R1 | R5 | R10 | MRR40 |
| Codex raw (w/o full-text) | 0.605 | 0.719 | 0.719 | 0.684 | 0.677 | 0.774 | 0.774 | 0.720 |
| Full-text | 0.711 | 0.772 | 0.772 | 0.763 | 0.645 | 0.774 | 0.774 | 0.695 |
| ScholarStack | 0.711 | 0.798 | 0.798 | 0.781 | 0.710 | 0.807 | 0.807 | 0.753 |
| Domain | SAGE queries | Source papers |
|---|---|---|
| Computer science | 29 | 1,754 |
| Natural science | 28 | 1,749 |
| Healthcare | 32 | 1,748 |
| Field | Count | Field | Count | Field | Count |
|---|---|---|---|---|---|
| Linguistics | 5 | Physics | 4 | Education | 2 |
| Medicine | 5 | Psychology | 4 | Art | 1 |
| Engineering | 5 | Business | 4 | Chemistry | 1 |
| Mathematics | 4 | Biology | 4 | Materials Sci. | 1 |
| Answerability | Generation (Rouge-L) | |||||||
| Method | Ans-F1 | Unans-F1 | Acc. | W-F1 | Macro-F1 | FF | AE | Tokens |
| Full-text | 0.806 | 0.507 | 0.721 | 0.738 | 0.656 | 0.155 | 0.205 | 12,577 |
| Human (reference) | 0.732 | – | – | – | – | 0.150 | 0.161 | 301 |
| RAG10 | 0.736 | 0.476 | 0.649 | 0.677 | 0.606 | 0.142 | 0.156 | 1,864 |
| RAG20 | 0.766 | 0.489 | 0.679 | 0.703 | 0.627 | 0.149 | 0.184 | 3,481 |
| ScholarStack | 0.749 | 0.478 | 0.661 | 0.687 | 0.613 | 0.145 | 0.156 | 1,350 |
| Method | Ans-F1 | Unans-F1 | Acc. | W-F1 | Macro-F1 | RgL | RgL |
|---|---|---|---|---|---|---|---|
| Machine learning (313 answerability / 158 generation) | |||||||
| Full-text | 0.808 | 0.517 | 0.725 | 0.747 | 0.662 | 0.142 | 0.199 |
| RAG | 0.722 | 0.460 | 0.633 | 0.666 | 0.591 | 0.130 | 0.149 |
| ScholarStack | 0.729 | 0.459 | 0.639 | 0.672 | 0.594 | 0.128 | 0.141 |
| NLP (133 / 69) | |||||||
| Full-text | 0.822 | 0.593 | 0.752 | 0.761 | 0.707 | 0.194 | 0.235 |
| Question | Gold | Full-text | RAG | ScholarStack |
|---|---|---|---|---|
| QASPER | ||||
| What inference models are used? | XGBoost, AdaBoost, Random Forest | XGBoost, AdaBoost and Random Forest | Gaussian Processes | XGBoost, AdaBoost, Random Forest |
| How many actions are present in the dataset? | 14,769 | 14,769 | No Answer | 14,769 |
| PeerQA | ||||
| Does the optimization happen only over map locations observed so far? | No; it also covers unobserved (hallucinated) regions | No; any region observed or hallucinated | No Answer | No; goals from any observed or hallucinated region |
| What values are represented in Table 2? | Zero-shot classification performance across 5 datasets and 14 language models | Zero-shot classification performance for all models and datasets | No Answer | Performance statistics for all models and datasets |
| Method | Corr. | Comp. | Synth. | Cal. | Dir. | Total | Tokens | Sec. |
|---|---|---|---|---|---|---|---|---|
| Full-text + Bare | 1.11 | 2.01 | 1.41 | 0.92 | 3.01 | 8.46 | 20.8k | 13.13 |
| ScholarStack (L1) + Bare | 3.90 | 3.56 | 2.99 | 3.78 | 3.82 | 18.04 | 4.1k | 9.83 |
| ScholarStack (L1) + Skill | 3.91 | 3.60 | 3.03 | 3.82 | 3.83 | 18.18 | 4.2k | 10.45 |
| ScholarStack (L1+L2) + Bare | 3.82 | 3.25 | 2.80 | 3.26 | 3.88 | 17.01 | 4.8k | 7.20 |
| ScholarStack (L1+L2) + Skill | 3.89 | 3.53 | 3.02 | 3.78 | 3.84 | 18.05 | 4.9k | 10.67 |
| ScholarStack (L1+L2+L3) + Bare | 3.89 | 3.35 | 3.48 | 3.90 | 3.80 | 18.41 | 5.5k | 7.54 |
| System | Input(k) | Output(k) | Cache hit(%) | Cost(USD) |
|---|---|---|---|---|
| Claude Code DR | 3,128 | 106 | 78.94 | 7.18 |
| ScholarStack | 10,533 | 117 | 96.02 | 3.49 |
| Field | Count | Field | Count | Field | Count |
|---|---|---|---|---|---|
| Medicine | 5 | Physics | 4 | Education | 2 |
| Engineering | 5 | Business | 4 | Art | 1 |
| Linguistics | 4 | Biology | 4 | Chemistry | 1 |
| Mathematics | 4 | Psychology | 3 | Materials Sci. | 1 |
| Domain | Count |
|---|---|
| AI and engineering | 14 |
| Chemistry and materials | 12 |
| Life sciences and medicine | 4 |
| Public health and social sciences | 3 |
| Environmental science | 2 |
| Total | 35 |
| Dimension | Definition |
|---|---|
| Novelty | Are the problems or approaches new? Is this a novel combination of familiar techniques? Is it clear how this work differs from previous contributions? Is related work adequately referenced? |
| Significance | Is the idea important? Are other people (practitioners or researchers) likely to use these ideas or build on them? Does the idea address a difficult problem in a better way than previous research? Does it provide a unique theoretical or pragmatic approach? |
| Clarity | Is the paper clearly written? Is it well-organized? Does it adequately inform the reader? |
| Feasibility | Can the idea be realized with existing technology or methods? Are there any technical difficulties or bottlenecks? Is the idea clear and logical? Are there any obvious errors or unreasonable parts in the idea, and can the experiments be designed normally according to this idea. |
| Effectiveness | How likely the proposed idea is going to work well (e.g., better than existing baselines). |