ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research
Organizations: Stanford University · Seoul National University · University of Washington · Carnegie Mellon University · Allen Institute for AI · MIT
Abstract
What makes great scientists great? Even as AI systems start to make progress on open problems, scientists remain far ahead of them at sensing which prior idea, buried in an ever-growing archive of research, a new problem needs. To study this skill, we draw on researchers who know firsthand which earlier work advanced their completed projects, with papers serving as pointers to the ideas within. Using our automated pipeline that makes author annotation scalable, we build ScholarCatalyst by having 184 lead authors of 207 recent computer science papers label which candidates did or could have advanced their project, each with a detailed rationale. We introduce a retrieval task with author-provided judgments: given an initial research question, retrieve these papers from only the literature available when the project began. Agentic search does no better than embedding retrieval (0.42 vs. 0.48 Recall@20) despite calling that same retriever as a tool. Even an agent built on Claude Fable 5.1, which may have seen the completed papers during training, reaches only 0.51 R@20. These results highlight the need for new training recipes that equip models with expert intuition for searching broad corpora. We envision ScholarCatalyst as a step toward scientific agents that can take a half-formed idea and point to the prior research it needs.
Figures & tables
| Benchmark | Relevance | Pre-discovery query | Author annotation | Undocumented relevance | Queries | Corpus |
|---|---|---|---|---|---|---|
| SciFact ( Wadden et al., 2020 ) | Claim support | ✗ | ✗ | ✗ | 1.4K | 5.2K |
| DORIS-MAE ( Wang et al., 2023 ) | Multi-aspect relevance | ✗ | ✗ | ✗ | 100 | 100 |
| ScholarQABench ( Asai et al., 2024 ) | Claim support | ✗ | ✗ | ✗ | 2,967 | 45M |
| LitSearch ( Ajith et al., 2024 ) | Citation relation | ✗ | ✓ | ✗ | 597 | 64K |
| MIR ( Garikaparthi et al., 2025 ) | Methodological citation intent | ✓ | ✗ | ✗ | 139 | 4.7K |
| ScholarCatalyst (ours) | Author-judged inspiration | ✓ | ✓ | ✓ | 894 | 191K |
| Query type | Queries | Avg. | Cited / Uncited Pos. (%) | Avg. | Avg. Length |
|---|---|---|---|---|---|
| Core Research Query | 207 | 3.19 | 95.5 / 4.5 | 10.01 | 131.8 |
| Subfield-specific Query | 687 | 6.09 | 56.4 / 43.6 | 6.88 | 63.5 |
| Core Research Query (n=207) | Subfield-specific Query (n=687) | |||||||
| Model | nDCG@20 | R@5 | R@20 | R@100 | nDCG@20 | R@5 | R@20 | R@100 |
| BM25 | 0.16 | 0.12 | 0.23 | 0.39 | 0.21 | 0.12 | 0.26 | 0.45 |
| LateOn 0.1B | 0.16 | 0.12 | 0.24 | 0.37 | 0.27 | 0.16 | 0.33 | 0.55 |
| SPECTER2 | 0.13 | 0.10 | 0.20 | 0.38 | 0.18 | 0.11 | 0.23 | 0.42 |
| OpenScholar | 0.16 | 0.13 | 0.23 | 0.40 | 0.19 | 0.12 | 0.24 | 0.46 |
| Qwen3-Emb-4B | 0.25 | 0.18 | 0.39 | 0.59 | 0.35 | 0.21 | 0.44 | 0.69 |
Appendix figures & tables35 assets
Supplementary material from the paper’s appendix.
Appendix
| Component | Information |
|---|---|
| Source paper ( ) | The completed paper defining the instance. Withheld from the retriever, so its solution and bibliography cannot be used to find prior work. |
| Search corpus ( ) | Papers published before (temporal cutoff). |
| Core research query CoreQ | Asks which prior ideas could help address the paper’s central research question. |
| Subfield-specific query SubQ | Asks which prior ideas from one specific research direction could help advance the paper. |
| Positive papers ( ) | Papers in that authors judge to offer an insight that could inspire work on . |
| Hard negatives ( ) | Papers in that are topically related to but judged not to offer such an idea. |
| Source Paper |
| Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models ( cs.RO ) |
| Core Research Query |
| Robots working in human-centric settings receive language that is far messier than the atomic commands studied in language-conditioned imitation learning. A line of work trains policies to map instructions such as ’pick up the coke can’ to actions and generalizes well within that regime, but real requests carry constraints, preferences, and mid-execution corrections ("that’s not trash", "I’m allergic to pickles") whose meaning depends on what the robot is currently seeing and doing. A separate line of work uses LLMs and VLMs off the shelf to decompose long-horizon tasks over predefined or hand-designed skills. These systems parse richer language but inherit the limited dexterity of their skill libraries and rarely react to a user who interrupts mid-task. Both lines leave a gap between the semantics a system can interpret and the physical behavior it can actually produce, and it is not obvious whether that gap is best closed by making language interpretation more capable or by making dexterous learned policies more promptable. How can a robot with dexterous, learned manipulation skills interpret open-ended prompts and real-time human feedback, grounded in its own observations, and act on them while the task is underway? |
| Subfield-Specific Query 1 |
| Vision-language-action models fine-tune pretrained VLMs to emit low-level continuous actions conditioned on images and a language instruction, and have shown strong dexterity and some language generalization on real hardware. However, their instruction supervision comes from demonstration annotations that are short and literal ("put the cup on the plate"), so it is unclear how much of their language-following ability survives when a prompt carries constraints, negations, preferences, or a correction issued partway through a rollout. What language can an end-to-end VLA actually be steered by at execution time, and where does its instruction-following break down as prompts depart from the literal commands seen in its training data? |
| Subfield-Specific Query 2 |
| Primary category | Count | % |
|---|---|---|
| cs.LG (Machine Learning) | 66 | 31.9% |
| cs.CV (Computer Vision) | 46 | 22.2% |
| cs.CL (Computation and Language) | 38 | 18.4% |
| cs.AI (Artificial Intelligence) | 23 | 11.1% |
| cs.RO (Robotics) | 20 | 9.7% |
| Other | 14 | 6.8% |
| BM25 Sim. | Emb. Sim. | |||
|---|---|---|---|---|
| Query | ||||
| CoreQ | 28.7 | 29.2 | 64.5 | 66.2 |
| SubQ | 34.5 | 34.4 | 65.1 | 64.2 |
| CoreQ . How can RL agents leverage the temporal dynamics of spiking neural networks when randomly initialized networks fail to collect sequences long enough to expose those dynamics? | ||
|---|---|---|
| CoreQ Positive | CoreQ Hard Negative | |
| Cited | Elucidating the theoretical underpinnings of surrogate gradient learning in SNNs; Neuromorphic Attitude Estimation and Control | Evolving Connectivity for Recurrent Spiking Neural Networks; Deep RL with Spiking Q-learning |
| Uncited | Jump-Start Reinforcement Learning; Revisiting the Minimalist Approach to Offline RL | 8 further topically adjacent papers, including surrogate-gradient variants and neuromorphic control work |
| Author rationale for uncited positives. Jump-Start Reinforcement Learning suggests using a secondary controller to bridge the warmup period, enabling the spiking network to learn from the start. Revisiting the Minimalist Approach to Offline RL supports combining secondary-controller demonstrations with the spiking actor’s own TD3 rollouts. | ||
| Inspiration type | Definition |
|---|---|
| Method | A specific method, mechanism, component, formulation, or principle from the source work is adopted or adapted. |
| Generalization | An idea, finding, or method from the source work is extended to a new domain, modality, task, setting, or broader scope. |
| Limitation | A limitation, failure mode, or unresolved issue in the source work motivates the target work to address it. |
| Framework | The target work builds on the source work’s overall framework, model, or system rather than a specific component. |
| Phenomenon | An empirical phenomenon or observation in the source work motivates the target work to explain, characterize, or formalize it. |
| Reframing | An insight from the source work provides a new perspective, abstraction, or interpretation of the target problem. |
| Source Paper Towards Understanding the Mechanisms of Classifier-Free Guidance ( cs.CV ) (NeurIPS 2025 Spotlight) |
|---|
| Core Research Question Classifier-free guidance (CFG) is a highly effective inference-time guidance method that significantly improves sample quality and condition alignment in diffusion models, but its underlying mechanisms remain poorly understood. Existing theoretical analyses mainly study isotropic Gaussian mixtures or one-dimensional distributions, where much of the complex structure of natural images is absent. This made us wonder (i) whether we could construct a simplified model in which the effects of CFG on image structures become directly visible, hence capturing the underlying mechanism of CFG more faithfully than the simplified settings considered in prior theoretical analyses and (ii) whether the mechanisms identified in this simplified setting extend beyond itself and can help explain, at least partially explain the CFG’s mechanism in real-world diffusion models? |
| Key Inspiration Paper: Abid, Zhang, Bagaria & Zou (2018). Exploring patterns enriched in a dataset with contrastive principal component analysis. Nature Communications 9, 2134. |
| Rationale Contribution: This paper shows that the eigendirections associated with differences between the covariance structures of two datasets can reveal features that are distinctive to one dataset relative to the other. These directions, referred to as contrastive principal components, characterize prominent structures that distinguish the target data from the background data. |
| New idea I brought into my own research: When analyzing CFG in our simplified linear model, I found that the guidance term involves the difference between two covariance matrices: the covariance of the conditional distribution and that of the unconditional distribution (CPC guidance component in equation (12) of my paper). This suggested that the effect of CFG might be understood through the structure of this covariance difference. I then searched for prior work studying such differences between covariance structures and came across contrastive PCA. This connection suggested that the relevant eigendirections could identify structures that are particularly prominent in the conditional distribution relative to the unconditional distribution, which we later show contribute to the CFG’s effect. |
| Why I expect that idea to work: The mathematical structure arising in our analysis of CFG closely resembles that underlying contrastive PCA: both are governed by differences between the covariance structures of two distributions. This similarity suggested that the interpretation of contrastive principal components could provide useful insight into how CFG emphasizes structures that distinguish the conditional distribution from the unconditional one. |
| Key Inspiration Paper: Self-Consistency Improves Chain of Thought Reasoning in Language Models ( cs.CL , arxiv:2203.11171 ) |
|---|
| Source Paper: Benefits and Limitations of Communication in Multi-Agent Reasoning ( cs.MA , arxiv:2510.13903 ) |
| Role: naive baseline. Proposes multiple CoT with a voting mechanism at the end to improve performance. This work seemed to be the perfect naive baseline as the simplest "multi-agent" protocol where multiple agents think and do not communicate during their thinking process. Here all agents have all the context and think the full CoT, in contrast with multi-agent approaches where context is chunked and CoT is distributed. We used this as a baseline in our empirical experiments section. We expected this to be a good naive baseline as it does not adaptively allocate resources and models are given full context. These were the two main levers by which we conceived multi-agent reasoning to be superior. |
| Source Paper Self-Consistency for LLM-Based Motion Trajectory Generation and Verification ( cs.CV , arxiv:2603.29301 ) |
| Role: generalization to a new modality. The key work that our approach extends from, we bring it to the visual domain, whereas this work only works for language-based tasks. |
| Source Paper Accelerated Test-Time Scaling with Model-Free Speculative Sampling ( cs.CL , arxiv:2506.04708 ) |
| Role: structural premise to exploit. This paper proposes parallel test-time scaling, which is one of the core bases of our work. We aim to accelerate test-time scaling, along the setting that the given paper suggested. We expected the idea to work because speculative decoding is already known to accelerate decoding, and we found redundancy to leverage. |
| Model | Size | Arch. | Max. | Max. | Instr. | Link |
| Sparse retriever | ||||||
| BM25 ( Robertson and Zaragoza, 2009 ) | — | Sparse | No | |||
| Multi-vector retriever | ||||||
| LateOn ( Sourty et al., 2026 ) | 149M | ColBERT | 32 | 300 | No | |
| Scientific embedding retrievers | ||||||
| SPECTER2 ( Singh et al., 2022 ) | 110M | Encoder | 512 | 512 | No | |
| Model | Knowledge Cutoff |
|---|---|
| Under the knowledge-cutoff rule | |
| GPT-4.1 ( OpenAI, 2026a ) | June 2024 |
| o3 ( OpenAI, 2026e ) | June 2024 |
| Analysis backbones | |
| Gemini 2.5 Flash ( Comanici et al., 2025 ) | January 2025 |
| Gemini 3.1 Pro ( Google DeepMind, 2026a ) | January 2025 |
| Mean | Median | Min | Max | ||
|---|---|---|---|---|---|
| Full corpus (no cutoff) | 190,896 | ||||
| CoreQ (cutoff) | 207 | 190,306 | 190,574 | 188,526 | 190,896 |
| SubQ (cutoff) | 687 | 190,300 | 190,574 | 188,526 | 190,896 |
| All queries (cutoff) | 894 | 190,302 | 190,574 | 188,526 | 190,896 |
| Core Research Query | Subfield-specific Query | |||||||
| Model | nDCG@20 | R@5 | R@20 | R@100 | nDCG@20 | R@5 | R@20 | R@100 |
| Sparse retriever | ||||||||
| BM25 | 0.16 | 0.12 | 0.23 | 0.39 | 0.21 | 0.12 | 0.26 | 0.45 |
| Multi-vector retriever | ||||||||
| LateOn 0.1B | 0.16 | 0.12 | 0.24 | 0.37 | 0.27 | 0.16 | 0.33 | 0.55 |
| Scientific embedding retrievers | ||||||||
| Core Research Query | Subfield-specific Query | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Verifier | Condition | ||||||||
| GPT-4.1 | Standard | 0.27 | 0.29 | 0.26 | 0.28 | 0.18 | 0.17 | 0.15 | 0.16 |
| Oracle | 0.38 | 0.32 | 0.29 | 0.29 | 0.26 | 0.20 | 0.18 | 0.17 | |
| Gemini 2.5 Flash | Standard | 0.22 | 0.24 | 0.20 | 0.24 | 0.14 | 0.14 | 0.15 | 0.15 |
| Oracle | 0.36 | 0.30 | 0.30 | 0.24 | 0.26 | 0.23 | 0.18 | 0.14 | |
| Gemini 3.1 Pro | Standard | 0.31 | 0.30 | 0.27 | 0.30 | 0.18 | 0.19 | 0.19 | 0.21 |
| Pre-cutoff slice | Post-cutoff slice | |||||||
| System | R@5 | R@20 | R@100 | Traj. | R@5 | R@20 | R@100 | Traj. |
| Gemini-Emb-2 retriever | 0.23 | 0.35 | 0.49 | 0.25 | 0.41 | 0.63 | ||
| Tool-calling agent | ||||||||
| GPT-4.1 (June 2024) | 0.20 | 0.32 | 0.44 | 0.33 | 0.25 | 0.39 | 0.60 | 0.42 |
| Gemini 3.1 Pro (January 2025) | 0.21 | 0.36 | 0.47 | 0.40 | 0.26 | 0.37 | 0.53 | 0.45 |
| GPT-5.6 Sol (February 2026) | 0.21 | 0.42 | 0.50 | 0.49 | 0.25 | 0.47 | 0.64 | 0.58 |
| Pre-cutoff slice | Post-cutoff slice | |||||||
|---|---|---|---|---|---|---|---|---|
| Condition | R@5 | R@20 | R@100 | Traj. | R@5 | R@20 | R@100 | Traj. |
| Local corpus only | 0.29 | 0.37 | 0.49 | 0.38 | 0.28 | 0.41 | 0.59 | 0.41 |
| + source paper (title and abstract) | 0.34 | 0.48 | 0.57 | 0.51 | 0.34 | 0.52 | 0.69 | 0.56 |
| + web search | 0.29 | 0.34 | 0.48 | 0.37 | 0.28 | 0.42 | 0.61 | 0.44 |
| + source-paper bibliography | 0.36 | 0.68 | 0.93 | 0.93 | 0.42 | 0.81 | 0.98 | 0.98 |
| Source-paper bibliography alone | 0.23 | 0.60 | 0.93 | 0.93 | 0.18 | 0.60 | 0.98 | 0.98 |
| Original query | Enhancing the reasoning capabilities of large language models traditionally involves supervised fine-tuning followed by reinforcement learning applied to reasoning-specific datasets. Recent advancements have shown that applying reinforcement learning directly to base models can autonomously elicit sophisticated reasoning behaviors without relying on supervised fine-tuning data. However, these methods still depend on rule-based verifiers, verified golden answers, or pre-trained reward models to generate reward signals during the reinforcement learning phase. Relying on such explicit feedback limits scalability in domains where ground-truth verification is costly or unavailable. Can we enhance the reasoning capability of LLMs in a fully unsupervised manner? |
|---|---|
| Single-query expansion | Enhancing the reasoning capabilities of large language models traditionally involves supervised fine-tuning followed by reinforcement learning applied to reasoning-specific datasets. Recent advancements have shown that applying reinforcement learning directly to base models can autonomously elicit sophisticated reasoning behaviors without relying on supervised fine-tuning data. However, these methods still depend on rule-based verifiers, verified golden answers, or pre-trained reward models to generate reward signals during the reinforcement learning phase. Relying on such explicit feedback limits scalability in domains where ground-truth verification is costly or unavailable. Can we enhance the reasoning capability of LLMs in a fully unsupervised manner? I am particularly interested in prior work exploring self-supervised objectives, intrinsic motivation, or unsupervised reward shaping techniques for language models, as well as approaches that leverage emergent properties or self-consistency signals to improve reasoning without external supervision. |
| Multi-query generation | (1) Enhancing the reasoning capability of LLMs in a fully unsupervised manner? I am particularly interested in approaches that leverage self-supervised objectives or intrinsic motivation signals to guide reasoning skill acquisition without any external reward models. (2) Enhancing the reasoning capability of LLMs in a fully unsupervised manner? I would like to explore whether emergent reasoning behaviors can be fostered through unsupervised pre-training on synthetic or automatically generated reasoning tasks. (3) Enhancing the reasoning capability of LLMs in a fully unsupervised manner? I am curious about prior work that investigates the use of mutual information maximization or contrastive learning techniques to improve reasoning without labeled data. (4) Enhancing the reasoning capability of LLMs in a fully unsupervised manner? I am seeking studies that examine the role of curriculum learning or self-play in enabling models to develop reasoning skills autonomously. (5) Enhancing the reasoning capability of LLMs in a fully unsupervised manner? I am interested in research that analyzes the emergence of reasoning abilities in large-scale language models trained solely on raw text corpora, without any explicit reasoning supervision. |
| Original query | Enhancing the reasoning capabilities of large language models traditionally involves supervised fine-tuning followed by reinforcement learning applied to reasoning-specific datasets. Recent advancements have shown that applying reinforcement learning directly to base models can autonomously elicit sophisticated reasoning behaviors without relying on supervised fine-tuning data. However, these methods still depend on rule-based verifiers, verified golden answers, or pre-trained reward models to generate reward signals during the reinforcement learning phase. Relying on such explicit feedback limits scalability in domains where ground-truth verification is costly or unavailable. Can we enhance the reasoning capability of LLMs in a fully unsupervised manner? |
|---|---|
| Hypothetical abstract 1 | Emergent Reasoning in Language Models via Self-Consistency Training. Recent work has demonstrated that large language models can exhibit improved reasoning abilities when trained to produce self-consistent outputs. In this study, we investigate a self-consistency training paradigm where models are encouraged to generate multiple solutions to a given problem and are rewarded for producing answers that agree with the majority of their own outputs. We apply this approach to arithmetic and commonsense reasoning tasks, showing that self-consistency training leads to modest improvements in accuracy without requiring external supervision or gold-standard answers. Our analysis suggests that self-consistency acts as a form of internal regularization, promoting the emergence of more robust reasoning patterns. These findings highlight the potential of leveraging a model’s own outputs as a weak supervisory signal for enhancing reasoning capabilities. |
| Hypothetical abstract 2 | Unsupervised Skill Acquisition in Reinforcement Learning via Intrinsic Motivation. Traditional reinforcement learning relies on extrinsic rewards provided by the environment, which can be sparse or unavailable in many real-world scenarios. We propose an unsupervised RL framework where agents are driven by intrinsic motivation signals, such as curiosity or novelty, to autonomously acquire diverse skills. Our method employs a self-supervised objective that rewards agents for exploring novel states and behaviors, without access to task-specific rewards or external supervision. Experiments in simulated environments demonstrate that agents trained with intrinsic motivation develop a repertoire of complex behaviors, including multi-step reasoning and problem-solving strategies. This work suggests that unsupervised RL can serve as a foundation for skill acquisition in domains where explicit reward signals are difficult to define. |
| Hypothetical abstract 3 | Bootstrapping Commonsense Knowledge in Neural Language Models via Iterative Self-Labeling. Acquiring commonsense knowledge remains a challenge for neural language models, particularly in the absence of large-scale annotated datasets. We introduce an iterative self-labeling approach in which a language model generates candidate answers to commonsense questions and then refines its predictions by training on its own high-confidence outputs. Our method leverages confidence estimation to select pseudo-labels, enabling the model to bootstrap its knowledge without external supervision. We evaluate our approach on several commonsense reasoning benchmarks and observe consistent improvements over baseline models trained without self-labeling. Our results indicate that iterative self-labeling can be an effective strategy for enhancing the reasoning abilities of language models in a data-efficient and unsupervised manner. |
| Core Research Query | Subfield-specific Query | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Read policy | nDCG@20 | R@5 | R@20 | R@100 | nDCG@20 | R@5 | R@20 | R@100 | Traj. |
| Gemini-Emb-2 retriever | 0.25 | 0.19 | 0.35 | 0.61 | 0.43 | 0.27 | 0.55 | 0.79 | |
| Abstract only (no read tool) | 0.25 | 0.22 | 0.35 | 0.58 | 0.41 | 0.32 | 0.46 | 0.76 | 0.37 / 0.52 |
| Abstract + intro | 0.26 | 0.22 | 0.34 | 0.61 | 0.41 | 0.29 | 0.48 | 0.78 | 0.36 / 0.51 |
| Abstract + method | 0.24 | 0.22 | 0.32 | 0.61 | 0.41 | 0.28 | 0.47 | 0.78 | 0.33 / 0.53 |
| Abstract + bib | 0.26 | 0.23 | 0.34 | 0.61 | 0.41 | 0.30 | 0.48 | 0.79 | 0.36 / 0.52 |
| All positives (%) | Agent top-20 rate (%) | |||||
| Retriever top 100 | Outside top 100 | Retriever top 100 | Outside top 100 | |||
| Agent | Top 20 | Missed | Top 20 | Missed | Top 20 | Top 20 |
| Tool-calling agent (GPT-4.1) | 40% | 30% | 1% | 29% | 57% | 4% |
| Deep research agent (o3) | 34% | 36% | 2% | 28% | 48% | 6% |
| Grep agent (GPT-4.1) | 6% | 64% | 1% | 29% | 8% | 4% |
| Verdict | Author’s rationale | Backbone’s rationale (GPT-4.1) |
| TTRL: Test-Time Reinforcement Learning | ||
| Same relationship | This is the most significant paper that impacted my initial experiments. I was following their setup and used their code to run test-time training with unsupervised labels, but I found the correctness of the labels seems to not matter too much on the Qwen-Math models they used. […] | The TTRL paper’s main contribution is the demonstration that reinforcement learning can be effectively performed on reasoning tasks in large language models using only unlabeled data , by leveraging majority-voted answers as a proxy for ground-truth rewards. […] |
| Judge: Both rationales identify that TTRL demonstrates the effectiveness of using majority-voted, potentially noisy, unsupervised labels as rewards for reinforcement learning in LLMs, directly informing the author’s […] | ||
| Consent in Crisis: The Rapid Decline of the AI Data Commons | ||
| Same relationship | The paper did a large-scale, longitudinal audit of the consent protocols for the web domains underlying AI training corpora, and pointed out there is a decline in data owners’ willingness to contribute data to AI training. This paper made an important observation […] | The key contribution of “Consent in Crisis: The Rapid Decline of the AI Data Commons” is its rigorous, longitudinal audit quantifying how quickly and extensively web data sources are restricting AI training use, particularly […] |
| Judge: Both rationales identify the paper’s empirical audit of declining data consent as foundational for understanding the real-world consequences of respecting opt-outs on LLM training, emphasizing the importance of the […] | ||
| Agent | Backbone | Tool calls | Backbone calls | Cost ($) | Total cost ($) | Median time (s) |
|---|---|---|---|---|---|---|
| Grep | GPT-4.1 | 20.3 | 18.5 | 0.37 | 332 | 22 |
| Tool-calling | GPT-4.1 | 5.8 | 44.3 | 0.05 | 42 | 25 |
| Deep research | o3 | 14.3 | 18.3 | 0.39 | 345 | 60 |
| Grep | Claude Fable 5.1 | 19.3 | 20.2 | 1.14 | 1,016 | 139 |
| Tool-calling | Claude Fable 5.1 | 13.9 | 63.2 | 1.32 | 1,178 | 88 |
| Deep research | Claude Fable 5.1 | 22.7 | 12.4 | 1.39 | 1,240 | 146 |
| Category | Metric | Value |
| Reconstructed query quality | CoreQ rated fully accurate | 60.9% |
| CoreQ rated accurate or mostly accurate | 98.1% | |
| SubQ rated fully accurate | 64.9% | |
| SubQ rated accurate or mostly accurate | 95.9% | |
| Candidate pool judgement | Original citation retained as positive | 85.1% |
| Pooled candidate promoted to positive | 35.1% |