EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights
Authors: Jiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, +7 more
Organizations: Carnegie Mellon University, Language Technologies Institute · Stanford University, Department of Computer Science · Yale University, Quantitative Biology Institute · Stanford University, Department of Geophysics · Carnegie Mellon University, Department of Chemistry · Massachusetts Institute of Technology, Plasma Science and Fusion Center · Columbia University, Department of Applied Physics and Applied Mathematics · Flatiron Institute, Center for Computational Astrophysics · Princeton University, Department of Astrophysical Sciences · Princeton University, Department of Computer Science · Engram
When Isaac Newton discovered the law of gravitation, he did so through an iterative process of analyzing observed data such as planetary patterns, finding the underlying mechanisms by describing patterns in mathematical equations, and refining his theory against the Moon's orbit, revealing the startling insight that the same force governs both falling apples and orbiting planets. Would it be possible for AI agents to make similar discoveries? To measure this ability, we introduce EurekaBench, a cross-domain benchmark that tests AI agents' ability to conduct long-horizon experiments and discover mechanisms that explain observations. We evaluate these mechanisms by the scientific insights that can be derived from them. EurekaBench contains an expert-verified set of 26 long-horizon tasks across neuroscience, computer science, chemistry, astrophysics, geophysics, and plasma physics, with a total of 306 scientific insights that the discovered mechanisms are expected to support. Our evaluation framework tests three axes of scientific discovery: agents' ability to follow known scientific constraints, the predictive accuracy of the discovered mechanisms, and whether these mechanisms yield scientific insights or inform future research. Our results show that current AI agents often overly fixate on predictive accuracy optimization, surpassing human scientists, while falling substantially short in deriving scientific insights.
Figures & tables
Figure 1: The scientific discovery process in Eurekabench . In this example from plasma physics, scientists observe that different transport models predict different fusion power outputs for the same reactor under the same operating conditions (observation) . Through iterative experiments, they find that differences in turbulent transport change temperature and density distributions, explaining the different power predictions (mechanism) . This helps them understand how turbulent transport may affect fusion power in larger reactors or under stronger magnetic fields (insight) . Discovering such insights is a “eureka moment” in science.
Figure 2: Initial observations, evaluation, and results in Eurekabench . Left: Examples of initial observations from six domains: (a) neuroscience, (b) geophysics, (c) plasma physics, (d) astrophysics, (e) computer science, and (f) chemistry. Top right: An ice shelf flow-law example illustrates how an agent uses iterative experiments to propose a mechanism, which is evaluated for scientific constraints, predictive accuracy, and scientific insights. SSA denotes the shallow-shelf approximation. Bottom right: Predictive accuracy and scientific insight scores for seven AI agents and the human reference. The strongest agents approach human predictive accuracy while attaining substantially lower insight scores.
Figure 3: Scientist agent workflow in Eurekabench . Given a task specification and a documented environment interface, the scientist agent conducts iterative experiments to discover a mechanism and submits its description, implementation, and experimental records.
Figure 4: The 9 interpretable insight questions for the Antarctic ice shelf problem in geophysics in Eurekabench . These insights explain why ice deforms differently under compression and extension, how its resistance varies with direction, and which measurements determine its viscosity.
Figure 5: Problem distribution across domains.
Figure 6: Effects of goals and insight questions on mechanism discovery. We evaluate Claude Fable 5.1 with general guidance (+goal) or insight questions (+insights); stacked bars show estimated PA and SI contributions to the conditioned score, with Δ marking changes from baseline.
Figure 7: Insights discovered during the experimentation loop vs. those derived from the discovered mechanism. Both scores use the same problems and instances.
Figure 8: Expert review of mechanism discovery. Experts review experiment logs of Claude Fable 5.1 and GPT 6 Astra on neuroscience and plasma physics tasks, highlighting failures in mechanism discovery alongside successful hypothesis rejection and understanding of physical concepts.
Appendix figures & tables29 assets
Supplementary material from the paper’s appendix.
Appendix
Task specification
Context
You are a systems neuroscientist studying {problem_description} .
{scientific_background} {observations_block}
Initial observations
However, it is unclear {initial_observations} .
Task
Your task is to discover the mechanism that explains this phenomenon. {mechanism_requirements} To discover this, you interrogate {environment_interface} through controlled numerical experiments designed to identify {experimental_objectives} . The structure of your predictive form needs to follow from the mechanism you uncover: for each ingredient, explain why it is included, what it traces in the circuit, and how it shapes the neuron’s activity.
Evaluation
Every claim in your discovered mechanism needs to be backed by data you actually collected. {predictive_accuracy_evaluation} Your experimental findings will also be reviewed by an independent researcher who has access to the same simulator, to assess whether the mechanism you discover can yield new insights.
Deliverables
You need to produce three deliverables. If any of them is missing or does not follow the requirements below, your solution receives a score of 0. Deliverable 1: Write your discovered mechanism to {output_dir} /mechanisms/mechanism.md. State the mechanism as the causal chain from the coarse action state of the population to the activity of one neuron; its executable form, with every ingredient and fitted constant defined and the physical process each one traces identified; why its inputs jointly determine the neuron’s activity; the concrete evidence from your experiments that supports the mechanism and its form, including the evidence that let you reject the competing explanations you considered; and the insights about how population activity organizes the activity of individual neurons that follow from what you found.
Appendix
Table 3: Task specification template for the scientist agent (neuroscience).
Problem
Goal
Neural dynamics
A good mechanism is not only accurate but also extrapolatable, and it can even open up more interesting new research directions. Through such a mechanism people come to understand how the nervous system of C. elegans organizes its activity into a hierarchical locomotion structure, down to how the coarse action state shapes a single neuron.
ARC tokamak
A good mechanism is not only accurate but also extrapolatable, and it can even open up more interesting new research directions. Through such a mechanism people come to understand how the transport a model returns sets the profiles an ARC plasma settles into and the fusion power it produces, down to why two models that leave the plasma with the same energy confinement disagree by a factor of two in power.
Appendix
Table 4: Goal guidance used in the +goal condition.
Figure 9: PA and SI scores across domains. Predictive accuracy (a) and SI scores (b) for seven AI agents on 26 problems across six domains. Each cell shows one agent’s score on one problem, averaged over its instances using the same rules as Table 2 . Both panels use the same color scale. Gray dashes mark the excluded Kimi K3 run on catchment properties .
Judge
τb
Reversals (%)
p
Both judges (average)
0.20 [0.07, 0.32]
39.9 [33.2, 46.5]
0.008
Claude Opus 4.8
0.15 [0.03, 0.26]
41.8 [35.3, 48.4]
0.054
GPT 5.6 Sol
0.15 [0.01, 0.29]
42.8 [34.3, 51.9]
0.054
Appendix
Table 8: PA and SI rankings. Brackets show 95% bootstrap confidence intervals; p -values are computed using 100,000 within-problem permutations under the null hypothesis of no association and adjusted across the three rows using Holm’s method ( Phipson and Smyth, 2010 ; Holm, 1979 ) .
Figure 10: Missing tests of the target behavior. Original outputs from AI scientist agents.
Figure 11: Missing tests of the target behavior (continued).
Figure 12: Explaining only part of the target phenomenon. Original outputs from AI scientist agents.
Figure 13: Explaining only part of the target phenomenon (continued).
Figure 14: Explaining only part of the target phenomenon (continued).
Figure 15: Explaining only part of the target phenomenon (continued).
Figure 16: Prediction with limited explanation. Original outputs from AI scientist agents.
Figure 17: Prediction with limited explanation (continued).
Figure 18: Prediction with limited explanation (continued).
Figure 19: Prediction with limited explanation (continued).
Figure 20: Unclear explanations. Original outputs from AI scientist agents.
Figure 21: Unclear explanations (continued).
Figure 22: Overgeneralization. Original outputs from AI scientist agents.
Figure 23: Failing to identify and reduce uncertainty. Original outputs from AI scientist agents.
Figure 24: Treating task optimization as mechanism discovery. Original outputs from AI scientist agents.
Figure 25: Treating task optimization as mechanism discovery (continued).
Figure 26: Treating task optimization as mechanism discovery (continued).
Figure 27: Interesting observations without further exploration. Original outputs from AI scientist agents.
Figure 28: Conflicting results and conclusions. Original outputs from AI scientist agents.
Figure 29: Risk of circular validation. Original outputs from AI scientist agents.
Figure 30: Using controls to rule out alternative explanations. Original outputs from AI scientist agents.
Figure 31: Testing which model better predicts the observations. Original outputs from AI scientist agents.
Figure 32: Revising models after experimental feedback. Original outputs from AI scientist agents.
Figure 33: Informative observations. Original outputs from AI scientist agents.
AI agents are increasingly being developed to accelerate scientific discovery, yet their practical capabilities in real research settings remain poorly understood. Existing benchmarks for AI agents rarely capture the complexity, heterogeneity, and extended reasoning required by scientific work, whereas benchmarks for scientific tasks often reduce research to static, direct problems and provide limited support for interactive evaluation. Here, we introduce SciAgentArena, a systematic benchmark for evaluating AI agents in real-world scientific research scenarios drawn from emerging needs across multiple domains. SciAgentArena comprises approximately 200 tasks with stepwise verification and an interactive, agent-agnostic environment for assessing diverse AI agents. Using this benchmark, we find that current agents can contribute effectively to well-specified data-analysis workflows, particularly when the task structure and evaluation criteria are clear. However, their performance remains uneven across scientific contexts: agents struggle to generate genuinely novel insights, sustain self-directed exploration, and formulate robust solutions for open-ended research questions. We further characterize common failure modes across agents and identify opportunities for improving their reliability, autonomy, and scientific reasoning. Together, SciAgentArena provides a practical framework for measuring progress in AI agents for science and for guiding the design of future agents capable of addressing complex scientific challenges. Full codes, tasks, and datasets can be accessed via this link: https://sciagentarena.github.io/.
Tianyu Liu, Allen Xin Wang, Antonia Panescu +30
Yale University, CT, USA · The Pennsylvania State University, PA, USA · Northeastern University, MA, USA +7
Autonomous coding agents are increasingly proposed as AI-scientist systems that conduct analyses and write research reports, but executing a prescribed analysis is not the same as making a discovery. Existing benchmarks are configured for reproduction: tasks, data, and rubrics are built around a hidden target study, and recovery of its result is rewarded. We present TruthInsightBench, a benchmark configured for discovery. Its 40 blind tasks, drawn from 40 peer-reviewed studies across 10 scientific domains, expose only a neutral scientific objective and frozen data; source conclusions, expected values, and analysis paths are withheld, leaving the agent to determine what claim the data support. A fixed LLM-based judge scores the evidentiary maturity of an agent's own claims along six dimensions, operationalized as 29 artifact-grounded items, with automated, deterministic aggregation and no per-instance human grading, so evaluation can be repeated automatically as agents evolve. On one frozen base model, four coding agents form a narrow plateau (58.4-60.3 of 100) with no statistically reliable pairwise separation: they execute and document analyses competently, with comparatively strong evidence auditability and novelty, but largely lack the discriminating acts that establish a trustworthy claim (controls, robustness, falsifiability, and cross-dataset generalization). The bottleneck is scientific judgment rather than coding, and genuine discovery remains out of reach. TruthInsightBench makes this gap a measurable target; data and scoring code are at https://github.com/TruthInsight-stack/TruthInsightBench.
Scientific discovery requires not only recovering mathematical laws that describe observable behavior, but also identifying the mechanisms that generate them. Existing benchmarks for symbolic regression and scientific agents primarily evaluate phenomenal-law recovery, leaving mechanism discovery largely untested. We introduce MechBench, a benchmark that explicitly separates these two capabilities. Each task is defined by a mechanistic model, a structured set of scientifically meaningful relations whose joint consequences entail an observable phenomenal law, while agents receive only observational data and scientific context. We evaluate mechanism recovery through mechanism probes, which query internal scientific consequences that cannot be inferred from the phenomenal law alone. To reduce reliance on memorized textbook mechanisms, we construct unfamiliar variants through controlled, scientifically interpretable mutations of canonical mechanisms, and screen for mechanistic indistinguishability to exclude ambiguous instances admitting comparable competing mechanisms. Experiments across representative scientific agents reveal a substantial phenomenal--mechanism recovery gap: for Codex with GPT-5.6-sol, phenomenal-law accuracy reaches 35.00% on the Core-set while mechanism accuracy is only 13.75%, with mechanism recovery failing in 64.29% of cases where the phenomenal law is correctly recovered. The gap widens as mechanisms become increasingly mutated, and even providing the correct phenomenal law leaves mechanism recovery below 50%. These results reveal a substantial generalization gap in mechanistic reasoning and establish mechanism discovery as a distinct challenge beyond recovering observable scientific laws.
Zihan Yu, Jiadong Zhang, Jialin Cheng +2
Department of Electronic Engineering, BNRist, Tsinghua University, Beijing, China · Institute of Automation, Chinese Academy of Sciences, Beijing, China · Beijing Zhongguancun Academy, Beijing, China +2