EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights
Organizations: Carnegie Mellon University, Language Technologies Institute · Stanford University, Department of Computer Science · Yale University, Quantitative Biology Institute · Stanford University, Department of Geophysics · Carnegie Mellon University, Department of Chemistry · Massachusetts Institute of Technology, Plasma Science and Fusion Center · Columbia University, Department of Applied Physics and Applied Mathematics · Flatiron Institute, Center for Computational Astrophysics · Princeton University, Department of Astrophysical Sciences · Princeton University, Department of Computer Science · Engram
Abstract
When Isaac Newton discovered the law of gravitation, he did so through an iterative process of analyzing observed data such as planetary patterns, finding the underlying mechanisms by describing patterns in mathematical equations, and refining his theory against the Moon's orbit, revealing the startling insight that the same force governs both falling apples and orbiting planets. Would it be possible for AI agents to make similar discoveries? To measure this ability, we introduce EurekaBench, a cross-domain benchmark that tests AI agents' ability to conduct long-horizon experiments and discover mechanisms that explain observations. We evaluate these mechanisms by the scientific insights that can be derived from them. EurekaBench contains an expert-verified set of 26 long-horizon tasks across neuroscience, computer science, chemistry, astrophysics, geophysics, and plasma physics, with a total of 306 scientific insights that the discovered mechanisms are expected to support. Our evaluation framework tests three axes of scientific discovery: agents' ability to follow known scientific constraints, the predictive accuracy of the discovered mechanisms, and whether these mechanisms yield scientific insights or inform future research. Our results show that current AI agents often overly fixate on predictive accuracy optimization, surpassing human scientists, while falling substantially short in deriving scientific insights.
Figures & tables
Appendix figures & tables29 assets
Supplementary material from the paper’s appendix.
Appendix
| Task specification | |
|---|---|
| Context | You are a systems neuroscientist studying {problem_description} . |
| {scientific_background} {observations_block} | |
| Initial observations | However, it is unclear {initial_observations} . |
| Task | Your task is to discover the mechanism that explains this phenomenon. {mechanism_requirements} To discover this, you interrogate {environment_interface} through controlled numerical experiments designed to identify {experimental_objectives} . The structure of your predictive form needs to follow from the mechanism you uncover: for each ingredient, explain why it is included, what it traces in the circuit, and how it shapes the neuron’s activity. |
| Evaluation | Every claim in your discovered mechanism needs to be backed by data you actually collected. {predictive_accuracy_evaluation} Your experimental findings will also be reviewed by an independent researcher who has access to the same simulator, to assess whether the mechanism you discover can yield new insights. |
| Deliverables | You need to produce three deliverables. If any of them is missing or does not follow the requirements below, your solution receives a score of 0. Deliverable 1: Write your discovered mechanism to {output_dir} /mechanisms/mechanism.md. State the mechanism as the causal chain from the coarse action state of the population to the activity of one neuron; its executable form, with every ingredient and fitted constant defined and the physical process each one traces identified; why its inputs jointly determine the neuron’s activity; the concrete evidence from your experiments that supports the mechanism and its form, including the evidence that let you reject the competing explanations you considered; and the insights about how population activity organizes the activity of individual neurons that follow from what you found. |
| Problem | Goal |
| Neural dynamics | A good mechanism is not only accurate but also extrapolatable, and it can even open up more interesting new research directions. Through such a mechanism people come to understand how the nervous system of C. elegans organizes its activity into a hierarchical locomotion structure, down to how the coarse action state shapes a single neuron. |
| ARC tokamak | A good mechanism is not only accurate but also extrapolatable, and it can even open up more interesting new research directions. Through such a mechanism people come to understand how the transport a model returns sets the profiles an ARC plasma settles into and the fusion power it produces, down to why two models that leave the plasma with the same energy confinement disagree by a factor of two in power. |
| Judge | Reversals (%) | ||
| Both judges (average) | 0.20 [0.07, 0.32] | 39.9 [33.2, 46.5] | 0.008 |
| Claude Opus 4.8 | 0.15 [0.03, 0.26] | 41.8 [35.3, 48.4] | 0.054 |
| GPT 5.6 Sol | 0.15 [0.01, 0.29] | 42.8 [34.3, 51.9] | 0.054 |
| Scientific domains | Problem names | Sources |
| Neuroscience | Premotor states | Morrison and Young,2025 |
| Zigzag foraging | Zhao et al.,2024 | |
| Biological networks | Chen et al.,2025a | |
| Neural dynamics | Fieseler et al.,2025 | |
| Navigation strategies | Chen et al.,2026 | |
| Phase memory | Dunn et al.,2025 |