cs.CLSep 30, 2026

EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights

Authors: Jiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, +7 more

Organizations: Carnegie Mellon University, Language Technologies Institute · Stanford University, Department of Computer Science · Yale University, Quantitative Biology Institute · Stanford University, Department of Geophysics · Carnegie Mellon University, Department of Chemistry · Massachusetts Institute of Technology, Plasma Science and Fusion Center · Columbia University, Department of Applied Physics and Applied Mathematics · Flatiron Institute, Center for Computational Astrophysics · Princeton University, Department of Astrophysical Sciences · Princeton University, Department of Computer Science · Engram

Abstract

When Isaac Newton discovered the law of gravitation, he did so through an iterative process of analyzing observed data such as planetary patterns, finding the underlying mechanisms by describing patterns in mathematical equations, and refining his theory against the Moon's orbit, revealing the startling insight that the same force governs both falling apples and orbiting planets. Would it be possible for AI agents to make similar discoveries? To measure this ability, we introduce EurekaBench, a cross-domain benchmark that tests AI agents' ability to conduct long-horizon experiments and discover mechanisms that explain observations. We evaluate these mechanisms by the scientific insights that can be derived from them. EurekaBench contains an expert-verified set of 26 long-horizon tasks across neuroscience, computer science, chemistry, astrophysics, geophysics, and plasma physics, with a total of 306 scientific insights that the discovered mechanisms are expected to support. Our evaluation framework tests three axes of scientific discovery: agents' ability to follow known scientific constraints, the predictive accuracy of the discovered mechanisms, and whether these mechanisms yield scientific insights or inform future research. Our results show that current AI agents often overly fixate on predictive accuracy optimization, surpassing human scientists, while falling substantially short in deriving scientific insights.

Figures & tables

Appendix figures & tables29 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Benchmarking AI Agents for Addressing Scientific Challenges Across Scales

    Jun 10, 2026Tianyu Liu, Allen Xin Wang, Antonia Panescu +30Scientific DiscoveryAgentic Benchmarks

  2. MechBench: Can AI Scientific Agents Discover Mechanisms Beyond Phenomenal Laws?

    Sep 28, 2026Zihan Yu, Jiadong Zhang, Jialin Cheng +2Scientific Agents