Organizations: Department of Electronic Engineering, BNRist, Tsinghua University, Beijing, China · Institute of Automation, Chinese Academy of Sciences, Beijing, China · Beijing Zhongguancun Academy, Beijing, China · Xi’an Jiaotong University, Xi’an, China · Department of Earth System Science, Tsinghua University, Beijing, China
Scientific discovery requires not only recovering mathematical laws that describe observable behavior, but also identifying the mechanisms that generate them. Existing benchmarks for symbolic regression and scientific agents primarily evaluate phenomenal-law recovery, leaving mechanism discovery largely untested. We introduce MechBench, a benchmark that explicitly separates these two capabilities. Each task is defined by a mechanistic model, a structured set of scientifically meaningful relations whose joint consequences entail an observable phenomenal law, while agents receive only observational data and scientific context. We evaluate mechanism recovery through mechanism probes, which query internal scientific consequences that cannot be inferred from the phenomenal law alone. To reduce reliance on memorized textbook mechanisms, we construct unfamiliar variants through controlled, scientifically interpretable mutations of canonical mechanisms, and screen for mechanistic indistinguishability to exclude ambiguous instances admitting comparable competing mechanisms. Experiments across representative scientific agents reveal a substantial phenomenal--mechanism recovery gap: for Codex with GPT-5.6-sol, phenomenal-law accuracy reaches 35.00% on the Core-set while mechanism accuracy is only 13.75%, with mechanism recovery failing in 64.29% of cases where the phenomenal law is correctly recovered. The gap widens as mechanisms become increasingly mutated, and even providing the correct phenomenal law leaves mechanism recovery below 50%. These results reveal a substantial generalization gap in mechanistic reasoning and establish mechanism discovery as a distinct challenge beyond recovering observable scientific laws.
Figures & tables
Figure 1: From phenomenal-law recovery to mechanistic discovery.
Figure 2: Construction and evaluation of MechBench. (a) Canonical mechanisms are modified and composed to create unfamiliar variants. (b) Each mechanism induces a phenomenal law used to generate observations and held-out ID/OOD tests. (c) Candidate competing mechanisms are screened for mechanistic indistinguishability. (d) Submitted mechanisms are evaluated by both phenomenal-law recovery and mechanism probes.
Base Model
Agent
Core-set (n=80)
Full-set (n=512)
SA( P )↑
SA( M )↑
1−(M∣P)
SA( P )↑
SA( M )↑
1−(M∣P)
Deepseek-v4-flash-0731
PySR+Direct-Ask
0.00%
1.25%
/
0.20%
0.20%
/
GPT-5.6-sol
Codex
35.00%
13.75%
64.29%
16.02%
7.81%
57.32%
GLM-5.3-flash
Codex
6.25%
2.92%
67.22%
4.30%
1.76%
63.64%
GLM-5.3-flash
Claude Code
8.75%
3.75%
57.14%
/
/
/
Deepseek-v4-flash-0731
Codex
7.92%
3.75%
49.07%
/
/
/
Table 1: Symbolic recovery results on the Core-set and Full-set. SA(P) and SA(M) denote phenomenal-law and mechanism recovery, respectively.
Base Model
Core-set (n=80) · SA(M)↑
Full-set (n=512) · SA(M)↑
Standard discovery
Gold-P Direct-Ask
Standard discovery
Gold-P Direct-Ask
GPT-5.6-sol
13.75%
45.00%
7.81%
49.02%
GLM-5.3-flash
2.92%
6.25%
1.76%
3.91%
Deepseek-v4-flash-0731
3.75%
10.00%
/
4.10%
Table 2: Mechanism recovery under standard discovery and Gold-P Direct-Ask on the Core-set and Full-set.
Base model
Agent / condition
Physics
Chemistry
Biology
Materials
P
M
P
M
P
M
P
M
No LLM
PySR
0.00
–
0.00
–
0.00
–
0.00
–
DeepSeek-v4-flash-0731
PySR + Direct-Ask
–
5.00
–
0.00
–
0.00
–
0.00
GPT-5.6-sol
Codex
40.00
20.00
40.00
25.00
40.00
10.00
20.00
0.00
GLM-5.3-flash
Codex (3-run avg.)
11.67
11.67
5.00
0.00
3.33
0.00
5.00
0.00
GLM-5.3-flash
Claude Code
25.00
15.00
5.00
0.00
0.00
0.00
5.00
0.00
Table 3: Complete Core-set symbolic-recovery results by scientific domain. A dash denotes a metric that is not defined for the corresponding control or baseline.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Domain
Task family
Description
Physics
Circular Orbit
Models the orbital period of a test body using Newtonian inverse-square attraction, circular radial balance, and the speed–period relation.
Hall Response
Models the voltage measured across offset contacts using a classical carrier-conductivity tensor and an open transverse circuit.
Inclined Rolling
Models the downslope acceleration of a rigid body from coupled translational and rotational balances under no-slip rolling.
Terminal Settling
Models the terminal speed of a sphere from gravity, buoyancy, and linear Stokes drag at low Reynolds number.
Chemistry
Acid–Base Buffer Relaxation
Models transient pH response through acid–base buffer capacities and coupled proton-exchange population balances.
Competitive Inhibition Pulse Reaction
Models product response through mass-action balances among free enzyme, substrate-bound enzyme, and inhibitor-bound enzyme.
Appendix
Table 4: The 16 task families and the scientific constructions used for their Original reference mechanisms.
Mutation type
Scientific interpretation
Representative changes in the benchmark
Constitutive or component response
Change how a component property depends on its local state or an observable control.
Mobile carrier trapping in Hall Response; temperature-difference-dependent conductivity in Composite Heat Transport; pressure-dependent migration barrier in Vacancy Diffusion.
Internal state or storage
Add a scientifically meaningful population, compartment, reservoir, or relaxation mode.
Product-bound enzyme and surface-product intermediates; inactive channel state; slow thermal and anelastic reservoirs.
Additional transport or reaction channel
Open a parallel pathway that contributes a new flux while retaining the original pathway.
Finite-range orbital attraction, conductive heat bypass, fast grain-boundary path, and alternate vacancy-migration channel.
Source, sink, or balance modification
Add a force, production, loss, or exchange term to a governing balance.
Resisting torque in rolling, external lift in settling, direct Lindemann product formation, and feed-dependent mortality.
Geometric or kinematic constraint
Change how geometric quantities or motions are related without directly editing the final observable law.
Non-Euclidean circumferential radius, controlled rolling slip, conducting-area ratio, and pressure-dependent jump distance.
Appendix
Table 5: Scientific interpretations of the mechanism mutations.
Family
Mutation
Circular Orbit
Δ1 : Make relative inertia depend on radius.
Δ2 : Add a finite-range central attraction.
Δ3 : Modify circumferential geometry together with its radial derivative.
Δ4 : Add an external central trapping field.
Δ5 : Store tangential momentum in a coupled internal reservoir.
Hall Response
Δ1 : Trap a density-dependent fraction of electrons.
Appendix
Table 6: The five base mutations for each physics task family.
Family
Mutation
Acid–Base Buffer Relaxation
Δ1 : Split buffer sites between two acid families.
Δ2 : Introduce a slowly exchanging interior buffer region.
Δ3 : Add a diffusively coupled solvent pocket.
Δ4 : Add proton-catalyzed exchange without changing equilibrium capacity.
Δ5 : Insert a proton-storing relay between solution and primary buffer.
Competitive Inhibition Pulse Reaction
Δ1 : Add substrate-assisted catalytic release.
Appendix
Table 7: The five base mutations for each chemistry task family.
Family
Mutation
Allosteric Regulation
Δ1 : Add an inactive conformation binding three regulators.
Δ2 : Couple inactive-state substrate affinity to substrate exposure.
Δ3 : Add another productive active-state complex.
Δ4 : Partition regulator into a local compartment.
Table 11: Model–agent configurations, baselines, and controls used in the experiments.
Base model
Agent / condition
Physics
Chemistry
Biology
Materials
P
M
P
M
P
M
P
M
No LLM
PySR
0.78
–
0.00
–
0.00
–
0.00
–
DeepSeek-v4-flash-0731
PySR + Direct-Ask
0.78
0.78
0.00
0.00
0.00
0.00
0.00
0.00
GPT-5.6-sol
Codex
25.78
9.38
17.19
12.50
13.28
5.47
7.81
3.91
GLM-5.3-flash
Codex
10.16
3.12
0.78
0.00
2.34
2.34
3.91
1.56
Appendix
Table 12: Complete Full-set symbolic-recovery results by scientific domain.
Domain
Task family
Codex (GPT-5.6-sol)
Codex (GLM-5.3-flash)
P
M
P
M
Physics
Circular Orbit
21.88
15.63
9.38
3.12
Hall Response
25.00
12.50
3.12
6.25
Inclined Rolling
50.00
6.25
28.12
3.12
Terminal Settling
6.25
3.12
0.00
0.00
Chemistry
Acid–Base Buffer Relaxation
3.12
3.12
0.00
0.00
Appendix
Table 13: Full-set symbolic-recovery results by task family for the two standard discovery configurations evaluated on all 512 instances. Each family contains 32 instances.
Base model
Agent
Physics
Chemistry
Biology
Materials
ID
OOD
ID
OOD
ID
OOD
ID
OOD
No LLM
PySR
35.00
–
0.00
–
20.00
–
0.00
–
GPT-5.6-sol
Codex
55.00
45.00
50.00
50.00
60.00
45.00
55.00
35.00
GLM-5.3-flash
Codex (3-run avg.)
56.67
40.00
18.33
8.33
53.33
10.00
38.33
25.00
GLM-5.3-flash
Claude Code
60.00
50.00
20.00
15.00
35.00
10.00
35.00
25.00
DeepSeek-v4-flash-0731
Codex (3-run avg.)
45.00
36.67
18.33
11.67
46.67
16.67
38.33
23.33
Appendix
Table 14: Core-set numerical accuracy by scientific domain. Each entry is the percentage of tasks whose maximum relative error over the indicated split is at most 0.1 . A dash denotes an unavailable OOD result.
Base model
Agent
Core-set variants
Full-set variants
SA(P)
SA(M)
SA(P)
SA(M)
No LLM
PySR
0.00
–
0.00
–
DeepSeek-v4-flash-0731
PySR + Direct-Ask
–
0.00
–
0.00
GPT-5.6-sol
Codex
30.56
9.72
14.11
6.05
GLM-5.3-flash
Codex (3-run avg. / 1 run)
2.78
0.46
3.23
0.81
GLM-5.3-flash
Claude Code
4.17
1.39
–
–
Appendix
Table 15: Variant-only symbolic-recovery results after excluding Original instances. Codex with GLM-5.3-flash and Codex with DeepSeek-v4-flash-0731 use three runs only on the Core-set. A dash denotes a metric or set that is not evaluated for the corresponding configuration.
Base model
Agent
SA(P) by mutation count
SA(M) by mutation count
0
1
2
3
4
5
0
1
2
3
4
5
No LLM
PySR
0.00
0.00
0.00
0.00
0.00
0.00
–
–
–
–
–
–
DeepSeek-v4-flash-0731
PySR + Direct-Ask
–
–
–
–
–
–
12.50
0.00
0.00
0.00
0.00
0.00
GPT-5.6-sol
Codex
75.00
75.00
31.25
25.00
6.25
0.00
50.00
37.50
6.25
0.00
0.00
0.00
GLM-5.3-flash
Codex (3-run avg.)
37.50
8.33
0.00
4.17
0.00
0.00
25.00
2.08
0.00
0.00
0.00
0.00
GLM-5.3-flash
Claude Code
50.00
12.50
6.25
0.00
0.00
0.00
25.00
0.00
6.25
0.00
0.00
0.00
Appendix
Table 16: Core-set symbolic recovery stratified by the number of applied mutations. Mutation count zero denotes Original instances.
Base model
Agent
SA(P) by mutation count
SA(M) by mutation count
0
1
2
3
4
5
0
1
2
3
4
5
No LLM
PySR
6.25
0.00
0.00
0.00
0.00
0.00
–
–
–
–
–
–
DeepSeek-v4-flash-0731
PySR + Direct-Ask
–
–
–
–
–
–
6.25
0.00
0.00
0.00
0.00
0.00
GPT-5.6-sol
Codex
75.00
40.00
15.00
8.13
1.25
0.00
62.50
25.00
5.63
0.63
0.00
0.00
GLM-5.3-flash
Codex
37.50
13.75
2.50
0.63
0.00
0.00
31.25
5.00
0.00
0.00
0.00
0.00
Appendix
Table 17: Full-set symbolic recovery stratified by the number of applied mutations. Mutation count zero denotes Original instances.
Factor
Comparison or grouping
Metric
Rates (%)
p
Base model
GPT-5.6-sol vs GLM-5.3-flash (Codex)
SA(P)
35.00 / 6.25
<.001
SA(M)
13.75 / 2.92
.008
Acc0.1OOD
43.75 / 20.83
<.001
GPT-5.6-sol vs DeepSeek-v4-flash-0731 (Codex)
SA(P)
35.00 / 7.92
<.001
SA(M)
13.75 / 3.75
.008
Acc0.1OOD
43.75 / 22.08
<.001
Appendix
Table 18: Core-set statistical comparisons across base models, harnesses, and scientific domains. Rates follow the order of the groups named in each row. Codex rates for the two flash models are task-level averages over three runs except in the three-harness DeepSeek comparison, which uses the first Codex run alongside the single Claude Code and DeepSeek Harness runs.
When Isaac Newton discovered the law of gravitation, he did so through an iterative process of analyzing observed data such as planetary patterns, finding the underlying mechanisms by describing patterns in mathematical equations, and refining his theory against the Moon's orbit, revealing the startling insight that the same force governs both falling apples and orbiting planets. Would it be possible for AI agents to make similar discoveries? To measure this ability, we introduce EurekaBench, a cross-domain benchmark that tests AI agents' ability to conduct long-horizon experiments and discover mechanisms that explain observations. We evaluate these mechanisms by the scientific insights that can be derived from them. EurekaBench contains an expert-verified set of 26 long-horizon tasks across neuroscience, computer science, chemistry, astrophysics, geophysics, and plasma physics, with a total of 306 scientific insights that the discovered mechanisms are expected to support. Our evaluation framework tests three axes of scientific discovery: agents' ability to follow known scientific constraints, the predictive accuracy of the discovered mechanisms, and whether these mechanisms yield scientific insights or inform future research. Our results show that current AI agents often overly fixate on predictive accuracy optimization, surpassing human scientists, while falling substantially short in deriving scientific insights.
Jiayi Geng, Zhengxuan Wu, Kevin S. Chen +12
Carnegie Mellon University, Language Technologies Institute · Stanford University, Department of Computer Science · Yale University, Quantitative Biology Institute +8
Data-driven mechanistic hypotheses are essential to scientific discovery because they explain how underlying processes produce observed phenomena. AI agents and AI scientists increasingly support scientific data analysis. However, their ability to turn empirical findings into mechanistic hypotheses remains insufficiently examined. To address this gap, we introduce MechHypoBench, the first benchmark for evaluating whether AI agents and AI scientists can generate such hypotheses from empirical data. It combines paper-derived mechanisms from 14 scientific fields with real-world datasets containing 17.98 million records. The construction retains the observational complexity of empirical data while providing a specified underlying mechanism. Agents analyze the observations and propose open-form hypotheses. We develop an evaluation framework that assesses open-form mechanistic hypotheses through their consequences under withheld conditions. Experiments with general agents and AI scientists reveal a substantial gap between generated hypotheses and the underlying mechanisms.
Xiaxun Xie, Qingqing Long, Meng Xiao +4
Computer Network Information Center, Chinese Academy of Sciences · National University of Singapore · Sichuan University
We introduce NatureBench, a cross-discipline benchmark of 90 tasks distilled from peer-reviewed Nature-family publications, designed to evaluate whether AI coding agents can move beyond reproduction toward discovery on real scientific problems. NatureBench is built on NatureGym, an automated pipeline that constructs a standardized, per-task containerized environment from a source paper, addressing the environment-fragmentation problem that has limited the credibility of prior agent-on-research benchmarks. Evaluating ten frontier agent configurations under a strict web-search-disabled protocol, we find that the strongest model surpasses SOTA on only 17.8% of tasks under the g>0.1 criterion. Analysis of method pathways reveals that agents succeed primarily through methodological translation, converting scientific tasks into familiar supervised prediction problems, rather than through genuine scientific invention. Failures are dominated by wrong method choice and insufficient compute budget, not by task misunderstanding. We release the benchmark, the NatureGym pipeline, and a public leaderboard with maintainer-side reproduction. Code: https://github.com/FrontisAI/NatureBench