Large language model agents are being increasingly deployed as autonomous scientists, designing experiments and inferring mechanistic world models with minimal human oversight. Yet identifiability is often overlooked: when a plateau is reached, the agent needs to know whether it is not yet capable enough or the model simply is not identifiable from the data, in which case no amount of further experimentation of the same kind can help. We propose the Identifiability-Driven Experimental Agent (LLM-IDEA) for closed-loop discovery with an identifiability engine that returns a three-way plateau verdict: capability limit, resolvable within the design class, or certified exhausted. On ODEBench, 60 of the 62 systems with free constants are identifiable at round 0; the RC circuit is certified exhausted for every experiment that protocol can run, and a harvesting model is resolvable by one added initial condition. The identifiability engine reproduces known verdicts on Lotka-Volterra, Van der Pol, Lorenz, and a pharmacokinetic model, where it recommends the intravenous arm pharmacologists use, and it ranks the depth scorer of our own benchmark last among four observation designs. On the DiscoverPhysics benchmark, it finds two public worlds whose explanation rubric rewards a distinction no legal experiment can make, and every model there with accurate trajectories failed the explanation grade (15 of 15, against 5 of 9 in identifiable worlds, p = 0.012). On the Alien Universe, a two-body testbed we propose in which a force law switches between a provably non-identifiable and an identifiable protocol, LLM-IDEA on the identifiable protocol reaches discovery depth at least three on 8/8 seeds versus 1/8 without it. An autonomous discovery agent can thus compute, rather than guess, whether a plateau calls for more search, a better experiment of the same kind, or a different kind of experiment.
Figures & tables
Figure 1: The two-body problem in the Alien Universe. The bodies have masses and charges m1,s1 and m2,s2 .
System
Design
n_params
rank
identifiable? (initial conditions known)
cond. number
Lotka-Volterra (4-param)
full-state (x, y)
4
4
yes
21
Lotka-Volterra (4-param)
prey-only (x)
4
4
yes
169
Van der Pol
full-state (x, y)
1
1
yes
1
Van der Pol
position-only (x)
1
1
yes
1
Lorenz
full-state (x, y, z)
3
3
yes
23
Lorenz
x-only (x)
3
3
yes
25
Table 1: The diagnostic on textbook ODE systems with initial conditions known: rank and condition number of the sensitivity matrix at the nominal parameters.
Figure 2: Observability lift on the Alien Universe (EXP-076): the made-observable treatment reaches discovery depth ≥3 on 8/8 seeds against 1/8 for the single-configuration control. The treatment also carries the fitting engine; the factorial (EXP-089) attributes the cross-configuration layer to the family and the correction-exponent layer to the engine.
Figure 3: The cross-configuration sum gate (L3) on the second Alien Universe regime holds at 7/8 under a corrupted in-loop fitting engine (EXP-080), 7/8 with the probe battery corrected (EXP-081), and 6/8, its pre-registered bar, with both agent-visible channels corrected under clean weather (EXP-083b). The band-gated layers fall across the same runs. A fourth, outage-confounded run (EXP-083) is reported in the text and excluded here.
Figure 4: The identifiability loop closed by a live agent decision, both ways (EXP-078). Left: the director requests a multi-charge family, the deterministic ranker selects the two-distinct-product design, and the re-diagnosis returns full rank (SOLVED). Right: the director requests a new initial condition, no candidate restores rank, and the loop reports EXHAUSTED. With the certificate of Section 3.5 the right-hand branch terminates at round 0 by proof.
Figure 5: The engine diagnoses its own benchmark. With the correction exponent as a free parameter, the design ranker places the depth scorer’s own observation design last of four; its indifference band (orange) does not reach the gate band, while the top-ranked design (blue) excludes r−3 .
Figure 6: Terminal structural accuracy tracks the engine-feedback dose rather than the architecture as such: near-zero deliveries (N=15: two across the eight Council seeds) yield zero in-regime terminals, one delivery (N=40) about half, many (EXP-075) nearly all.
Figure 7: Selection criteria on identical candidates (EXP-262). For each ODEBench task, the recovery error of a criterion’s chosen experiment divided by the error of the best experiment available on that task (1 = chose the best), median over the 61 tasks per protocol with at least two full-rank candidates, with 95% bootstrap intervals. A descriptive view; the registered statistics are in the text.
world
verdict
what decides it
best expl.
TG & EF / TG
fractional
NONID
“fractional operator” ≡ power law (certified)
0.48
11/11
circle
NONID
the same equivalence at α=0.75
0.70
4/4
extra_dimensions
RESOLVABLE
Rc identifiable by 5 of 65 legal designs
0.54
6/6
gravity
RESOLVABLE
an alternative decay law separated only by design
1.00
3/7
yukawa
RESOLVABLE
screening length identifiable only by design
1.00
5/9
hubble
ID
—
1.00
2/4
Table 2: Identifiability audit of the DiscoverPhysics public worlds (EXP-263). Verdicts from the benchmark’s own simulator under the agent’s interface; outcomes from the pinned leaderboard (13 models, means over 5 seeds). TG: cells with geometric position error ≤0.1 ; TG & EF: of those, cells whose explanation grade is below 0.9.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
prompt
SHA-256 (first 12)
calls
runs
first to last date
aristotle_hypothesis 1.0
806417c6a5ab
875
93
2026-05-30 to 2026-09-09
diogenes_hypothesis 1.0
901289c03123
855
95
2026-05-30 to 2026-09-09
plato_hypothesis 1.0
8ed581c49a36
852
101
2026-05-30 to 2026-09-09
heraclitus_hypothesis 1.0
b364831eca2e
742
101
2026-05-30 to 2026-09-09
heraclitus_contrastive_distillation 1.0
6ba953a86de9
699
99
2026-05-30 to 2026-09-09
parmenides_hypothesis 1.0
f61781f63774
671
94
2026-05-30 to 2026-09-09
Appendix
Table 5: Model calls by system prompt in the run database, 2026-05-30 to 2026-09-09.
A primary goal of science is to learn mechanistic world models from limited experimental data, both to explain observations and to predict novel interventions. We introduce the Model Discovery Agent (MDA), which combines LLM proposals for M-open model discovery, experiment design based on Value of Information, and approximate Bayesian inference over model structures, parameters, and stochastic latent trajectories. We apply MDA to learn symbolic reaction rate laws for ChemBench \citep{kabra2026autoscilab}, partially observed ODE models for GlucoseBench \citep{xie2018simglucose,kovatchev2009insilico}, and partially observed SDE models for a new stochastic single-neuron simulator we create. In the appendix, we also show results on various other domains from BoxingGym \citep{gandhi2025boxinggym}. We show that MDA has improved sample efficiency compared to various baseline methods, and the learned models are good predictors but also provide interpretable abstractions of each domain.
Kevin Murphy
Dept. Computer Science Univ. British Columbia, Canada.
Frontier LLMs now perform strongly across a wide range of physics evaluations, but it is hard to disentangle genuine reasoning from recall of established science. We introduce DiscoverPhysics, an interactive benchmark that asks a LLM agent to discover the laws of motion of a simulated world whose physics deliberately deviates from our own. We construct 22 worlds governed by, among others, screened and fractional-power gravity, multi-species couplings, hidden dark-matter-like particles, non-coordinate-free physics, and time-varying interactions. Each world is generated on demand by an N-body simulator, for which the agent proposes several rounds of experiments, observes raw trajectory data, and ultimately submits both a natural-language explanation of the world's physics and a Python implementation of the inferred law. Because solving a world requires the agent to design informative experiments and revise its hypotheses, the benchmark probes long-horizon reasoning over an experimental history. We evaluate submissions along two complementary axes: trajectory MSE on held-out particles and an LLM-judged explanation score following an expert-written rubric assessing conceptual understanding of each world. Across eleven frontier models, we find that the strongest agents pass only half of the worlds and consistently fail on those where latent structure must be uncovered. Open-source models lag substantially behind commercial models, both in their ability to design informative experiments and in extracting conclusions from the data. We further find that good predictive accuracy does not guarantee high explanation quality and that conceptual understanding depends on hypothesis refinement through well-chosen experiments.
Matt L. Wiemann, Lindsay M. Smith, Peter Melchior +4
Princeton University · Work done during internship at NYU/Polymathic AI. · Boston University +3
We propose agentic automata learning to evaluate the extent to which tool-calling LLM agents can uncover hidden environments through interaction, a capability increasingly required in agentic tasks (e.g., reproducing an executable without access to its source code by interacting with it). In our setup, an agent should uncover a hidden deterministic finite automaton (DFA) by interacting with an oracle through (1) membership queries ("Does this string belong to the target language?'') and (2) equivalence queries ("Is this the target DFA?''). Agentic automata learning yields a scalable testbed with controlled task complexity, measurable interaction efficiency, and strong algorithms to compare against from the classic automata-learning literature. Evaluating state-of-the-art LLMs with a multi-turn agent scaffold, we find that while they are able to recover simple DFAs, their performance drops sharply as DFA size increases. Results improve with a more elaborate ReAct-style state-tracking scaffold, yet strong models still struggle with complex instances. Trajectory analyses reveal recurring failures in query planning, evidence integration, and hypothesis construction. These failures occur even though the models mention classic automata learning algorithms in their reasoning, and can implement and execute them when given access to a coding environment. Our results suggest that for state-of-the-art LLMs, identifying a solution to the problem is insufficient, and reliably executing the resulting plan still poses a distinct challenge.
Reef Menaged, Gili Lior, Shauli Ravfogel +2
The Hebrew University of Jerusalem · New York University · Google Research