cs.AIOct 8, 2026

LLM-IDEA: Identifiability-Driven Experimental Agent for Autonomous Discovery of Mechanistic World Models

Authors: Surya Shetty, Ulisses Braga-Neto

Organizations: Texas A&M University · Polara Labs Inc.

Abstract

Large language model agents are being increasingly deployed as autonomous scientists, designing experiments and inferring mechanistic world models with minimal human oversight. Yet identifiability is often overlooked: when a plateau is reached, the agent needs to know whether it is not yet capable enough or the model simply is not identifiable from the data, in which case no amount of further experimentation of the same kind can help. We propose the Identifiability-Driven Experimental Agent (LLM-IDEA) for closed-loop discovery with an identifiability engine that returns a three-way plateau verdict: capability limit, resolvable within the design class, or certified exhausted. On ODEBench, 60 of the 62 systems with free constants are identifiable at round 0; the RC circuit is certified exhausted for every experiment that protocol can run, and a harvesting model is resolvable by one added initial condition. The identifiability engine reproduces known verdicts on Lotka-Volterra, Van der Pol, Lorenz, and a pharmacokinetic model, where it recommends the intravenous arm pharmacologists use, and it ranks the depth scorer of our own benchmark last among four observation designs. On the DiscoverPhysics benchmark, it finds two public worlds whose explanation rubric rewards a distinction no legal experiment can make, and every model there with accurate trajectories failed the explanation grade (15 of 15, against 5 of 9 in identifiable worlds, p = 0.012). On the Alien Universe, a two-body testbed we propose in which a force law switches between a provably non-identifiable and an identifiable protocol, LLM-IDEA on the identifiable protocol reaches discovery depth at least three on 8/8 seeds versus 1/8 without it. An autonomous discovery agent can thus compute, rather than guess, whether a plateau calls for more search, a better experiment of the same kind, or a different kind of experiment.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Aug 10, 2026cs.AI

Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models

A primary goal of science is to learn mechanistic world models from limited experimental data, both to explain observations and to predict novel interventions. We introduce the Model Discovery Agent (MDA), which combines LLM proposals for M\mathcal M-open model discovery, experiment design based on Value of Information, and approximate Bayesian inference over model structures, parameters, and stochastic latent trajectories. We apply MDA to learn symbolic reaction rate laws for ChemBench \citep{kabra2026autoscilab}, partially observed ODE models for GlucoseBench \citep{xie2018simglucose,kovatchev2009insilico}, and partially observed SDE models for a new stochastic single-neuron simulator we create. In the appendix, we also show results on various other domains from BoxingGym \citep{gandhi2025boxinggym}. We show that MDA has improved sample efficiency compared to various baseline methods, and the learned models are good predictors but also provide interpretable abstractions of each domain.
May 25, 2026stat.ML

DiscoverPhysics: Benchmarking LLMs for Out-of-the-Box Scientific Thinking

Frontier LLMs now perform strongly across a wide range of physics evaluations, but it is hard to disentangle genuine reasoning from recall of established science. We introduce DiscoverPhysics, an interactive benchmark that asks a LLM agent to discover the laws of motion of a simulated world whose physics deliberately deviates from our own. We construct 22 worlds governed by, among others, screened and fractional-power gravity, multi-species couplings, hidden dark-matter-like particles, non-coordinate-free physics, and time-varying interactions. Each world is generated on demand by an N-body simulator, for which the agent proposes several rounds of experiments, observes raw trajectory data, and ultimately submits both a natural-language explanation of the world's physics and a Python implementation of the inferred law. Because solving a world requires the agent to design informative experiments and revise its hypotheses, the benchmark probes long-horizon reasoning over an experimental history. We evaluate submissions along two complementary axes: trajectory MSE on held-out particles and an LLM-judged explanation score following an expert-written rubric assessing conceptual understanding of each world. Across eleven frontier models, we find that the strongest agents pass only half of the worlds and consistently fail on those where latent structure must be uncovered. Open-source models lag substantially behind commercial models, both in their ability to design informative experiments and in extracting conclusions from the data. We further find that good predictive accuracy does not guarantee high explanation quality and that conceptual understanding depends on hypothesis refinement through well-chosen experiments.
Jun 15, 2026cs.CL

Can Agents Infer Environment from Interaction? Evidence from Agentic Automata Learning

We propose agentic automata learning to evaluate the extent to which tool-calling LLM agents can uncover hidden environments through interaction, a capability increasingly required in agentic tasks (e.g., reproducing an executable without access to its source code by interacting with it). In our setup, an agent should uncover a hidden deterministic finite automaton (DFA) by interacting with an oracle through (1) membership queries ("Does this string belong to the target language?'') and (2) equivalence queries ("Is this the target DFA?''). Agentic automata learning yields a scalable testbed with controlled task complexity, measurable interaction efficiency, and strong algorithms to compare against from the classic automata-learning literature. Evaluating state-of-the-art LLMs with a multi-turn agent scaffold, we find that while they are able to recover simple DFAs, their performance drops sharply as DFA size increases. Results improve with a more elaborate ReAct-style state-tracking scaffold, yet strong models still struggle with complex instances. Trajectory analyses reveal recurring failures in query planning, evidence integration, and hypothesis construction. These failures occur even though the models mention classic automata learning algorithms in their reasoning, and can implement and execute them when given access to a coding environment. Our results suggest that for state-of-the-art LLMs, identifying a solution to the problem is insufficient, and reliably executing the resulting plan still poses a distinct challenge.