Organizations: Lattice Lab, Toyota Motor Corporation, Japan · Department of Computer Science, University of Tokyo, Japan · EPIC Lab, College of Computing & Data Science, Nanyang Technological University, Singapore · RIKEN Center for Advanced Intelligence Project, Japan · Rational Intelligence Lab, CISPA, Helmholtz Center for Information Security, Germany
Autonomous scientific discovery with LLMs requires generating and testing hypotheses adaptively as evidence accumulates while maintaining statistical validity. Existing anytime-valid methods can handle data-dependent hypotheses, but open-ended discovery poses a deeper challenge: the best discovered hypothesis may still be the best of a bad lot, with better explanations yet undiscovered, while even background knowledge such as physical laws may require revision in light of new findings. In response, we formalize the problem as Abductive Autonomous Scientific Discovery (AASD) using possibility theory. We introduce abductive utility, a computable measure of discovery progress, and possibility frontier search, the first algorithm for AASD, which maintains anytime validity and achieves ε-optimal abductive utility asymptotically under suitable conditions. Experiments on synthetic and real-world scientific-discovery tasks show strong performance.
Figures & tables
Figure 1: Our possibilistic reasoning approach enables open-ended scientific discovery with anytime-valid guarantees under closed-/open-world uncertainty.
Figure 2: Two-stage assessment. (a) Persistent evidential possibility classifies claims as contradicted, open, or certified. (b) Guaranteed background possibility ranks certified claims by their minimum statewise compatibility. Smaller semantic claims are preferred when scores tie.
Figure 3: Possibility frontier search collects open hypotheses that can improve the current best Ut or ∣Ht∣ , then selects experiments to falsify or certify them.
Figure 4: Synthetic discovery setup: The environment Wθ⋆ differs from the background model WK (dashed). At t=0 (left), none of the initial hypotheses H∈H0 contains θ⋆ . For t>0 (right), the hypothesis space is expanded to discover a hypothesis containing θ⋆ .
Figure 5: Precision (upper) and abductive utility (lower) under complete (left) incomplete H0 . Ours returns a certified hypothesis at t=2 or 3 with empirical precision 1.0 , and its abductive utility subsequently improves and converges to U⋆ . Shaded regions denote 95% confidence intervals.
Figure 6: Ablation study: (a) experimental design and (b) stopping condition for Ut and (c) Ut−Ut . Ours reaches U⋆ faster than random actions, and larger tolerance ε stops earlier following the gap.
Figure 7: MLE and MAP remain confined to the initial set H0 . AutoDiscovery expands the search toward hypotheses that are most surprising relative to the prior, whereas our method rapidly identifies a certified hypothesis H∈Ct requiring minimal revision of the background knowledge. Orange box denotes H≤T .
Figure 8: Natural-language discovery results. Left: HMS. Middle: abductive utility. Right: cardinality ∣Ht∣ . Shaded regions indicate ± standard error.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 9: As evidence accumulates, hypotheses are certified only when sufficiently supported, with the false-certification rate remaining below the target level δ=0.1 . Shaded regions show 95% confidence intervals
Figure 10: Validity under post-hoc hypothesis selection. The certification criteria control both error types while retaining competitive power. The necessity requirement distinguishes compatibility from sufficient evidential support.
Figure 11: Empirical error (top) and power (bottom) under post-hoc hypothesis selection for different mixture scales τ . Our method maintains error control at δ=0.05 with power comparable to naive reuse across varying τ .
Open-ended scientific discovery with large language models (LLMs) increasingly operates as a long-horizon loop of hypothesis search and verification, where a reward signal guides which hypotheses to test next. A notable recent example is AutoDiscovery, which uses "Bayesian surprise" - the belief shift an LLM undergoes after observing evidence for a hypothesis - as both a discovery metric and a reward for search. We first observe that AutoDiscovery treats surprisal as a static quantity, while surprisal in human reasoning is non-stationary - it is defined relative to beliefs that evolve with experience, a prerequisite for continual scientific discovery. We address this mismatch with evidence-informed LLM beliefs: priors updated with evidence from previous hypotheses to compute non-stationary surprisal for new hypotheses. We compare in-context belief-updating mechanisms and find that embedding-based retrieval-augmented generation over prior discoveries best anticipates eventual posteriors, identifying 37.5% of static surprisals as spurious. We then modify search to avoid these spurious rewards and prioritize hypotheses that remain surprising under non-stationary beliefs. Concretely, we introduce two complementary changes to the original search procedure: belief-update filtering and diversity maximization. Across five discovery domains, our method increases accumulated non-stationary surprisal by 30.62% on average compared to the original search procedure, demonstrating that continual scientific discovery with LLMs requires not only better belief measurement but also search procedures that avoid redundancy and encourage diversity.
Dhruv Agarwal, Reece Adamson, Andrew McCallum +3
University of Massachusetts Amherst · Allen Institute for AI
Autonomous scientific discovery systems increasingly use large language models (LLMs) to propose new hypotheses, but many such systems condition primarily on experimental memory: archives of high-scoring candidates or heuristic summaries of recent trials. We argue that discovery agents should instead maintain explicit, uncertainty-aware beliefs about hypothesis quality. We introduce BayesEvolve, a belief-guided discovery framework that converts experimental evidence into a predictive belief state and uses this belief to guide future experimentation. As a controlled testbed for belief-guided discovery, we evaluate BayesEvolve on shifted BBOB-style black-box optimization tasks, leaving program and laboratory discovery domains to future work. BayesEvolve improves sample efficiency over memory- and archive-guided LLM baselines under a fixed evaluation budget. We further show that the belief state is predictive on held-out candidate pools, that controlled decision-rule ablations favor belief-guided selection with an annealed uncertainty bonus, and that BayesEvolve exhibits productive late-stage concentration rather than unfocused exploration.
Xuening Wu, Shan Yu, Qianya Xu +1
Pfizer, Shanghai, China · Independent Researcher, Hangzhou, China · University of California San Diego, La Jolla, CA, USA +1
Scientific discovery is a closed-loop process in which hypotheses guide data acquisition and observations refine the hypothesis space. Yet most approaches reduce discovery to supervised learning over fixed datasets, where limited observations can support multiple plausible mechanisms that fit locally but fail to generalize. Thus, the key challenge is selecting informative observations to resolve uncertainty, shifting the focus from static inference to adaptive data acquisition. To address this, we propose LLM-AutoSciLab, a closed-loop framework that couples hypothesis generation with hypothesis-conditioned experiment selection and mechanism refinement. Rather than fitting models to passively collected data, LLM-AutoSciLab iteratively proposes plausible hypotheses, selects informative experiments to distinguish or refine them, and updates its state using the resulting evidence. To evaluate dynamic, closed-loop scientific discovery with active data acquisition, we introduce ActiveSciBench, comprising two datasets: ActiveSciBench-Chem with 57 enzyme-kinetics tasks and ActiveSciBench-GRN with 45 gene-regulatory-network tasks. These datasets model discovery as a budget-constrained process requiring adaptive experiment design, variable selection, and recovery of true mechanisms. Across NewtonBench, ActiveSciBench-Chem, and ActiveSciBench-GRN, LLM-AutoSciLab outperforms prior methods, achieving 67.6% and 35.1% symbolic accuracy on NewtonBench and ActiveSciBench-Chem, respectively, and 31.1% exact graph recovery on ActiveSciBench-GRN. Moreover, hypothesis-guided experimentation is 2-5x more sample-efficient than the strongest competing baselines. Code and data are available at: https://github.com/scientific-discovery/LLM-AutoSciLab