Organizations: Lattice Lab, Toyota Motor Corporation, Japan · Department of Computer Science, University of Tokyo, Japan · EPIC Lab, College of Computing & Data Science, Nanyang Technological University, Singapore · RIKEN Center for Advanced Intelligence Project, Japan · Rational Intelligence Lab, CISPA, Helmholtz Center for Information Security, Germany
Autonomous scientific discovery with LLMs requires generating and testing hypotheses adaptively as evidence accumulates while maintaining statistical validity. Existing anytime-valid methods can handle data-dependent hypotheses, but open-ended discovery poses a deeper challenge: the best discovered hypothesis may still be the best of a bad lot, with better explanations yet undiscovered, while even background knowledge such as physical laws may require revision in light of new findings. In response, we formalize the problem as Abductive Autonomous Scientific Discovery (AASD) using possibility theory. We introduce abductive utility, a computable measure of discovery progress, and possibility frontier search, the first algorithm for AASD, which maintains anytime validity and achieves ε-optimal abductive utility asymptotically under suitable conditions. Experiments on synthetic and real-world scientific-discovery tasks show strong performance.
Figures & tables
Figure 1: Our possibilistic reasoning approach enables open-ended scientific discovery with anytime-valid guarantees under closed-/open-world uncertainty.
Figure 2: Two-stage assessment. (a) Persistent evidential possibility classifies claims as contradicted, open, or certified. (b) Guaranteed background possibility ranks certified claims by their minimum statewise compatibility. Smaller semantic claims are preferred when scores tie.
Figure 3: Possibility frontier search collects open hypotheses that can improve the current best Ut or ∣Ht∣ , then selects experiments to falsify or certify them.
Figure 4: Synthetic discovery setup: The environment Wθ⋆ differs from the background model WK (dashed). At t=0 (left), none of the initial hypotheses H∈H0 contains θ⋆ . For t>0 (right), the hypothesis space is expanded to discover a hypothesis containing θ⋆ .
Figure 5: Precision (upper) and abductive utility (lower) under complete (left) incomplete H0 . Ours returns a certified hypothesis at t=2 or 3 with empirical precision 1.0 , and its abductive utility subsequently improves and converges to U⋆ . Shaded regions denote 95% confidence intervals.
Figure 6: Ablation study: (a) experimental design and (b) stopping condition for Ut and (c) Ut−Ut . Ours reaches U⋆ faster than random actions, and larger tolerance ε stops earlier following the gap.
Figure 7: MLE and MAP remain confined to the initial set H0 . AutoDiscovery expands the search toward hypotheses that are most surprising relative to the prior, whereas our method rapidly identifies a certified hypothesis H∈Ct requiring minimal revision of the background knowledge. Orange box denotes H≤T .
Figure 8: Natural-language discovery results. Left: HMS. Middle: abductive utility. Right: cardinality ∣Ht∣ . Shaded regions indicate ± standard error.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 9: As evidence accumulates, hypotheses are certified only when sufficiently supported, with the false-certification rate remaining below the target level δ=0.1 . Shaded regions show 95% confidence intervals
Figure 10: Validity under post-hoc hypothesis selection. The certification criteria control both error types while retaining competitive power. The necessity requirement distinguishes compatibility from sufficient evidential support.
Figure 11: Empirical error (top) and power (bottom) under post-hoc hypothesis selection for different mixture scales τ . Our method maintains error control at δ=0.05 with power comparable to naive reuse across varying τ .