Organizations: Department of Physics, Princeton University, Princeton, NJ 08540, USA · Center for Gravitational Physics, University of Texas at Austin, Austin, TX 78712, USA · School of Natural Sciences, Institute for Advanced Study, 1 Einstein Drive, Princeton, NJ 08540, USA
We study how AI agents discover and validate scientific formulas using a controlled case study of the hydrotope, a recently discovered geometric formula that combines the different polynomial pieces of nonlinear surface-wave scattering into one global expression. This problem is deceptively difficult: simple formulas can hold within individual frequency regions, but the global result must identify their boundaries and combine exponentially many potentially active terms. We reconstruct how the formula was originally discovered through human--agent collaboration and analyze 18 single-prompt rediscovery runs under no hint and two forms of human guidance: a false hint representing an incorrect prior and a true hint representing domain-informed insight. Only four recover the formula across all kinematic chambers (i.e., regions in which a single polynomial form applies), while most unsuccessful runs find correct chamber polynomials but fail to combine them or test their full domain. Conventional and LLM-assisted symbolic regression and standard machine-learning regressors likewise fail to recover the global formula in our experiments. Guided by these failure modes, we test a PI+two-student workflow in which a coordinating lead agent assigns complementary analytic and numerical tasks to two research agents and independently evaluates their results. The PI+two-student team successfully rediscovers the complete hydrotope formula, while the same workflow applied to the harder three negative wavenumber problem discovers a new independent verified analytic expression for the six-point amplitude A6.
Figures & tables
Figure 1: This paper asks how AI agents can infer an analytic scattering formula from numerical values provided at different input frequencies. (a) Berends–Giele (BG) recursion numerically calculates the exact answer at any allowed point but does not reveal the analytic formula by itself. A simple formula can work in one frequency chamber and fail in another, so the task is to find one expression valid in every chamber and for an arbitrary number of waves. (b) Fifteen of 18 single-agent runs find a correct formula in at least one chamber, but only four combine the chamber formulas and test the result in new chambers. (c) Two students try different approaches and write their formulas, tests, and failures in a shared record. A PI combines their work and checks it on new examples. Neither of six single-agent instances finds the complete all- n hydrotope formula, whereas both PI + student agentic teams do. The same team also produces a new six-point three-minus formula that passes 140 exact tests in 58 chambers.
Figure 2: Geometric interpretation of the two-minus hydrotope, the known formula for the sector in which two waves have negative wavenumber. A chamber-dependent scattering amplitude can be represented by a slice of a box whose edge lengths are set by the wave frequencies. Crossing a chamber boundary switches a term (x)+=max(x,0) in Eq. 3 on or off, so the sum adds the active contributions and subtracts their overlaps. At n=5 , a plane cuts a three-dimensional box into four allowed polygon types, with the conservation laws excluding the central hexagon; at n=6 , the analogous cut is in four dimensions and gives different three-dimensional solids in different chambers. One formula built from max(x,0) therefore describes every chamber-dependent cross-section.
package
equation guidance
recommended tests
false hint
one ratio of polynomials for all frequencies; no chamber split
comparable frequencies; avoid separated scales and chamber boundaries
true hint
different polynomials in different chambers, with fixed scaling
find the chamber boundaries and test on both sides
no hint
no proposed equation class
test n=4,5,6,7 , including widely separated frequency scales
Table 1: The repeated runs vary both the suggested type of equation and the recommended tests. Appendix A reproduces the complete prompt descriptions.
Figure 3: As a benchmark to compare the performance of the agentic method, we test symbolic regression and classical ML methods (see section 3.2 ) to see if they can learn the relationship predicted by the exact (i.e., zero error) five-point hydrotope formula discovered by the multi-agent system (see Eq. ( 5 )). We use the same training dataset used in the agentic runs: wave frequencies as input and corresponding output amplitude from the recursion (BG) relation. Bars show median held-out mean absolute error (MAE), open markers show individual runs, and the triangle marks an off-scale run retained in the benchmark table. Operon has the lowest fitted median MAE (1.36), but no conventional, LLM-assisted, or prediction method recovers exact formula.
condition
complete
partial sum
one chamber
hint rejected
incorrect
false hint
0
1
3
1
1
true hint
4
1
0
0
1
no hint
0
2
4
0
0
total
4
4
7
1
2
Table 2: Each of the three prompts is tested with six agent configurations. “Complete” means correct for every chamber and every n . “Partial sum” means that the agent finds the alternating-sum structure but misses a term, an on/off condition, or a test case. “One chamber” means that the final all- n formula applies only in one chamber. All four complete results use the true hint that says a different polynomial applies in each chamber.
Figure 4: The rows show eight key steps in order of completeness of the solution. These range from sampling exact values from the BG recursion (K1) to testing and verifying the output analytic formula holds in new chambers (K8). Dark, light, and pale cells mean complete, partial, and not reached. The result column uses C for a complete formula, P for a partial sum, O for a one-chamber formula, R for a rejected false hint, and I for an incorrect result. The final column gives the mean absolute error (MAE) of the run’s n=8 formula on a separate set of 250 test points; a dash means that the final answer contains no formula that can be evaluated at n=8 . The largest difference among runs appears at K6–K7, where they must find every boundary and combine the chamber terms.
Figure 5: Four representative runs without human hints all recover the one-term formula in the principal chamber, where no chamber-boundary correction has yet switched on, and find counterexamples outside that chamber. Codex 5.5 and Fugu then produce short but incomplete subset sums. The Claude runs also attempt one formula for all chambers: Claude max’s sorting-based formula misses six of 233 tests, while Claude ultra cannot combine its separate chamber expressions. Both then report only the principal-chamber formula. The bottom expression omits the common factor i2n−1g3−nω1ω2 . Thus, these runs diagnose chamber dependence but fail to synthesize and verify one globally valid formula.
Figure 6: Workflow of our PI+two student (S1, S2) architectures with different base models. The two students construct formulas, while a PI assigns work and checks the result independently. The top row shows the six no-hint single-agent results for comparison, and each box records one team round. The Opus team proposes the complete formula in round 1 and verifies it in round 2; the Codex 5.5 team uses its first failure to divide round 2 between the special n=4 limit and the complete formula, then the PI independently tests both results and accepts them in round 3. Thus both coordinated teams succeed where none of the six no-hint single agents does, although the runs are not cost-matched and are too few to measure success rates. A written record, two simultaneous searches, and separate final checks help the teams complete the formula.
Figure 7: After rediscovering the known hydrotope formula for two negative wavenumbers, the agents tackle the harder sector with three negative wavenumbers and propose a new analytic expression for the six-point amplitude A6 intended to cover every chamber (see Eq. ( 18 )). In this eight-round run, rounds 1–4 determine the formula’s ratio structure: a common frequency-dependent denominator and a polynomial numerator carrying the remaining frequency dependence. A failed test in round 5 reveals that the numerator needs corrections involving pairs of chamber boundaries. Rounds 6–7 add those terms, and independently written BG recursion code finds exact agreement on 140 tests in 58 chambers. Round 8 studies n=7 , whose explicit numerator remains unknown. The failed cross-chamber test drives the correction that completes the supported A6 expression, but the finite tests do not provide an analytic proof and the numerator for n≥7 remains open.
Figure 8: AI agents often find formulas that work in one frequency chamber but fail to combine them into a global expression. We compare four architectures on the harder discovery problem introduced in Figure 7 : deriving an explicit formula for the new six-point amplitude A6 in the three-negative-wavenumber sector and testing it in unseen chambers. The single-agent run returns a tree expansion that hides the chamber changes, while the PI + one-student team finds separate chamber formulas that fail in a new chamber. Both PI + two-student setups, one with an additional verifier, complete and independently check the global expression. Because each architecture was run once and the runs are not cost-matched, these outcomes identify failure modes rather than success rates. In these runs, only teams with two complementary student roles carry both formula construction and cross-chamber testing through to a complete A6 formula.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 9: This figure repeats the comparison in Figure 3 for the n=7 case. The formula increases from eight max(x,0) terms at n=5 to 32 such terms. Bars show median held-out MAE from raw frequencies and open markers show runs; gplearn has the lowest fitted median, 373.49, but slightly negative held-out R2 and does not recover the target equation. The research team’s exact formula remains at zero, while every fitted method misses terms that switch on in some chambers. A few unusually large values in the fixed split also make absolute errors larger than at n=5 . Increasing the number of terms that can switch on or off makes exact recovery more difficult under the tested budgets.
Figure 10: Additional methods test whether searches that allow a wider range of equations, including separate equations for different regions, recover the five-point formula. Bars show median held-out MAE and open markers show individual runs for alternative symbolic-regression systems, combinations of region-finding and equation search, archived LLM-assisted variants, SciExplorer, and additional ML methods. PS-Tree improves approximate prediction by learning seven regions with separate symbolic expressions, but it does not recover the compact hydrotope formula; the region-first Operon hybrid likewise misses the true chamber boundaries. SciExplorer misses the max(x,0) terms. Allowing separate formulas in different regions can improve predictive accuracy, but none of these searches recovers the complete formula found by our research team.
prompt condition
agent configuration
first page
true hint
Codex 5.5 xhigh
F.1
no hint
Claude Opus 4.8 max
F.2
no hint
Claude Opus 4.8 ultra
F.3
no hint
Codex 5.5 xhigh
F.4
no hint
Fugu ultra
F.5
Appendix
Table 3: Index of the five complete run records included below.
Deriving governing equations from empirical observations is a longstanding challenge in science. Although artificial intelligence (AI) has demonstrated substantial capabilities in function approximation, the discovery of explainable and extrapolatable equations remains a fundamental limitation of modern AI, posing a central bottleneck for AI-driven scientific discovery. Here, we present machine collective intelligence, a unified paradigm that integrates two fundamental yet distinct traditions in computational intelligence--symbolism and metaheuristics--to enable autonomous and evolutionary discovery of governing equations. It orchestrates multiple reasoning agents to evolve their symbolic hypotheses through coordinated generation, evaluation, critique, and consolidation, enabling scientific discovery beyond single-agent inference. Across scientific systems governed by deterministic, stochastic, or previously uncharacterized dynamics, machine collective intelligence autonomously recovered the underlying governing equations without relying on hand-crafted domain knowledge. Furthermore, the resulting equations reduced extrapolation error by up to six orders of magnitude relative to deep neural networks, while condensing 0.5-1 million model parameters into just 5-40 interpretable parameters. This study marks an important shift in AI toward the autonomous discovery of principled scientific equations.
Gyoung S. Na, Chanyoung Park
Korea Research Institute of Chemical Technology (KRICT) · Korea Advanced Institute of Science and Technology (KAIST)
Scientific discovery is often a collective process: researchers share partial results, inspect failed attempts, and build on each other's ideas over long time horizons. Recent AI systems have shown that language-model-based agents can make meaningful progress on open scientific problems, but most existing systems operate in isolation. In this paper, we present EinsteinArena, an agent-native platform for open distributed research and discovery. EinsteinArena provides agents with a live set of open problems, each with a solid verifier, public leaderboard, and problem-specific discussion forum where agents can ask questions and share insights. We focus on mathematical tasks that have garnered substantial research interest, where progress can be measured unambiguously. As of May 2026, agents on EinsteinArena have discovered 12 new state-of-the-art results better than any previous human or AI solutions. One notable example is the kissing number problem in dimension 11, where the platform improved the best known lower bound from 593 to 604. This advance did not come from a single agent or isolated run. Rather it arose through a sequence of submissions, public discussion, verifier refinement, and subsequent agent-to-agent borrowing of ideas. These results provide evidence that decentralized scientific discovery can emerge from open interaction among autonomous agents in the wild, demonstrating a new paradigm for collective AI-driven research.
Scientific law discovery has long been central to scientific progress, proceeding through iterative cycles of generating hypotheses, testing them against empirical evidence, and refining them under scientific constraints. As large language models (LLMs) become increasingly involved in scientific research, whether they can discover scientific laws and how to evaluate this ability remain open questions. A central evaluation challenge is to move beyond familiar published equations while keeping discovery tasks grounded in scientific data and constraints. We introduce SciLaws-Bench, a curated collection of scientific task packages grounded in the source literature, each linking a scientific problem, supporting data, published reference equations, and scientific-validity rubrics. Through agent-assisted curation and human verification, we assemble 118 problems spanning six disciplines, drawing on 381 papers, 291 candidate laws, and roughly 8M data points. Each problem supports two complementary evaluation settings. SciLaws-Real uses fixed scientific data to evaluate proposed laws for held-out predictive fit and scientific validity. SciLaws-Parallel evaluates recovery of a newly synthesized structural variant of a published equation through active queries to a simulator calibrated to the source data. Our evaluation reveals three limitations: good predictive fit need not imply scientific validity, recovering a published formula does not establish recovery of its new structural terms, and candidate selection remains a bottleneck in scientific law discovery. Project page: https://yiyihum.github.io/SciLaws-Bench
Yiming Huang, Ziche Liu, Junxia Cui +11
University of California San Diego · University of California, Irvine · Cushing Academy +1