Scientific discovery begins where known problems end. There, AI systems must engage in exploration: framing hypotheses, designing experiments, and iterating on the results. However, evaluating this ability is difficult: (1) how to verify whether a genuinely new hypothesis holds, and (2) how to determine whether a system has discovered it through exploration or merely recalled related knowledge from pre-training data. To this end, we introduce ExplorationBench, which turns the wicked problem of evaluating scientific exploration into a concrete and tractable framework built on verifiable Alien Worlds: their rules are executable, so every answer can be checked exactly, and they conflict with familiar knowledge, so recall alone cannot solve the tasks. The benchmark contains two sandboxes, AlienCode (31 discovery targets, 70 tasks) and AlienLogic (24 discovery targets, 70 tasks). Each sandbox provides a flawed manual, task-specific environmental feedback, and a dedicated tool-call schema. Systems use these resources to explore the sandbox, then solve held-out tasks. We evaluate 10 AI systems and find that the strongest systems can acquire and apply unfamiliar rules, while performance varies substantially across trajectories and continued exploration can stall or reverse earlier gains. ExplorationBench represents a step towards AI systems that can acquire and apply genuinely new knowledge through exploration in unknown environments.
Figures & tables
Figure 1: Held-out accuracy over autonomous exploration. Top: AlienCode ; bottom: AlienLogic . Each cell highlights one system’s M0 – M4 curve; the other nine are grey. The upper-left and lower-right values are M4 and M0 . Cells are ordered by M4 within each sandbox. Each curve is the Best@ 3 trajectory, and each milestone is the mean of three answers per held-out question.
Figure 2: ExplorationBench : what a system is given, what it may do about it, and how it is scored. (A) Two executable sandboxes. AlienCode has 31 discovery targets and AlienLogic 24, so the manual each ships with is wrong in ways no amount of prior knowledge recovers – the changes have to be found from evidence. (B) Every system starts from the same flawed manual and the same fixed worked examples, then runs four rounds in which it chooses its own probes, runs them in the environment, and adds the results to its exploration history. Nothing else enters the context. (C) At each milestone Mt , the system is tested without tools: it answers the 70 held-out tasks and separately reports the rules it believes it has found. A trajectory is scored on where it ends, and a system on its best complete trajectory.
Figure 3: Held-out task composition. The 70 tasks of each sandbox by family.
Table 1: Held-out accuracy under autonomous exploration. M0 and M4 of each system’s Best@3 trajectory ( eq. 3 ), the mean and the lowest M4 over its three trajectories, and accuracy under open-book answering, with the complete rule set supplied before exploration (O@ M0 ) or after the Best@3 trajectory’s exploration (A4+O). All values are percentages, each the mean of three answers per question; rows are ordered by Best@3 M4 .
Figure 4: Held-out accuracy at M4 under five exploration conditions. Each dot is one system; the horizontal line in each column marks the median over the ten systems. Autonomous exploration is shown at each system’s Best@3 trajectory; hindsight exploration replays the probes of that same trajectory, and fixed-probe exploration issues model-independent probes at matched volume. Shaded columns receive no environment feedback, and from left to right the system takes a larger part in designing the experiments. Control conditions carry one trajectory per system, and every value is the mean of three answers per question.
Figure 5: Best@3 M4 in the two sandboxes. Each row is one system, ordered by its AlienCode score; the right column gives its rank in AlienCode and then in AlienLogic . Red marks a fall of three or more places, blue a rise of three or more.
Figure 6: Finding out versus being told. Held-out accuracy at M4 under open-book answering before exploration (O@ M0 , hatched), after autonomous exploration alone (Best@3, solid), and under open-book answering after that exploration (A4+O, outlined). Systems are ordered by Best@3 M4 within each sandbox.
Figure 7: Accuracy jumps when the keystone rules are found. Each block is one system and each row one of its three AlienCode trajectories; columns are milestones, and each cell gives held-out accuracy (%). Triangles mark the milestone at which the trajectory’s rule report first states R15 ( CARVE slice shift, pointing up) and R14 ( PLUCK index shift, pointing down) correctly.
Figure 8: Knowing a rule is not being able to use it. For each system, pooled over its three AlienCode trajectories, filled dots give accuracy at M4 on tasks whose required rules the M4 rule report states correctly, and open dots on tasks with at least one required rule missing or stated incorrectly. The dashed segment up to 100% is the remaining execution gap; n counts task–trajectory pairs in the first group.
Figure 9: Outcomes vary between trajectories, not between answers. Each row is one system. Dots are its three trajectories’ M4 (filled for Best@3), the coloured band spans them, and the short vertical line marks Mean@3. The grey halo around each dot extends one standard deviation of that trajectory’s three answers per question to either side.
Figure 10: When the gain arrives. For each system’s Best@3 trajectory, bubble area is the change in held-out accuracy from the previous milestone ( Mt−Mt−1 ). The solid bubble marks the largest step and is labelled with its size; a red ring marks a milestone at which accuracy fell. The right column gives M4 .
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Learning setup
What is measured
Benchmark
Target
Evidence
Novelty control
Progress
Transfer
Ground truth
CL-bench [ 10 ]
Context learning
Provided context
Expert-authored content
Final score
Same context
Expert rubrics, LLM verifier
EvaLearn [ 8 ]
Sequential learning
Prior solved tasks
Authored task sequences
Learning curve
Later related tasks
Rubrics, LLM verifier
SE-Bench [ 37 ]
Weight internalization
Docs, training tasks
Obfuscated APIs
Pre/post score
Closed-book held-out
Tests, AST checks
SWE-bench [ 16 ]
Software repair
Issue, codebase
Real GitHub issues
Final patch
Same repository
Test suites
DiscoveryWorld [ 15 ]
Scientific investigation
Agent actions
Fictional worlds
Final score
Same world
World state
Appendix
Table 2: Where the evidence comes from, and where competence is tested. Each row states a benchmark’s primary protocol rather than ranking design choices. ExplorationBench measures exploration by combining autonomous probe selection, milestone measurement, and unseen-task evaluation after interaction ends; no single column alone defines that distinction.
AlienCode
AlienLogic
System
C
P
T (M)
C=P
T (M)
Claude Opus 5
48
168
11.59
48
0.67
GPT-5.6 Sol
48
131
3.62
48
0.91
Gemini 3.8 Flash
45
124
2.41
48
1.08
Kimi K3
45
216
2.01
48
0.38
Grok 4.6
48
197
0.50
48
0.15
Appendix
Table 3: Exploration budget of each system’s Best@3 trajectory. C is tool calls, P probe units, and T exploration tokens in millions, summed over the four rounds. In AlienLogic , C=P . Rows follow the AlienCode order of table 1 .
Figure 11: Task-family gains for every AlienCode system. Cells show M4−M0 within each task family, in percentage points. Tasks are split by compositional depth and by representation, so each task enters one column of each. Autonomous rows average three trajectories; control rows are one trajectory each.
Figure 12: Task-family gains for every AlienLogic system. Cells show M4−M0 within each task family, in percentage points. Columns are the five families of fig. 3 : provable theorems by proof length, unprovable ones by how deeply the goal nests. Autonomous rows average three trajectories; control rows are one trajectory each.
Figure 13: Exploration trajectories for every AlienCode system. Curves show cumulative gain Mt−M0 . The autonomous curve is the mean of three trajectories, and its shading is ± population SD across them; each control curve is a single trajectory.
Figure 14: Exploration trajectories for every AlienLogic system. Curves show cumulative gain Mt−M0 . The autonomous curve is the mean of three trajectories, and its shading is ± population SD across them; each control curve is a single trajectory.
We introduce NatureBench, a cross-discipline benchmark of 90 tasks distilled from peer-reviewed Nature-family publications, designed to evaluate whether AI coding agents can move beyond reproduction toward discovery on real scientific problems. NatureBench is built on NatureGym, an automated pipeline that constructs a standardized, per-task containerized environment from a source paper, addressing the environment-fragmentation problem that has limited the credibility of prior agent-on-research benchmarks. Evaluating ten frontier agent configurations under a strict web-search-disabled protocol, we find that the strongest model surpasses SOTA on only 17.8% of tasks under the g>0.1 criterion. Analysis of method pathways reveals that agents succeed primarily through methodological translation, converting scientific tasks into familiar supervised prediction problems, rather than through genuine scientific invention. Failures are dominated by wrong method choice and insufficient compute budget, not by task misunderstanding. We release the benchmark, the NatureGym pipeline, and a public leaderboard with maintainer-side reproduction. Code: https://github.com/FrontisAI/NatureBench
Reliable automated research requires agents to vet data, verify analyses, and generate hypotheses grounded in trustworthy evidence, potentially reducing routine scientific workload while allowing scientists to focus on interpretation and discovery. Existing benchmarks often only assess analytical task completion or hypothesis generation separately rather than testing whether reliable evidence supports valid and novel claims. We introduce DISCERN (Data Integrity and Scientific Capability: Evidence, Reasoning, and Novelty), a controlled benchmark on real, publicly available datasets that evaluates three key levels of an automated research workflow. The first two levels test data integrity and analysis verification under confounds and tool traps, while the third tests hypothesis generation and revision under adversarial review, including counterfactual cases in which evidence consistent with real data and documented scientific phenomena conflicts with established expectations, motivating alternative explanations and testable hypotheses. Across 203 tasks, eight life-science tracks, and eight models, DISCERN shows that strong aggregate performance can mask level-specific weaknesses. Agents earn perfect scores in only 60.8% of Level 1, 34.2% of Level 2, and 0.6% of Level 3 evaluations, with penalties attributed to rejection of sound data, failure to carry recognized limitations into conclusions, and wide variation in hypothesis production. Cross-track rankings by token and code use are substantially more stable than rankings by evidence judgment, suggesting greater consistency in computational effort than in evidence-based reasoning. These profiles identify opportunities for supervised scientific assistance, but current agents do not yet demonstrate reliable autonomous analysis or discovery. Code and data: https://huggingface.co/datasets/discern-bench-anon/discern-benchmark
Nan Huang, Mario Tapia-Pacheco, Kun Zhou +4
University of California San Diego · Independent Researcher
When Isaac Newton discovered the law of gravitation, he did so through an iterative process of analyzing observed data such as planetary patterns, finding the underlying mechanisms by describing patterns in mathematical equations, and refining his theory against the Moon's orbit, revealing the startling insight that the same force governs both falling apples and orbiting planets. Would it be possible for AI agents to make similar discoveries? To measure this ability, we introduce EurekaBench, a cross-domain benchmark that tests AI agents' ability to conduct long-horizon experiments and discover mechanisms that explain observations. We evaluate these mechanisms by the scientific insights that can be derived from them. EurekaBench contains an expert-verified set of 26 long-horizon tasks across neuroscience, computer science, chemistry, astrophysics, geophysics, and plasma physics, with a total of 306 scientific insights that the discovered mechanisms are expected to support. Our evaluation framework tests three axes of scientific discovery: agents' ability to follow known scientific constraints, the predictive accuracy of the discovered mechanisms, and whether these mechanisms yield scientific insights or inform future research. Our results show that current AI agents often overly fixate on predictive accuracy optimization, surpassing human scientists, while falling substantially short in deriving scientific insights.
Jiayi Geng, Zhengxuan Wu, Kevin S. Chen +12
Carnegie Mellon University, Language Technologies Institute · Stanford University, Department of Computer Science · Yale University, Quantitative Biology Institute +8