Can AI Agents Make Open-Ended Scientific Discovery? Evidence from Station
Organizations: DualverseAI · University of Hong Kong · University of Cambridge
Abstract
Recent AI systems have made rapid progress in scientific discovery when given well-defined metrics, but whether they can autonomously undertake open-ended scientific discovery remains unclear. We investigate AI's ability to tackle open-ended tasks in Station, an open-world environment in which multiple agents simulate a scientific ecosystem. To tackle challenges specific to open-ended tasks, we propose augmenting Station with two mechanisms: a Supervisor mechanism and periodic Meta Reflection, which encourage persistent exploration even when intermediate metrics are lacking. We construct open-ended tasks from three recent oral papers presented at ICLR. We give agents the main research question studied in each paper while withholding the paper's results and disabling web access. We then measure how many of the original findings-partitioned into individual criteria-agents rediscover. We find that Station rediscovers 62.7% of the criteria on average, compared with 15.4% for Codex Multiagent-v2 and 14.4-20.6% for AI Scientist-v2. Ablation and behavioral analyses indicate that adding the two mechanisms together improves research coverage and continuity. We further evaluate Station on two open-ended tasks without oracle papers and find that some of the discoveries made by the agents closely match discoveries reported by researchers after the knowledge cutoff date. Together, these results indicate that a suitable environment can enable agents to autonomously make meaningful progress in open-ended scientific discovery.
Figures & tables
| Task | Task type | Main research question | Experimental setup |
|---|---|---|---|
| Emergent planning in RL agents [ 11 ] | Rediscovery | Do model-free RL agents form internal plans, and how do these plans guide their behavior? | RL environment and a trained RL agent checkpoint. |
| Low-rank structure in LLM outputs [ 12 ] | Rediscovery | Do LLM outputs exhibit low-rank structure that can support prediction or generation? | A trained LLM checkpoint and an environment for analyzing its outputs. |
| Temporal representations in RNNs [ 13 ] | Rediscovery | How do RNNs represent temporal information and balance spatial and temporal memory demands? | A -delay task and an environment for training and analyzing RNNs. |
| Subliminal learning in LLMs | Exploratory | How do behavioral traits transfer between LLMs through semantically unrelated data? | A teacher–student fine-tuning setup and a brief summary of prior work. |
| Hallucination in VLMs | Exploratory | How do knowledge-dependent and visual hallucinations arise, and how can they be reduced? | A VLM evaluation setup and a proposal distinguishing the two types of hallucinations. |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| System | Output tokens (M) | Cumulative experiment time (h) |
|---|---|---|
| Emergent Planning in RL Agents | ||
| Station (Tick 100) | 3.2–3.5 | 10.7–29.8 |
| Station (Tick 300) | 7.5–9.0 | 34.4–57.3 |
| Codex Multiagent-v2 (GPT) | 3.3–4.3 | 50.2–300.5 |
| AI Scientist-v2 (GPT) | 8.9–10.8 | 70.3–167.2 |
| AI Scientist-v2 (Gemini) | 5.3–5.9 | 46.0–131.8 |
| System | Probing | Planning features | Causal intervention | Total |
|---|---|---|---|---|
| Emergent Planning in RL Agents | ||||
| Station (Tick 300) | 99.1 [97.2, 100.0] | 63.9 [50.0, 72.2] | 60.5 [51.9, 66.7] | 75.8 [71.7, 80.8] |
| Station (Tick 100) | 99.1 [97.2, 100.0] | 52.8 [25.0, 75.0] | 63.0 [55.6, 66.7] | 72.4 [62.6, 78.8] |
| Codex Multiagent-v2 (GPT) | 16.7 [0.0, 25.0] | 0.0 [0.0, 0.0] | 0.0 [0.0, 0.0] | 6.1 [0.0, 9.1] |
| AI Scientist-v2 (Claude) | 52.8 [50.0, 58.3] | 0.0 [0.0, 0.0] | 0.0 [0.0, 0.0] | 19.2 [18.2, 21.2] |
| AI Scientist-v2 (Gemini) | 41.7 [0.0, 75.0] | 8.3 [0.0, 25.0] | 0.0 [0.0, 0.0] | 18.2 [0.0, 27.3] |
| System | Recovery (%, mean SE) |
|---|---|
| Station (Tick 300) | |
| Station (Tick 100) | |
| Codex Multiagent-v2 (GPT) | |
| AI Scientist-v2 (Claude) | |
| AI Scientist-v2 (Gemini) | |
| AI Scientist-v2 (GPT) |
| Trait | Rank | Base (%) | All-layer LoRA (%) | Early-layer LoRA (%) |
|---|---|---|---|---|
| Cat | 128 | 5.72 | 4.43 0.29 | 30.23 8.66 |
| Cat | 256 | 5.72 | 5.54 0.27 | 18.81 5.52 |
| Eagle | 128 | 1.06 | 2.79 0.95 | 64.79 19.93 |
| Eagle | 256 | 1.06 | 2.38 0.84 | 62.38 20.40 |
| Owl | 128 | 0.28 | 1.07 0.34 | 24.00 6.70 |
| Owl | 256 | 0.28 | 0.75 0.08 | 13.95 6.07 |
| Early layers | Seed 0 (%) | Seed 1 (%) | Seed 2 (%) | Mean s.d. (%) |
|---|---|---|---|---|
| 6 | 41.86 | 29.44 | 51.42 | |
| 8 | 52.78 | 40.36 | 49.74 | |
| 10 | 43.24 | 32.06 | 46.28 | |
| 12 | 41.90 | 33.60 | 39.44 | |
| 14 (default) | 37.68 | 32.28 | 20.72 | |
| 16 | 29.02 | 23.22 | 25.42 |
| Model | Baseline accuracy (%) | Intervention accuracy (%) | Accuracy gain (pp) |
|---|---|---|---|
| InternVL3.5-4B | 75.27 0.51 | 78.97 0.45 | +3.70 |
| InternVL3.5-8B | 88.37 0.83 | 89.17 0.32 | +0.80 |
| Qwen3-VL-4B | 89.80 0.10 | 91.03 0.12 | +1.23 |
| Qwen3-VL-8B | 58.53 0.15 | 71.23 0.15 | +12.70 |
| MLP scale | Seed 1 (%) | Seed 2 (%) | Seed 3 (%) | Mean s.d. (%) |
|---|---|---|---|---|
| 1.00 (Baseline) | 89.3 | 88.1 | 87.7 | 88.37 0.83 |
| 0.75 | 89.7 | 89.3 | 89.3 | 89.43 0.23 |
| 0.50 | 89.4 | 89.3 | 88.8 | 89.17 0.32 |
| 0.25 | 86.7 | 86.2 | 86.3 | 86.40 0.26 |