cs.AIOct 6, 2026

Can AI Agents Make Open-Ended Scientific Discovery? Evidence from Station

Authors: Wenyu Du, Stephen Chung

Organizations: DualverseAI · University of Hong Kong · University of Cambridge

Abstract

Recent AI systems have made rapid progress in scientific discovery when given well-defined metrics, but whether they can autonomously undertake open-ended scientific discovery remains unclear. We investigate AI's ability to tackle open-ended tasks in Station, an open-world environment in which multiple agents simulate a scientific ecosystem. To tackle challenges specific to open-ended tasks, we propose augmenting Station with two mechanisms: a Supervisor mechanism and periodic Meta Reflection, which encourage persistent exploration even when intermediate metrics are lacking. We construct open-ended tasks from three recent oral papers presented at ICLR. We give agents the main research question studied in each paper while withholding the paper's results and disabling web access. We then measure how many of the original findings-partitioned into individual criteria-agents rediscover. We find that Station rediscovers 62.7% of the criteria on average, compared with 15.4% for Codex Multiagent-v2 and 14.4-20.6% for AI Scientist-v2. Ablation and behavioral analyses indicate that adding the two mechanisms together improves research coverage and continuity. We further evaluate Station on two open-ended tasks without oracle papers and find that some of the discoveries made by the agents closely match discoveries reported by researchers after the knowledge cutoff date. Together, these results indicate that a suitable environment can enable agents to autonomously make meaningful progress in open-ended scientific discovery.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights

    Sep 30, 2026Jiayi Geng, Zhengxuan Wu, Kevin S. Chen +12Scientific DiscoveryAgentic Discovery

  2. Benchmarking AI Agents for Addressing Scientific Challenges Across Scales

    Jun 10, 2026Tianyu Liu, Allen Xin Wang, Antonia Panescu +30Scientific DiscoveryAgentic Benchmarks

  3. Harnessing the Collective Intelligence of AI Agents in the Wild for New Discoveries

    Jun 9, 2026Federico Bianchi, Yongchan Kwon, Aneesh Pappu +1Scientific DiscoveryArtificial Intelligence Agents