cs.MAOct 1, 2026

Can AI Scientists Coordinate at Runtime?

Authors: Zijian Liu, Yangzhixin Luo, Junyu Lu, Yi Li, Yu Chen, David Xu, William F. Shen, Xinchi Qiu, +1 more

Organizations: University of Oxford · King Abdullah University of Science and Technology · University of Sydney · University of Cambridge

Abstract

Multi-agent AI scientists have shown improving performance across a diverse range of tasks. Yet a common approach is design-time agentic orchestration, which typically relies on fixed workflows. In contrast, human scientists coordinate and adjust their division of labor at runtime. We therefore ask: can AI scientists also coordinate at runtime? To this end, we introduce Runtime Agent Coordination (RAC), which selects agents from existing AI-scientist hosts during execution, assigns scoped work contracts, and provides artifact-grounded verification. Verification informs subsequent agents without blocking transitions or discarding artifacts. We conduct a single-seed exploratory evaluation across Agent Laboratory, EvoScientist, and ARK on ResearchClawBench, preserving host models, tools, and permissions under host-calibrated budgets. Four cumulative conditions separate native execution, runtime communication, runtime selection, and the combined addition of contracts and verification. Runtime selection yields the highest observed mean score for each host; adding contracts and verification reduces these means, with host-dependent outcomes relative to native execution. These results motivate runtime coordination while exposing the limits of additional coordination mechanisms under constrained budgets. Code is available at https://github.com/systemind-team/Runtime-AI-Scientist.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Science Earth: Towards A Planet-Scale Operating System for AI-Native Scientific Discovery

    May 31, 2026Zhe Zhao, Haibin Wen, Yingcheng Wu +10DiscoveryArtificial Intelligence Systems

  2. Cross-domain benchmarks reveal when coordinated AI agents improve scientific inference from partial evidence

    May 21, 2026Fiona Y. Wong, Markus J. BuehlerScientific DiscoveryExoplanet Atmospheres