Can AI Scientists Coordinate at Runtime?
Organizations: University of Oxford · King Abdullah University of Science and Technology · University of Sydney · University of Cambridge
Abstract
Multi-agent AI scientists have shown improving performance across a diverse range of tasks. Yet a common approach is design-time agentic orchestration, which typically relies on fixed workflows. In contrast, human scientists coordinate and adjust their division of labor at runtime. We therefore ask: can AI scientists also coordinate at runtime? To this end, we introduce Runtime Agent Coordination (RAC), which selects agents from existing AI-scientist hosts during execution, assigns scoped work contracts, and provides artifact-grounded verification. Verification informs subsequent agents without blocking transitions or discarding artifacts. We conduct a single-seed exploratory evaluation across Agent Laboratory, EvoScientist, and ARK on ResearchClawBench, preserving host models, tools, and permissions under host-calibrated budgets. Four cumulative conditions separate native execution, runtime communication, runtime selection, and the combined addition of contracts and verification. Runtime selection yields the highest observed mean score for each host; adding contracts and verification reduces these means, with host-dependent outcomes relative to native execution. These results motivate runtime coordination while exposing the limits of additional coordination mechanisms under constrained budgets. Code is available at https://github.com/systemind-team/Runtime-AI-Scientist.
Figures & tables
| ID | Configuration | Newly enabled mechanism |
|---|---|---|
| N0 | Native lifecycle | Native execution without additional communication mechanisms or a RAC phase loop. |
| R1 | Communication | Inter-agent communication with the native successor at each handoff. |
| R2 | + Runtime selection | Select a host capability from a fresh checkpoint using the shared policy. |
| R3 | + Contracts and verification | Scope the task and check artifacts; pass supported, refuted, or inconclusive feedback to the next selected agent. Retain state and transitions for all verdicts. |
| Host | Condition | Mean score | Mean total input tokens (M) | Est. cost (USD) |
|---|---|---|---|---|
| ARK | Native (N0) | 16.66 | 48.85 | 69.31 |
| ARK | Full RAC (R3) | 17.98 | 68.76 | 96.69 |
| Agent Laboratory | Native (N0) | 5.47 | 4.26 | 7.09 |
| Agent Laboratory | Full RAC (R3) | 9.88 | 2.56 | 4.48 |
| EvoScientist | Native (N0) | 15.99 | 11.49 | 15.44 |
| EvoScientist | Full RAC (R3) | 15.96 | 38.58 | 51.50 |
| Shared configuration | ARK | Agent Lab. | EvoScientist |
|---|---|---|---|
| N0: Native lifecycle | 16.66 | 5.47 | 15.99 |
| R1: Runtime communication | 17.40 | 9.63 | 12.07 |
| R2: + Runtime selection | 18.42 | 12.08 | 18.53 |
| R3: + Contracts and verification | 17.98 | 9.88 | 15.96 |
| Host | Condition | Neuroscience 000 | Energy 001 | Life 001 | Neuroscience 002 | Physics 002 |
|---|---|---|---|---|---|---|
| ARK | N0 | 10.00 | 29.30 | 9.80 | 6.60 | 27.60 |
| R1 | 7.60 [-2.40] | 27.60 [-1.70] | 9.40 [-0.40] | 5.70 [-0.90] | 36.70 [+9.10] | |
| R2 | 8.40 [-1.60] | 31.90 [+2.60] | 14.10 [+4.30] | 5.70 [-0.90] | 32.00 [+4.40] | |
| R3 | 13.80 [+3.80] | 27.90 [-1.40] | 10.45 [+0.65] | 3.15 [-3.45] | 34.60 [+7.00] | |
| Agent Lab. | N0 | 9.00 | 9.20 | 1.30 | 3.30 | 0.00 |
| R1 | 8.00 [-1.00] | 10.50 [+1.30] | 4.65 [+3.35] | 5.40 [+2.10] | 6.60 [+6.60] |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Record | Contents | Use |
|---|---|---|
| Checkpoint | Run and hop identifiers, objective, native state, artifacts, unresolved problems, remaining budget, and available capabilities. | State from which R2–R3 select the next authorized capability. |
| work.request | Structured request to the selected agent. | Hands off work through the communication channel. |
| work.result | Returned result with artifacts and usage records. | Returns work and charges usage to the run. |
| Contract (R3) | Objective, readable and writable artifacts, and required output. | Scopes the selected agent’s task. |
| work.verification (R3) | Verdict ( supported , refuted , or inconclusive ), checks, reasons, and feedback status. | Feedback for the next selected agent’s prompt. |
| Provenance | Communication backend and non-secret room identifier; never invite tokens. | Attributes each trace to its channel. |
| Host | Condition | Input (M) | Output (K) | Est. cost (USD) | |
|---|---|---|---|---|---|
| Agent Laboratory | N0 | 10/10 | 4.26 | 369.42 | 7.09 |
| Agent Laboratory | R1 | 10/10 | 2.90 | 269.91 | 4.89 |
| Agent Laboratory | R2 | 10/10 | 2.36 | 308.18 | 4.33 |
| Agent Laboratory | R3 | 9/10 | 2.56 | 279.24 | 4.48 |
| EvoScientist | N0 | 10/10 | 11.49 | 70.34 | 15.44 |
| EvoScientist | R1 | 8/10 | 18.31 | 126.37 | 24.67 |
| ID | Task | Host | N0 | R1 | R2 | R3 |
|---|---|---|---|---|---|---|
| Pair 1 | nls_ses / metadata_0 | Agent Laboratory | 0.00 | 13.33 | 22.22 | 100.00 |
| Pair 2 | nls_ses / metadata_0 | EvoScientist | 12.50 | 44.44 | 44.44 | 0.00 |
| Pair 3 | nls_incarceration / metadata_2 | Agent Laboratory | 0.00 | 20.00 | 0.00 | 66.67 |
| Pair 4 | meta_regression / metadata_0 | Agent Laboratory | 9.38 | 12.00 | 25.00 | 0.00 |
| Pair 5 | archaeology / metadata_13 | Agent Laboratory | 0.00 | 50.00 | 28.57 | 28.57 |
| Pair 6 | requirements_engineering_ for_ML_enabled_systems / metadata_0 | Agent Laboratory | 0.00 | 0.00 | 19.05 | 0.00 |
| Neuroscience-000 Weights: 20%, 20%, 20%, 20%, 20%. Host Cond. C1 C2 C3 C4 C5 Final AL N0 12 5 0 0 28 9.00 AL R1 25 3 0 0 12 8.00 AL R2 5 0 0 0 12 3.40 AL R3 24 0 0 0 28 10.40 Evo N0 24 0 0 0 18 8.40 Evo R1 28 0 0 0 5 6.60 Evo R2 18 0 0 0 12 6.00 Evo R3 28 22 0 0 18 13.60 ARK N0 24 3 5 0 18 10.00 ARK R1 8 8 2 8 12 7.60 ARK R2 22 8 0 0 12 8.40 ARK R3 22 12 5 2 28 13.80 | Energy-001 Weights: 10%, 30%, 20%, 20%, 20%. Host Cond. C1 C2 C3 C4 C5 Final AL N0 2 0 8 12 25 9.20 AL R1 15 8 3 2 28 10.50 AL R2 21 3 0 21 25 12.20 AL R3 3 0 0 0 28 5.90 Evo N0 24 25 8 3 28 17.70 Evo R1 35 12 12 18 34 19.90 Evo R2 8 28 37 18 28 25.80 Evo R3 85 22 18 18 31 28.50 ARK N0 45 38 18 18 31 29.30 ARK R1 35 35 32 8 28 27.60 ARK R2 25 42 34 22 28 31.90 ARK R3 25 42 18 22 24 27.90 |
| Life-001 Weights: 35%, 20%, 15%, 30%. Host Cond. C1 C2 C3 C4 C5 Final AL N0 2 0 0 2 – 1.30 AL R1 3 0 12 6 – 4.65 AL R2 12 3 12 1 – 6.90 AL R3 2 2 12 0 – 2.90 Evo N0 18 24 34 0 – 16.20 Evo R1 18 12 32 0 – 13.50 Evo R2 8 18 41 4 – 13.75 Evo R3 18 18 28 0 – 14.10 ARK N0 8 8 28 4 – 9.80 ARK R1 8 12 28 0 – 9.40 ARK R2 18 12 32 2 – 14.10 ARK R3 12 8 31 0 – 10.45 | Neuroscience-002 Weights: 15%, 25%, 25%, 20%, 15%. Host Cond. C1 C2 C3 C4 C5 Final AL N0 2 12 0 0 0 3.30 AL R1 0 12 0 12 0 5.40 AL R2 0 12 0 12 2 5.70 AL R3 0 12 0 8 3 5.05 Evo N0 0 12 0 18 8 7.80 Evo R1 8 12 0 2 0 4.60 Evo R2 0 14 0 18 8 8.30 Evo R3 0 12 0 18 3 7.05 ARK N0 0 12 0 12 8 6.60 ARK R1 2 12 0 12 0 5.70 ARK R2 0 12 0 12 2 5.70 ARK R3 1 12 0 0 0 3.15 |
| Math-000 Weights: 15%, 15%, 70%. Host Cond. C1 C2 C3 C4 C5 Final AL N0 45 32 8 – – 17.15 AL R1 32 38 12 – – 18.90 AL R2 44 36 15 – – 22.50 AL R3 34 32 12 – – 18.30 Evo N0 24 27 12 – – 16.05 Evo R1 3 5 12 – – 9.60 Evo R2 24 36 18 – – 21.60 Evo R3 28 37 18 – – 22.35 | Math-001 Weights: 40%, 30%, 30%. Host Cond. C1 C2 C3 C4 C5 Final AL N0 0 8 29 – – 11.10 AL R1 0 46 31 – – 23.10 AL R2 8 39 42 – – 27.50 AL R3 23 38 31 – – 29.90 Evo N0 6 22 26 – – 16.80 Evo R1 35 34 28 – – 32.60 Evo R2 32 45 33 – – 36.20 Evo R3 35 43 36 – – 37.70 |
| Math-002 Weights: 40%, 30%, 30%. Host Cond. C1 C2 C3 C4 C5 Final AL N0 0 0 0 – – 0.00 AL R1 3 5 3 – – 3.60 AL R2 12 0 12 – – 8.40 AL R3 12 0 0 – – 4.80 Evo N0 7 8 18 – – 10.60 Evo R1 5 6 5 – – 5.30 Evo R2 14 3 15 – – 11.00 Evo R3 5 2 2 – – 3.20 | Math-003 Weights: 40%, 35%, 25%. Host Cond. C1 C2 C3 C4 C5 Final AL N0 0 0 0 – – 0.00 AL R1 12 15 0 – – 10.05 AL R2 6 12 0 – – 6.60 AL R3 12 12 0 – – 9.00 Evo N0 31 9 3 – – 16.30 Evo R1 21 3 0 – – 9.45 Evo R2 15 15 0 – – 11.25 Evo R3 18 5 3 – – 9.70 |
| Physics-002 Weights: 20%, 20%, 15%, 20%, 25%. Host Cond. C1 C2 C3 C4 C5 Final AL N0 0 0 0 0 0 0.00 AL R1 5 12 2 12 2 6.60 AL R2 18 32 3 28 34 24.55 AL R3 2 3 2 15 0 4.30 Evo N0 32 45 28 38 48 39.20 Evo R1 6 8 5 28 5 10.40 Evo R2 32 36 24 32 78 43.10 Evo R3 12 25 8 8 25 16.45 ARK N0 28 35 24 12 36 27.60 ARK R1 38 42 34 18 48 36.70 ARK R2 31 32 22 28 42 32.00 ARK R3 32 42 28 18 48 34.60 | Information-003 Weights: 15%, 20%, 30%, 20%, 15%. Host Cond. C1 C2 C3 C4 C5 Final AL N0 8 8 0 4 0 3.60 AL R1 15 8 0 8 0 5.45 AL R2 8 0 6 0 0 3.00 AL R3 12 6 12 8 0 8.20 Evo N0 8 22 12 8 0 10.80 Evo R1 18 6 12 6 0 8.70 Evo R2 12 22 3 6 0 8.30 Evo R3 14 3 12 3 0 6.90 |
| Hop | Agent | Reported work and outcome | Recorded observed next |
|---|---|---|---|
| Initial research and first refinement (hops 0–6) | |||
| 0 | Researcher | Established the RCS fidelity-estimation project context. | Exp. Exp. |
| 1 | Experimenter | Executed the initial five experiments and summarized results. | Wri. Wri. |
| 2 | Writer | Wrote and compiled the initial manuscript. | Rev. Rev. |
| 3 | Reviewer | No summary available; reviewer invocation retained. | Rev. Plan. |
| 4 | Planner | Prepared an action plan from the available review state. | Exp. Exp. |
| Hop | Agent | Reported work and outcome | Recorded observed next |
|---|---|---|---|
| Initial research and report (hops 0–2) | |||
| 0 | Researcher | Established the vaccine-design project context. | Exp. Exp. |
| 1 | Experimenter | Ran five experiments; reported repairing pivot indexing and verifying 17 output files. | Wri. Wri. |
| 2 | Writer | Wrote and compiled the initial manuscript. | Rev. Rev. |
| Direct reviewer–writer corrections (hops 3–20) | |||
| 3 | Reviewer | Reviewed the method–result mismatch and missing evidence. | Plan. Wri. |