Collaborative Principle Evolution via Evidence Transfer for Scientific Discovery
Organizations: Zhejiang University, Hangzhou, Zhejiang, China · Westlake University, Hangzhou, Zhejiang, China
Abstract
Large Language Model (LLM)-based agents promise to automate scientific discovery, yet exploring the vast hypothesis space remains costly. Existing principle-evolution methods accelerate this loop, but operate sequentially, which caps exploration breadth and wastes wall-clock time on challenging problems. To address this, we formulate collaborative scientific discovery as evidence transfer between parallel principle-evolution branches. We present COEVOLVE, which realizes this transfer through a coordination core over parallel branches. By integrating value-of-information-gated routing and context-discounted likelihood injection, COEVOLVE enables branches to collaborate through shared measurements while keeping their principle posteriors separate. Across six scientific-discovery tasks under a matched evaluation budget, COEVOLVE attains a mean solution quality of 66.5% versus 57.0% for single-branch principle evolution, with a 1.80x mean wall-clock speedup on the GPT-5.6-Terra backbone; on five auto-research tasks delegated to an autonomous research harness, it is the only arm whose mean stays above the published SOTA anchor on every task. These results establish when evidence sharing accelerates parallel discovery and when transfer safeguards are necessary to limit negative or inert transfers
Figures & tables
| Model / Method | MBO | NHO | SPO | TMC | AMP | Promoter | Average |
| Gemma-4-31B-IT | |||||||
| Vanilla MAS | 39.2 0.8 | 77.2 1.4 | 0.0 0.0 0.0x | 74.4 0.0 1.3x | 27.4 2.3 1.1x | 84.4 0.3 4.5x | 50.4 1.2x |
| The AI Scientist v1 ( Lu et al., 2024 ) | 48.6 2.2 1.2x | 78.3 4.2 1.0x | 0.0 0.0 0.0x | 84.5 0.0 1.5x | 24.1 1.5 | 18.7 32.4 | 42.4 1.0x |
| The AI Scientist v2 ( Yamada et al., 2025 ) | 48.5 5.1 1.2x | 84.3 3.9 1.1x | 8.2 1.6 | 83.7 0.8 1.5x | 27.7 3.0 1.2x | 0.0 0.0 0.0x | 42.1 |
| AI Researcher ( Tang et al., 2025 ) | 50.0 2.4 1.3x | 88.3 1.8 1.1x | 24.6 2.0 3.0x | 81.5 6.1 1.5x | 26.7 14.6 1.1x | 19.3 33.5 1.0x | 48.4 1.2x |
| EvoScientist ( Lyu et al., 2026 ) | 45.3 1.5 1.2x | 93.3 0.7 1.2x | 9.2 0.7 1.1x | 85.4 0.0 1.5x | 29.7 2.6 1.2x | 54.0 7.8 2.9x | 52.8 1.3x |
| Model / Method | MBO | NHO | SPO | TMC | AMP | Promoter | Average | |||||||
| APD | AUOC | APD | AUOC | APD | AUOC | APD | AUOC | APD | AUOC | APD | AUOC | Avg APD | Avg AUOC | |
| Gemma-4-31B-IT | ||||||||||||||
| Vanilla MAS | 79.1 1.1 | 37.5 0.9 | 24.0 1.1 | 76.5 1.0 | 0.0 0.0 | 0.0 0.0 | 29.8 4.8 | 73.7 0.5 | 76.1 0.9 | 25.7 2.6 | 36.9 1.1 | 80.2 4.4 | 41.0 | 48.9 |
| The AI Scientist v1 ( Lu et al., 2024 ) | 21.7 17.0 | 38.2 3.4 | 5.9 3.4 | 76.8 3.6 | 0.0 0.0 | 0.0 0.0 | 53.1 1.1 | 81.5 1.0 | 58.5 3.0 | 17.2 1.1 | 0.0 0.0 | 18.7 32.4 | 23.2 0.6x | 38.7 0.8x |
| The AI Scientist v2 ( Yamada et al., 2025 ) | 34.7 12.5 | 43.6 3.5 | 12.9 2.2 | 80.5 3.0 | 4.6 2.9 | 7.7 1.0 | 47.4 4.2 | 78.5 0.7 | 34.3 6.2 | 25.7 3.5 | 0.0 0.0 | 0.0 0.0 | 22.3 0.5x | 39.3 0.8x |
| AI Researcher ( Tang et al., 2025 ) | 46.5 9.0 | 46.1 2.1 | 16.6 4.8 | 84.5 1.4 | 4.7 0.1 | 21.2 1.9 | 52.3 5.3 | 81.1 5.9 | 73.1 8.7 | 23.5 14.8 | 18.3 31.7 | 18.3 31.7 | 35.3 0.9x | 45.8 0.9x |
| Arm | ALDE | Deconv | InvScat | D2D | MolEdit | Average |
|---|---|---|---|---|---|---|
| Claude Code | +2.3 0.2 | +14.0 7.3 | 5.8 7.6 | +9.0 0.3 | +0.2 1.7 | +3.9 |
| Codex | +2.2 1.0 | +14.3 5.4 | 5.9 6.7 | +11.2 0.9 | +1.2 1.9 | +4.6 |
| Arbor | +1.3 1.9 | +11.7 23.8 | 6.5 5.9 | +15.3 0.1 | +0.5 0.6 | +4.5 |
| PiEvo | +2.4 2.0 | +15.4 28.6 | 5.3 7.9 | +15.4 0.1 | +1.3 1.6 | +5.8 |
| CoEvolve | +6.6 0.7 | +45.4 1.4 | +11.9 3.7 | +15.5 0.0 | +4.1 0.6 | +16.7 |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Symbol | Default | Role |
|---|---|---|
| safe-VoI gate margin (Eq. ( 2 )) | ||
| routing quota per branch per round | ||
| redundancy cap (token-Jaccard) | ||
| import cost weight (Eq. ( 1 )) | ||
| total false-import budget (alpha-spending) | ||
| parsing cost |
| Worst-branch SQ (%) | Best-of-seed SQ (%) | ||||||
|---|---|---|---|---|---|---|---|
| Task | CoEvolve shared | Isolated | CoEvolve shared | Isolated | Independent | ||
| MBO | {\color[rgb]{0.4883,0.1641,0.8203}+9.4} | {\color[rgb]{0.4883,0.1641,0.8203}+13.6} | |||||
| NHO | {\color[rgb]{0.4883,0.1641,0.8203}+0.9} | {\color[rgb]{0.4883,0.1641,0.8203}+0.5} | |||||
| SPO | {\color[rgb]{0.4883,0.1641,0.8203}+1.3} | {\color[rgb]{0.4883,0.1641,0.8203}+2.5} | |||||
| TMC | {\color[rgb]{0.8203,0.3008,0.3008}-4.7} | {\color[rgb]{0.4883,0.1641,0.8203}+0.5} | |||||
| AMP | {\color[rgb]{0.4883,0.1641,0.8203}+8.9} ⋆ | {\color[rgb]{0.4883,0.1641,0.8203}+7.9} ⋆ | |||||
| Arm | knocked-out module | SQ | worst-br. | SQ | imports | |
|---|---|---|---|---|---|---|
| CoEvolve (full) | all modules on (main-table arm) | – | 42 | |||
| VoI gate | safe-VoI gate off ( ): every routed import accepted | {\color[rgb]{0.8203,0.3008,0.3008}-2.1} | 137 | |||
| discounting | : imports applied at full weight | {\color[rgb]{0.8203,0.3008,0.3008}-19.0} | 58 | |||
| redundancy control | : redundancy cap disabled | {\color[rgb]{0.8203,0.3008,0.3008}-9.6} | 34 | |||
| sharing (Isolated) | coordination on, zero evidence flow | {\color[rgb]{0.8203,0.3008,0.3008}-13.6} | 0 | – |
| Per-branch SQ (%) | Imports | |||||||
|---|---|---|---|---|---|---|---|---|
| Task | CoEvolve worst-branch | PiEvo | accepted | net unique | dedup | uplift (SQ pts) | inert | |
| MBO | {\color[rgb]{0.8203,0.3008,0.3008}-7.0} | 120 | 99 | |||||
| NHO | {\color[rgb]{0.8203,0.3008,0.3008}-4.1} | 19 | 17 | |||||
| SPO | {\color[rgb]{0.4883,0.1641,0.8203}+1.1} | 6 | 6 | |||||
| TMC | {\color[rgb]{0.8203,0.3008,0.3008}-4.6} | 2 | 2 | |||||
| AMP | {\color[rgb]{0.8203,0.3008,0.3008}-2.7} | 96 | 80 | |||||
| Task | Per-seed SQ differences | Paired -value |
|---|---|---|
| MBO | ||
| NHO | ||
| SPO | ||
| TMC | ||
| AMP | ||
| Promoter |
| vs PiEvo | vs InternAgent-1.5 | |||
|---|---|---|---|---|
| Task | Gemma | Terra | Gemma | Terra |
| MBO | ||||
| NHO | ||||
| SPO | ||||
| TMC | ||||
| AMP | ||||
| Billable | All-in | ||||
|---|---|---|---|---|---|
| Task | CoEvolve | PiEvo | ratio | CoEvolve | PiEvo |
| ALDE | |||||
| Deconv | |||||
| InvScat | |||||
| D2D | |||||
| MolEdit | |||||
| Task | TTT PiEvo | TTT | ratio | ||||
|---|---|---|---|---|---|---|---|
| ALDE | +0.024 | 32.5 | 20.8 | 1.56 | 6.64 | 7.67 | 0.9 |
| Deconv | +0.154 | 22.4 | 12.1 | 1.85 | 1.38 | 1.11 | 1.2 |
| InvScat | -0.053 | 47.7 | 59.7 | 0.80 | 11.87 | 15.17 | 0.8 |
| D2D | +0.154 | 31.7 | 38.0 | 0.83 | 17.05 | 13.45 | 1.3 |
| MolEdit | +0.013 | 45.9 | 30.8 | 1.49 | 11.52 | 6.03 | 1.9 |
| Task | Domain | Search space | Surrogate score | Anti-memorization gate | Ref. scale |
|---|---|---|---|---|---|
| NHO | nanophotonics | 7-dim continuous | -factor | surrogate-defined optimum | |
| MBO | bio-chemistry | SMILES (drug-like) | pChEMBL | surrogate-defined optimum; Lipinski + PAINS | |
| AMP | biology | peptide 12–50 aa | AMP probability | net charge (cationic recipe banned) | |
| Promoter | biology | 50-nt DNA | expression bin | yeast GRF + GC-box banned | |
| TMC | chemistry | 4-of-49 ligand comb. | polarizability ( ) | pool-computed optimum; charge neutrality, distinct ligands | |
| SPO | materials | cuprate formulas | (K) | record/textbook backbones banned (Hg/Tl, BSCCO, YBCO) |
| Gemma-4-31B-IT | GPT-5.6-Terra | ||||
|---|---|---|---|---|---|
| Task | Unit (scale) | CoEvolve | PiEvo | CoEvolve | PiEvo |
| NHO | -factor ( ) | ||||
| MBO | pChEMBL ( ) | ||||
| AMP | probability ( ) | ||||
| Promoter | expression bin ( ) | ||||
| TMC | polarizability ( ) | ||||