The increasing adoption of Large Language Models (LLMs) has enabled AI scientists to perform complex end-to-end scientific discovery tasks requiring coordination of specialized roles, including idea generation and experimental execution. However, most state-of-the-art AI scientist systems rely on static, hand-designed pipelines and fail to adapt based on accumulated interaction histories. As a result, these systems overlook promising research directions, repeat failed experiments, and pursue infeasible ideas. To address this, we introduce EvoScientist, an evolving multi-agent AI scientist framework that continuously improves research strategies through persistent memory and self-evolution. EvoScientist comprises three specialized agents: a Researcher Agent (RA) for scientific idea generation, an Engineer Agent (EA) for experiment implementation and execution, and an Evolution Manager Agent (EMA) that distills insights from prior interactions into reusable knowledge. EvoScientist contains two persistent memory modules: (i) an ideation memory, which summarizes feasible research directions from top-ranked ideas while recording previously unsuccessful directions; and (ii) an experimentation memory, which captures effective data processing and model training strategies derived from code search trajectories and best-performing implementations. These modules enable the RA and EA to retrieve relevant prior strategies, improving idea quality and code execution success rates over time. Experiments show that EvoScientist outperforms 7 open-source and commercial state-of-the-art systems in scientific idea generation, achieving higher novelty, feasibility, relevance, and clarity via automatic and human evaluation. EvoScientist also substantially improves code execution success rates through multi-agent evolution, demonstrating persistent memory's effectiveness for end-to-end scientific discovery.
Figures & tables
Figure 1. Overview of EvoScientist, a self-evolving multi-agent system for end-to-end scientific discovery. EvoScientist consists of a researcher agent (RA), an engineer agent (EA), and an evolution manager agent (EMA). The EMA distills interaction histories into two persistent memories, an ideation memory MI and an experimentation memory ME , which are retrieved by the RA and EA to enable continuous improvement in idea quality and execution success rates across tasks.
Novelty
Feasibility
Relevance
Clarity
Method
Win
Tie
Lose
Win
Tie
Lose
Win
Tie
Lose
Win
Tie
Lose
Avg. gap
Open-sourced Systems
EvoScientist vs Virtual Scientist
96.67 ∗
0 3.33
0 0.00
93.33 ∗
0 6.67
0 0.00
90.00 ∗
0 6.67
0 3.33
96.67 ∗
0 3.33
0 0.00
+93.34
EvoScientist vs AI-Researcher
96.67 ∗
0 3.33
0 0.00
90.00 ∗
0 0.00
0 10.00
86.67 ∗
10.00
0 3.33
93.34 ∗
0 3.33
0 3.33
+87.50
EvoScientist vs InternAgent
73.33 ∗
16.67
10.00
93.33 ∗
0 0.00
0 6.67
86.67 ∗
13.33
0 0.00
96.67 ∗
0 3.33
0 0.00
+83.33
EvoScientist vs AI Scientist-v2
63.33 ∗
16.67
20.00
53.33 ∗
0 6.67
40.00
36.67 ∗
50.00
13.33
56.67 ∗
23.33
20.00
+29.17
Table 1. Comparison of EvoScientist with baseline systems on scientific idea generation, evaluated by Gemini-3-flash. The scores marked with ∗ mean EvoScientist outperforms the baseline significantly with p -value < 0.05 (sign. test).
Novelty
Feasibility
Relevance
Clarity
Method
Win
Tie
Lose
Win
Tie
Lose
Win
Tie
Lose
Win
Tie
Lose
Avg. gap
Open-sourced Systems
EvoScientist vs InternAgent
66.67 ∗
23.33
10.00
96.67 ∗
0 3.33
0 0.00
90.00 ∗
0 0.00
10.00
93.33 ∗
0 6.67
0 0.00
+84.17
EvoScientist vs AI Scientist-v2
73.33 ∗
10.00
16.67
50.00 ∗
16.67
33.33
43.33 ∗
50.00
6.67
53.33 ∗
20.00
26.67
+34.16
Commercial Systems
EvoScientist vs Novix
93.33 ∗
0 0.00
0 6.67
56.67 ∗
6.66
36.67
36.67 ∗
60.00
0 3.33
73.33 ∗
10.00
16.67
+49.17
Table 2. Comparison of EvoScientist with baseline systems on scientific idea generation, evaluated by human experts. The scores marked with ∗ mean EvoScientist outperforms the baseline significantly with p -value < 0.05 (sign. test).
Figure 2. Mean execution success rate across four experiment stages, before and after experiment strategy evolution (ESE).
Novelty
Feasibility
Relevance
Clarity
Method
Win
Tie
Lose
Win
Tie
Lose
Win
Tie
Lose
Win
Tie
Lose
Avg. gap
-IDE vs EvoScientist
16.67
16.67
66.67
20.00
30.00
50.00
23.33
50.00
26.67
23.33
46.67
30.00
-22.50
-IVE vs EvoScientist
30.00
26.67
43.33
10.00
26.67
63.33
30.00
46.67
23.33
16.67
46.67
36.67
-20.00
-all vs EvoScientist
10.00
10.00
80.00
0 3.33
13.33
83.33
16.67
46.67
36.67
20.00
46.67
33.33
-45.83
Table 3. Ablation study on scientific idea generation, evaluated by Gemini-3-flash.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3. Research queries used for evaluation.
Figure 4. Part 1. Prompt for LLM-based idea generation evaluation.
Figure 5. Part 2. Prompt for LLM-based idea generation evaluation.
Figure 6. Part 3. Prompt for LLM-based idea generation evaluation.
Figure 7. Part 4. Prompt for LLM-based idea generation evaluation.
Figure 8. Part 5. Prompt for LLM-based idea generation evaluation.
Figure 9. Instructions for human evaluation of idea generation.
Figure 10. Prompts for idea direction evolution.
Figure 11. Prompts for idea validation evolution.
Figure 12. Prompts for experiment strategy evolution.
Title
Review Results
Adaptive Evidential Meta-Learning with Hyper-Conditioned Priors for Calibrated ECG Personalisation
Best Paper Award
Hierarchical Change Signature Analysis: A Framework for Online Discrimination of Incipient Faults and Benign Drifts in Industrial Time Series
AI Reviewer’s Appraisal Award
Robust Zero-Shot NER for Crises via Iterative Knowledge Distillation and Confidence-Gated Induction
Accepted
Adaptive Log Anomaly Detection through Data–Centric Drift Characterization and Policy-Driven Lifelong Learning
Accepted
ConFIT: A Robust Knowledge-Guided Contrastive Framework for Financial Extraction
Accepted
Hierarchical Adaptive Normalization: A Placement-Conditioned Cascade for Robust Wearable Activity Recognition
Figure 13. Review evidence for Adaptive Evidential Meta-Learning with Hyper-Conditioned Priors for Calibrated ECG Personalisation (Best Paper Award, AI Scientist Track). Original meta-review and decision page: Airaxiv link .
Figure 14. Review evidence for Hierarchical Adaptive Normalization: A Placement-Conditioned Cascade for Robust Wearable Activity Recognition (AI Reviewer’s Appraisal Award, AI Scientist Track). Original meta-review and decision page: Airaxiv link .
Large language models (LLMs), have shown strong potential in scientific discovery, yet existing methods still face substantial challenges in the design of research workflows and multi-role collaboration mechanisms. To mitigate these issues, we propose EvoSci, a multi-agent scientific collaboration framework, which integrates bio-inspired evolution with knowledge graph modeling. To iteratively generate, evaluate, and refine research ideas, EvoSci incorporates multiple role-based agents, including mentor, researcher, and reviewer. By combining collaborative reasoning, shared memory, and evolutionary feedback, EvoSci significantly enhances the coherence and creativity of scientific exploration. Experiments on real-world research topics demonstrate that EvoSci significantly outperforms strong baselines in LLM-based structured peer-review and comparative ranking evaluations, achieving the highest overall peer-review score (ICLR 4.90) and top ranking (Top-10 = 54). These results suggest its superiority in both scientific idea generation and continuous discovery.
Xiaoyu Xiong, Yuqi Ren, Deyi Xiong
TJUNLP Lab, School of Computer Science and Technology, Tianjin University, China
Scientific research proceeds through iterative cycles of hypothesis generation, experiment design, execution, and revision. AI agents can automate parts of this process, but existing approaches typically follow a single research trajectory or coordinate through a central planner with fixed objectives. As a result, they struggle to sustain parallel exploration, adapt as experimental evidence changes, or preserve knowledge of failed directions over long-running experiments. We introduce AutoScientists, a decentralized team of AI agents for long-running computational scientific experimentation. Agents interpret a shared experimental state, self-organize into teams around promising hypotheses, critique proposals before using experimental compute, and share successes and failures to reduce redundant exploration. Under matched experimental budgets, AutoScientists improves over prior AI agents across biomedical machine learning, language-model training optimization, and protein fitness prediction. On BioML-Bench, spanning biomedical imaging, protein engineering, single-cell omics, and drug discovery, AutoScientists achieves a mean leaderboard percentile of 74.4% across 24 tasks, improving over the strongest AI agent by +8.33%. On GPT training optimization, AutoScientists reaches a target validation bits-per-byte 1.9x faster than Autoresearch and continues discovering improvements from a starting champion where the single-agent approach finds none (7 vs. 0 accepted improvements). On ProteinGym fitness prediction, AutoScientists discovers a method for ACE2-Spike binding that improves over the current state-of-the-art model by +12.5% in Spearman correlation. Applied without modification across all 217 ProteinGym assays, the same method improves over the prior state of the art by +6.5% (Spearman correlation).
The convergence of large language models and agents is catalyzing a new era of scientific discovery: Agentic Science. However, common agent infrastructure is repeatedly rebuilt across scientific fields (loop fragmentation) and useful evidence and experience are lost in long-horizon research (loop discontinuity), bringing obstacles to Agentic Science at Scale. We introduce EvoMaster, a foundational evolving agent framework for Agentic Science at Scale. EvoMaster handles loop fragmentation and loop discontinuity by implementing Loop Research in which external evidence persists and improves later decisions. Through three nested loops, Execution, Exploration and Evolution (E3), loop research connects research within runs, across experiments and across studies. Across ten benchmarks spanning scientific research, coding and reasoning, EvoMaster achieves the best score among four agents using GPT-5.4, reaching a mean score of 58.02%, and outperforms the strongest competing agent Codex(40.29%) while costing 35.6% less. These results show that a shared loop-research foundation can support diverse scientific agents at scale.
Xinyu Zhu, Yuzhu Cai, Zexi Liu +20
School of Artificial Intelligence, Shanghai Jiao Tong University, Shanghai, China. · SciLand, Shanghai, China. · DP Technology, Beijing, China.