The increasing adoption of Large Language Models (LLMs) has enabled AI scientists to perform complex end-to-end scientific discovery tasks requiring coordination of specialized roles, including idea generation and experimental execution. However, most state-of-the-art AI scientist systems rely on static, hand-designed pipelines and fail to adapt based on accumulated interaction histories. As a result, these systems overlook promising research directions, repeat failed experiments, and pursue infeasible ideas. To address this, we introduce EvoScientist, an evolving multi-agent AI scientist framework that continuously improves research strategies through persistent memory and self-evolution. EvoScientist comprises three specialized agents: a Researcher Agent (RA) for scientific idea generation, an Engineer Agent (EA) for experiment implementation and execution, and an Evolution Manager Agent (EMA) that distills insights from prior interactions into reusable knowledge. EvoScientist contains two persistent memory modules: (i) an ideation memory, which summarizes feasible research directions from top-ranked ideas while recording previously unsuccessful directions; and (ii) an experimentation memory, which captures effective data processing and model training strategies derived from code search trajectories and best-performing implementations. These modules enable the RA and EA to retrieve relevant prior strategies, improving idea quality and code execution success rates over time. Experiments show that EvoScientist outperforms 7 open-source and commercial state-of-the-art systems in scientific idea generation, achieving higher novelty, feasibility, relevance, and clarity via automatic and human evaluation. EvoScientist also substantially improves code execution success rates through multi-agent evolution, demonstrating persistent memory's effectiveness for end-to-end scientific discovery.
Figures & tables
Figure 1. Overview of EvoScientist, a self-evolving multi-agent system for end-to-end scientific discovery. EvoScientist consists of a researcher agent (RA), an engineer agent (EA), and an evolution manager agent (EMA). The EMA distills interaction histories into two persistent memories, an ideation memory MI and an experimentation memory ME , which are retrieved by the RA and EA to enable continuous improvement in idea quality and execution success rates across tasks.
Novelty
Feasibility
Relevance
Clarity
Method
Win
Tie
Lose
Win
Tie
Lose
Win
Tie
Lose
Win
Tie
Lose
Avg. gap
Open-sourced Systems
EvoScientist vs Virtual Scientist
96.67 ∗
0 3.33
0 0.00
93.33 ∗
0 6.67
0 0.00
90.00 ∗
0 6.67
0 3.33
96.67 ∗
0 3.33
0 0.00
+93.34
EvoScientist vs AI-Researcher
96.67 ∗
0 3.33
0 0.00
90.00 ∗
0 0.00
0 10.00
86.67 ∗
10.00
0 3.33
93.34 ∗
0 3.33
0 3.33
+87.50
EvoScientist vs InternAgent
73.33 ∗
16.67
10.00
93.33 ∗
0 0.00
0 6.67
86.67 ∗
13.33
0 0.00
96.67 ∗
0 3.33
0 0.00
+83.33
EvoScientist vs AI Scientist-v2
63.33 ∗
16.67
20.00
53.33 ∗
0 6.67
40.00
36.67 ∗
50.00
13.33
56.67 ∗
23.33
20.00
+29.17
Table 1. Comparison of EvoScientist with baseline systems on scientific idea generation, evaluated by Gemini-3-flash. The scores marked with ∗ mean EvoScientist outperforms the baseline significantly with p -value < 0.05 (sign. test).
Novelty
Feasibility
Relevance
Clarity
Method
Win
Tie
Lose
Win
Tie
Lose
Win
Tie
Lose
Win
Tie
Lose
Avg. gap
Open-sourced Systems
EvoScientist vs InternAgent
66.67 ∗
23.33
10.00
96.67 ∗
0 3.33
0 0.00
90.00 ∗
0 0.00
10.00
93.33 ∗
0 6.67
0 0.00
+84.17
EvoScientist vs AI Scientist-v2
73.33 ∗
10.00
16.67
50.00 ∗
16.67
33.33
43.33 ∗
50.00
6.67
53.33 ∗
20.00
26.67
+34.16
Commercial Systems
EvoScientist vs Novix
93.33 ∗
0 0.00
0 6.67
56.67 ∗
6.66
36.67
36.67 ∗
60.00
0 3.33
73.33 ∗
10.00
16.67
+49.17
Table 2. Comparison of EvoScientist with baseline systems on scientific idea generation, evaluated by human experts. The scores marked with ∗ mean EvoScientist outperforms the baseline significantly with p -value < 0.05 (sign. test).
Figure 2. Mean execution success rate across four experiment stages, before and after experiment strategy evolution (ESE).
Novelty
Feasibility
Relevance
Clarity
Method
Win
Tie
Lose
Win
Tie
Lose
Win
Tie
Lose
Win
Tie
Lose
Avg. gap
-IDE vs EvoScientist
16.67
16.67
66.67
20.00
30.00
50.00
23.33
50.00
26.67
23.33
46.67
30.00
-22.50
-IVE vs EvoScientist
30.00
26.67
43.33
10.00
26.67
63.33
30.00
46.67
23.33
16.67
46.67
36.67
-20.00
-all vs EvoScientist
10.00
10.00
80.00
0 3.33
13.33
83.33
16.67
46.67
36.67
20.00
46.67
33.33
-45.83
Table 3. Ablation study on scientific idea generation, evaluated by Gemini-3-flash.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3. Research queries used for evaluation.
Figure 4. Part 1. Prompt for LLM-based idea generation evaluation.
Figure 5. Part 2. Prompt for LLM-based idea generation evaluation.
Figure 6. Part 3. Prompt for LLM-based idea generation evaluation.
Figure 7. Part 4. Prompt for LLM-based idea generation evaluation.
Figure 8. Part 5. Prompt for LLM-based idea generation evaluation.
Figure 9. Instructions for human evaluation of idea generation.
Figure 10. Prompts for idea direction evolution.
Figure 11. Prompts for idea validation evolution.
Figure 12. Prompts for experiment strategy evolution.
Title
Review Results
Adaptive Evidential Meta-Learning with Hyper-Conditioned Priors for Calibrated ECG Personalisation
Best Paper Award
Hierarchical Change Signature Analysis: A Framework for Online Discrimination of Incipient Faults and Benign Drifts in Industrial Time Series
AI Reviewer’s Appraisal Award
Robust Zero-Shot NER for Crises via Iterative Knowledge Distillation and Confidence-Gated Induction
Accepted
Adaptive Log Anomaly Detection through Data–Centric Drift Characterization and Policy-Driven Lifelong Learning
Accepted
ConFIT: A Robust Knowledge-Guided Contrastive Framework for Financial Extraction
Accepted
Hierarchical Adaptive Normalization: A Placement-Conditioned Cascade for Robust Wearable Activity Recognition
Figure 13. Review evidence for Adaptive Evidential Meta-Learning with Hyper-Conditioned Priors for Calibrated ECG Personalisation (Best Paper Award, AI Scientist Track). Original meta-review and decision page: Airaxiv link .
Figure 14. Review evidence for Hierarchical Adaptive Normalization: A Placement-Conditioned Cascade for Robust Wearable Activity Recognition (AI Reviewer’s Appraisal Award, AI Scientist Track). Original meta-review and decision page: Airaxiv link .