Organizations: Department of Health Data Science and AI McWilliams School of Biomedical Informatics University of Texas Health Science Center at Houston
AI-driven scientific discovery accelerates research by autonomously developing solutions and designs. Large language model (LLM) agents support this process through iterative generation and evaluation. Yet these iterations alone do not ensure cumulative progress or establish which directions to pursue next. Costly evaluation further constrains the scope of exploration. Neural architecture discovery brings these challenges together, coupling open-ended design with resource-intensive experimentation. We introduce EvoTreeNAD, a genealogy-guided evolutionary algorithm that constructs trainable architectures without a supplied seed or a hand-specified search space. Starting from an empty root, it grows a persistent genealogy in which each new node represents a complete architecture. Top-percentile values computed from each node and its descendants guide lineage selection. Using the selected design history, an Idea Agent proposes a variant and a Code Agent implements it. Each evaluated variant becomes a child node, expanding the genealogy while providing evidence for subsequent lineage selection. Our theoretical analysis establishes the existence of stationary variation regimes as the genealogy grows. Under specified variation assumptions, sustained top-percentile family values quantify the probability of generating high-reward architectures in these regimes. EvoTreeNAD discovers architectures that outperform the compared NAS and NAD baselines, achieving CIFAR-10/100 test errors of 2.05±0.06% and 15.09±0.22%. On all six MedMNIST-v2 tasks, the discovered architectures surpass the strongest listed baselines. A controlled CIFAR-10 study further shows that EvoTreeNAD outperforms direct generation, best-of-N greedy continuation, and full-family-mean routing.
Figures & tables
Figure 1: Overview of EvoTreeNAD.
Approach
CIFAR-10
CIFAR-100
Method
Space
Top-1 err. (%) ↓
Params (M)
GD
Top-1 err. (%) ↓
Params (M)
GD
Classical NAS and differentiable NAS
NASNet-A ( Zoph et al., 2018 )
2.65
3.3
2000
N/R
N/R
N/R
RL
NASNet
ENAS ( Pham et al., 2018 )
2.89
4.6
0.45
N/R
N/R
N/R
RL
NASNet
AmoebaNet-B ( Real et al., 2019 )
2.55 ± 0.05
2.8
3150
N/R
N/R
N/R
EA
NASNet
Random Search ( Li & Talwalkar, 2019 )
2.85 ± 0.08
4.3
9.7
N/R
N/R
N/R
Random
DARTS
Table 1: CIFAR-10/100 test performance. For each dataset, we report the best-performing discovered architecture and a smaller competitive architecture from a different discovery run. EvoTreeNAD architectures are trained from scratch under fixed full-fidelity recipes; mean ± std is computed over four random seeds. Results for other methods are taken from their publications.
Strategy
Lineage-conditioned variation
Branch reconsideration
Selection signal
Run 1
Run 2
Run 3
Best
Mean
Repeated direct generation
–
–
–
96.94
96.59
96.60
96.94
96.71
Best-of- N greedy continuation
✓
–
Immediate child reward
96.34
96.92
97.06
97.06
96.77
Full-family mean
✓
✓
Mean-based family value
96.81
95.66
96.49
96.81
96.32
EvoTreeNAD
✓
✓
Top-percentile family value
97.39
97.26
97.11
97.39
97.25
Table 2: CIFAR-10 comparison of architecture-discovery strategies. Each strategy has three independent 100-iteration runs. The Run 1–3 columns report the highest full-fidelity accuracy (%) among three candidates selected by discovery reward within each run; Best and Mean summarize these three run-level accuracies.
Figure 2: Architecture development along EvoTreeNAD principal lineages on CIFAR-10.
Table 5
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Notation
Description
Φ
Rule mapping a descendant-family reward multiset to a family value.
ΩtΦ,GtΦ
Process state and its genealogy after t expansions under rule Φ .
HtΦ(v)
Reward multiset of node v and all its descendants in the current genealogy.
VtΦ(v)
Empirical family value Φ(HtΦ(v)) used for routing.
τ,Vttop(v)
Fixed top-percentile fraction and corresponding empirical family value.
LtΦ
Selected root-to-node lineage; its endpoint receives the next child.
Appendix
Table 5: Notation Table
Name
Symbol
Description
Representative setting
Child capacity
W(s)
Maximum number of direct children of node s . Capacities may differ across nodes; for example, W(ρ)=15 at the root and W(s)=7 at non-root nodes.
task-dep.
Code realization attempts
Amax
Maximum Code-Agent attempts within one proposed variation.
1–3
Rejection tolerance
Ptolerant
Failed expansions skipped per parent before unresolved failures are recorded.
4–8
Top-percentile fraction
τ
Fraction of the highest rewards averaged over a node and all its descendants (Eq. 3 ).
0.1–0.2
Idea ancestors
kidea
Maximum number of recent ancestors provided to the Idea Agent as context.
1–3
Code ancestors
kcode
Maximum number of recent ancestors provided to the Code Agent as context.
1–3
Appendix
Table 6: EvoTreeNAD hyperparameters, definitions, and representative settings. Task-dependent values vary by dataset and budget.
Dataset
Data Modality
Task Type
Input Shape
#(Train)
#(Val)
#(Test)
Total
PathMNIST
Colon Pathology (Histology)
Multi-Class (9)
3×64×64
89,996
10,004
7,180
107,180
OCTMNIST
Retinal OCT
Multi-Class (4)
1×64×64
97,477
10,832
1,000
109,309
TissueMNIST
Kidney Cortex Microscopy
Multi-Class (8)
1×64×64
165,466
23,640
47,280
236,386
VesselMNIST3D
Brain MRA (Shape)
Binary (2)
1×64×64×64
1,335
191
382
1,908
SynapseMNIST3D
Electron Microscopy
Binary (2)
1×64×64×64
1,230
177
352
1,759
OrganMNIST3D
Abdominal CT
Multi-Class (11)
1×64×64×64
971
161
610
1,742
Appendix
Table 7: MedMNIST-v2 Dataset Statistics.
Dataset
Epochs
Batch Size
Learning Rate
CIFAR-10
500
128
0.002
CIFAR-100
500
128
0.001
MedMNIST-v2 2D
60
160
0.001
MedMNIST-v2 3D
200
64
0.0005
CIFAR-10/100: DropPath, MixUp, and CutMix.
MedMNIST-v2 2D: DropPath and MixUp; 3D: DropPath.
Appendix
Table 8: Full-fidelity training settings for discovered models. All runs use the AdamW optimizer and a cosine-decay learning-rate schedule.
Task family
Representative discovered structures
CIFAR-10
Dynamic routing with stage transformers and cross-stage fusion; fractal MBConv with pyramid pooling.
CIFAR-100
Mixed-kernel MBConv with FFT context, ODE blocks, and cross-stage concatenation; Res2Net-style branches with gated expansion and channel attention.
MedMNIST-v2 2D
Selective-kernel layers with GeM pooling; inverted-residual hierarchies with squeeze-and-excitation.
MedMNIST-v2 3D
Depthwise-separable or MBConv-style 3D residual networks with task-specific frequency blocks, normalization, and anti-aliased downsampling.
Appendix
Table 9: Representative architectures discovered by EvoTreeNAD from an empty root.
Figure 3: Correlation between short-budget and full-fidelity performance on selected CIFAR-10 principal-lineage architectures.
Figure 4: Illustrative genealogies under the four discovery strategies. Colors represent reward levels.
Figure 5: Best full-fidelity accuracy per run for five Idea and Code Agent configurations on CIFAR-10. Each configuration has three independent runs.
Idea Agent
Code Agent
Attempts
Reject Rate
Early Stop Rate
Best Proxy Score
GPU Days
Idea Tokens (M)
Code Tokens (M)
Idea API cost (USD)
Code API cost (USD)
Total API cost (USD)
No
OSS20B
121.3
18.2%
47.6%
0.8571
0.28
0
1.18
0
0
0
No
GPT-5
115.3
12.7%
39.1%
0.8438
0.36
0
0.87
0
3.10
3.10
OSS20B
OSS20B
138.7
39.4%
28.8%
0.8658
0.35
0.45
1.57
0
0
0
GPT-4.1
OSS20B
144.0
37.5%
28.0%
0.8665
0.37
0.30
1.61
0.84
0
0.84
GPT-4.1
GPT-5
126.7
23.7%
36.5%
0.8747
0.27
0.25
0.90
0.70
3.04
3.74
Appendix
Table 10: Process statistics and costs across Idea and Code Agent configurations. Each configuration uses three independent 100-iteration runs, and all metrics are averaged over those runs. Code-Agent realizations, including retries, are counted as attempts. Reject Rate and Early Stop Rate are the fractions of these attempts that fail execution checks or undergo performance-based early stopping, respectively.
Current neural architecture search (NAS) methods are often limited by their predefined, restrictive search spaces. While recent large language model (LLM)-assisted NAS methods enable open-ended search spaces, they often suffer from inefficient exploration due to biased or low-quality design ideas. To address these issues, we propose to semi-automatically structure model design knowledge to guide the search process. Our approach first defines a high-level structural template of architectural attributes. An LLM then populates this template by analyzing papers, creating a rich and diverse search space that embodies this structured design knowledge. To efficiently explore this vast space, we introduce FairNAD, using a multi-type mutation that enables broad exploration through mutation with fair idea sampling, Pareto-aware mutation, LLM-driven iterative mutation, and a fine-grained feedback loop. We demonstrate the effectiveness of FairNAD in discovering high-performing architectures that yield 0.84, 2.17, and 2.35 points improvement on CIFAR-10, CIFAR-100, and ImageNet16-120, respectively, compared to current state-of-the-art methods.
Yuiko Sakuma, Masakazu Yoshimura, Marcel Gröpl +4
Sony Group Corporation, Tokyo, Japan · ETH Zurich, Switzerland
Neural Architecture Search (NAS) aims to automatically discover high-performing deep neural network (DNN) architectures. However, conventional algorithm-driven NAS relies on carefully hand-crafted search spaces to ensure executability, which restricts open-ended exploration. Recent coding-based agentic approaches using large language models (LLMs) reduce manual design, but current LLMs struggle to reliably generate complex, valid architectures, and their proposals are often biased toward a narrow set of patterns observed in their training data. To bridge reliable algorithmic search with powerful LLM assistance, we propose LLMasTool, a hierarchical tree-based NAS framework for stable and open-ended model evolution. Our method automatically extracts reusable modules from arbitrary source code and represents full architectures as hierarchical trees, enabling evolution through reliable tree transformations rather than code generation. At each evolution step, coarse-level planning is governed by a diversity-guided algorithm that leverages Bayesian modeling to improve exploration efficiency, while the LLM resolves the remaining degrees of freedom to ensure a meaningful evolutionary trajectory and an executable generated architecture. With this formulation, instead of fully agentic LLM approaches, our method explores diverse directions beyond the inherent biases in the LLM. Our method improves over existing NAS methods by 0.69, 1.83, and 2.68 points on CIFAR-10, CIFAR-100, and ImageNet16-120, demonstrating its effectiveness.
This paper focuses on a key challenge in Neural Architecture Search (NAS): integrating established architectural knowledge while exploring new designs under expensive evaluations. Large language models (LLMs) are a promising assistant for NAS because they can translate rich architectural and coding priors into executable code edits. However, in practice, seemingly local revisions often propagate into non-local behavioral and performance shifts because a single edit can inadvertently couple multiple interacting functional factors, a phenomenon we refer to as functional entanglement. To make LLM knowledge usable under such entanglement, we propose Structured Progressive Knowledge Activation (SPARK), which activates relevant priors by explicitly selecting the functional factor to modify and conditioning the edit on that factor. This factor-conditioned editing reduces entangled side effects and yields more targeted, reliable architecture modifications. On CLRS-DFS, SPARK achieves a 28.1x sample-efficient architecture evolution speedup and yields a 22.9% relative improvement in OOD accuracy. Our code is available at https://github.com/AIM-ResearchLab/SPARK.
Zhen Liu, Yuhan Liu, Jinjun Wang +3
State Key Laboratory of Human-Machine Hybrid Augmented Intelligence, Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University. · Zhongguancun Academy, Beijing, China. · MiLM Plus, Xiaomi Inc. +1