EvoTreeNAD: Genealogy-Guided Evolution for LLM-Driven Neural Architecture Discovery
Organizations: Department of Health Data Science and AI McWilliams School of Biomedical Informatics University of Texas Health Science Center at Houston
Abstract
AI-driven scientific discovery accelerates research by autonomously developing solutions and designs. Large language model (LLM) agents support this process through iterative generation and evaluation. Yet these iterations alone do not ensure cumulative progress or establish which directions to pursue next. Costly evaluation further constrains the scope of exploration. Neural architecture discovery brings these challenges together, coupling open-ended design with resource-intensive experimentation. We introduce EvoTreeNAD, a genealogy-guided evolutionary algorithm that constructs trainable architectures without a supplied seed or a hand-specified search space. Starting from an empty root, it grows a persistent genealogy in which each new node represents a complete architecture. Top-percentile values computed from each node and its descendants guide lineage selection. Using the selected design history, an Idea Agent proposes a variant and a Code Agent implements it. Each evaluated variant becomes a child node, expanding the genealogy while providing evidence for subsequent lineage selection. Our theoretical analysis establishes the existence of stationary variation regimes as the genealogy grows. Under specified variation assumptions, sustained top-percentile family values quantify the probability of generating high-reward architectures in these regimes. EvoTreeNAD discovers architectures that outperform the compared NAS and NAD baselines, achieving CIFAR-10/100 test errors of and . On all six MedMNIST-v2 tasks, the discovered architectures surpass the strongest listed baselines. A controlled CIFAR-10 study further shows that EvoTreeNAD outperforms direct generation, best-of- greedy continuation, and full-family-mean routing.
Figures & tables
| Approach | CIFAR-10 | CIFAR-100 | Method | Space | ||||
|---|---|---|---|---|---|---|---|---|
| Top-1 err. (%) | Params (M) | GD | Top-1 err. (%) | Params (M) | GD | |||
| Classical NAS and differentiable NAS | ||||||||
| NASNet-A ( Zoph et al., 2018 ) | 2.65 | 3.3 | 2000 | N/R | N/R | N/R | RL | NASNet |
| ENAS ( Pham et al., 2018 ) | 2.89 | 4.6 | 0.45 | N/R | N/R | N/R | RL | NASNet |
| AmoebaNet-B ( Real et al., 2019 ) | 2.55 0.05 | 2.8 | 3150 | N/R | N/R | N/R | EA | NASNet |
| Random Search ( Li & Talwalkar, 2019 ) | 2.85 0.08 | 4.3 | 9.7 | N/R | N/R | N/R | Random | DARTS |
| Strategy | Lineage-conditioned variation | Branch reconsideration | Selection signal | Run 1 | Run 2 | Run 3 | Best | Mean |
|---|---|---|---|---|---|---|---|---|
| Repeated direct generation | – | – | – | 96.94 | 96.59 | 96.60 | 96.94 | 96.71 |
| Best-of- greedy continuation | ✓ | – | Immediate child reward | 96.34 | 96.92 | 97.06 | 97.06 | 96.77 |
| Full-family mean | ✓ | ✓ | Mean-based family value | 96.81 | 95.66 | 96.49 | 96.81 | 96.32 |
| EvoTreeNAD | ✓ | ✓ | Top-percentile family value | 97.39 | 97.26 | 97.11 | 97.39 | 97.25 |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Notation | Description |
|---|---|
| Rule mapping a descendant-family reward multiset to a family value. | |
| Process state and its genealogy after expansions under rule . | |
| Reward multiset of node and all its descendants in the current genealogy. | |
| Empirical family value used for routing. | |
| Fixed top-percentile fraction and corresponding empirical family value. | |
| Selected root-to-node lineage; its endpoint receives the next child. |
| Name | Symbol | Description | Representative setting |
|---|---|---|---|
| Child capacity | Maximum number of direct children of node . Capacities may differ across nodes; for example, at the root and at non-root nodes. | task-dep. | |
| Code realization attempts | Maximum Code-Agent attempts within one proposed variation. | 1–3 | |
| Rejection tolerance | Failed expansions skipped per parent before unresolved failures are recorded. | 4–8 | |
| Top-percentile fraction | Fraction of the highest rewards averaged over a node and all its descendants (Eq. 3 ). | 0.1–0.2 | |
| Idea ancestors | Maximum number of recent ancestors provided to the Idea Agent as context. | 1–3 | |
| Code ancestors | Maximum number of recent ancestors provided to the Code Agent as context. | 1–3 |
| Dataset | Data Modality | Task Type | Input Shape | #(Train) | #(Val) | #(Test) | Total |
|---|---|---|---|---|---|---|---|
| PathMNIST | Colon Pathology (Histology) | Multi-Class (9) | 89,996 | 10,004 | 7,180 | 107,180 | |
| OCTMNIST | Retinal OCT | Multi-Class (4) | 97,477 | 10,832 | 1,000 | 109,309 | |
| TissueMNIST | Kidney Cortex Microscopy | Multi-Class (8) | 165,466 | 23,640 | 47,280 | 236,386 | |
| VesselMNIST3D | Brain MRA (Shape) | Binary (2) | 1,335 | 191 | 382 | 1,908 | |
| SynapseMNIST3D | Electron Microscopy | Binary (2) | 1,230 | 177 | 352 | 1,759 | |
| OrganMNIST3D | Abdominal CT | Multi-Class (11) | 971 | 161 | 610 | 1,742 |
| Dataset | Epochs | Batch Size | Learning Rate |
| CIFAR-10 | 500 | 128 | 0.002 |
| CIFAR-100 | 500 | 128 | 0.001 |
| MedMNIST-v2 2D | 60 | 160 | 0.001 |
| MedMNIST-v2 3D | 200 | 64 | 0.0005 |
| CIFAR-10/100: DropPath, MixUp, and CutMix. | |||
| MedMNIST-v2 2D: DropPath and MixUp; 3D: DropPath. | |||
| Task family | Representative discovered structures |
|---|---|
| CIFAR-10 | Dynamic routing with stage transformers and cross-stage fusion; fractal MBConv with pyramid pooling. |
| CIFAR-100 | Mixed-kernel MBConv with FFT context, ODE blocks, and cross-stage concatenation; Res2Net-style branches with gated expansion and channel attention. |
| MedMNIST-v2 2D | Selective-kernel layers with GeM pooling; inverted-residual hierarchies with squeeze-and-excitation. |
| MedMNIST-v2 3D | Depthwise-separable or MBConv-style 3D residual networks with task-specific frequency blocks, normalization, and anti-aliased downsampling. |
| Idea Agent | Code Agent | Attempts | Reject Rate | Early Stop Rate | Best Proxy Score | GPU Days | Idea Tokens (M) | Code Tokens (M) | Idea API cost (USD) | Code API cost (USD) | Total API cost (USD) |
|---|---|---|---|---|---|---|---|---|---|---|---|
| No | OSS20B | 121.3 | 18.2% | 47.6% | 0.8571 | 0.28 | 0 | 1.18 | 0 | 0 | 0 |
| No | GPT-5 | 115.3 | 12.7% | 39.1% | 0.8438 | 0.36 | 0 | 0.87 | 0 | 3.10 | 3.10 |
| OSS20B | OSS20B | 138.7 | 39.4% | 28.8% | 0.8658 | 0.35 | 0.45 | 1.57 | 0 | 0 | 0 |
| GPT-4.1 | OSS20B | 144.0 | 37.5% | 28.0% | 0.8665 | 0.37 | 0.30 | 1.61 | 0.84 | 0 | 0.84 |
| GPT-4.1 | GPT-5 | 126.7 | 23.7% | 36.5% | 0.8747 | 0.27 | 0.25 | 0.90 | 0.70 | 3.04 | 3.74 |