Bayesian Optimization in Sequence-to-Architecture Latent Space for Zero-Shot NAS
Organizations: Department of Cybernetics FEE Czech Technical University
Abstract
Zero-shot Neural Architecture Search removes the prohibitive cost of traditional NAS, but its search process is typically based on the evolutionary algorithm (EA); lacking an explicit model of the objective, it often resorts to a near-random search through mutation. Bayesian Optimization offers a principled alternative by modeling the objective and aggregating information across iterations, but scales poorly to the high-dimensional, discrete, graph-structured spaces of modern NAS, restricting its use to only small networks. In this paper, we bring Bayesian Optimization to zero-shot NAS for large-scale architectures by learning a latent space via a Variational Autoencoder trained to reconstruct a novel prefix encoding of architectures and propose a proxy scalarization that combines several zero-shot proxies into a single Bayesian Optimization objective. After only 10,000 iterations of the proposed search algorithm (8 hours on a single GPU), our method found a network architecture which under the given model parameter count constraints achieves state-of-the-art results on three separate tasks -- image classification, object detection and semantic segmentation.
Figures & tables
| Model | Params | FLOPs | Top-1 (%) |
|---|---|---|---|
| ResNet He et al. (2016) | 24.8M | 9.0G | 75.19 |
| EfficientNet Tan and Le (2019) | 24.0M | 6.0G | 80.52 |
| ViT Dosovitskiy et al. (2020) ; Dai et al. (2021) | 26.4M | 18.4G | 80.64 |
| CoAtNet Dai et al. (2021) | 26.9M | 8.44G | 80.58 |
| ZenNAS Lin et al. (2021) | 26.3M | 18.6G | 80.48 |
| AZ-NAS Lee and Ham (2024) | 26.7M | 19.0G | 80.39 |
| Method | Search strategy | Top-1 (%) | GPU days |
|---|---|---|---|
| -DARTS† Ye et al. (2022) | Differentiable (DARTS space) | 76.1 | 0.4 |
| GP-NAS† Li et al. (2020) | Bayesian Opt. w/ hand-crafted encoding | 73.4 | 0.9 |
| BONAS† Shi et al. (2020) | Bayesian Opt. w/ learned encoding | 75.4 | 10.0 |
| UniNAS-A Tybl and Neumann (2026) | Evolutionary Algorithm | 81.15 | 0.5 |
| ours | Bayesian Opt. w/ seq-to-arch latent enc. | 81.51 | 0.33 |
| Object Detection and Seg. | Semantic Segmentation | |||||||
|---|---|---|---|---|---|---|---|---|
| Model | Params | AP b | AP m | FPS | FLOPs | mIoU | FPS | FLOPs |
| ResNet He et al. (2016) | 24.8M | 37.7 | 34.7 | 102.8 | 184G | 39.3 | 481.2 | 47G |
| EfficientNet Tan and Le (2019) | 24.0M | 39.0 | 35.8 | 43.1 | 124G | 37.0 | 152.6 | 32G |
| CoAtNet Dai et al. (2021) | 26.9M | 41.3 | 38.4 | 14.4 | 296G | 42.4 | 61.7 | 51G |
| UniNAS-A Tybl and Neumann (2026) | 26.8M | 42.4 | 39.0 | 14.2 | 297G | 45.6 | 88.2 | 51G |
| ours | 27.0M | 42.9 | 39.8 | 13.5 | 305G | 46.1 | 62.4 | 51G |
| Variant | Search | Latent space | Proxy | Top-1 (%) | |
|---|---|---|---|---|---|
| Ours (full) | BO | prefix-based | scalarized | 81.5 | – |
| w/o latent space | EA | none | scalarized | 80.0 | |
| w/o proxy scalarization | BO | prefix-based | VKDNW (single) | 77.2 | |
| w/o prefix encoding | BO | graph-based (VGAE) | scalarized | diverged † | – |
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
| Letter | Node | Arity | Status | Description |
| Leaf nodes (no children) — 13 retained from UniNAS | ||||
| BatchNorm | – | used | batch normalisation | |
| LayerNorm | – | used | layer normalisation | |
| Conv1 | – | used | convolution | |
| Conv3 | – | used | convolution | |
| ConvDepth3 | – | used | depthwise convolution | |