The progressive growth of neural networks requires deciding when the current representation remains sufficient for optimization and when it should be expanded. CAGE-NAS formulates this decision in function space through an admissibility criterion on approximations of the functional gradient. As long as a representation enables a certified Functional Gradient Descent step, the architecture remains fixed; when the criterion fails, a function-preserving expansion is applied and the resulting representation is evaluated again. As the main instance, we study the family induced by the tangent space, using a regularized projection of the functional gradient. In a controlled setting with exact certification, CAGE-NAS produces architectures positioned above the 99.8th performance percentile by held-out RMSE among all admissible alternatives within the same parameter budget, without enumerating them during the growth trajectory.
Figures & tables
Figure 1: CAGE-NAS discovers compact models in one train-and-grow trajectory. The percentile is measured by held-out RMSE against all screened models satisfying the strict 600 -parameter budget.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Seed
Architecture
Grow test RMSE
0
[16,19,9]
0.0642
1
[9,18,18]
0.0680
2
[13,16,15]
0.0585
3
[11,17,17]
0.0544
Appendix
Table 1: Final architectures and test RMSE values of the CAGE-NAS train-and-grow runs.
Figure 2: Training and architectural-growth behavior of CAGE-NAS on the controlled synthetic N=1024 benchmark for seed 0 over 70 epochs.
Figure 3: Test behavior of CAGE-NAS on the controlled synthetic N=1024 benchmark for seed 0 over 70 epochs.
Seed
Growth RMSE
Grid RMSE
0
0.0642
0.0538
1
0.0680
0.0674
2
0.0585
0.0589
3
0.0544
0.0641
Appendix
Table 2: Comparison between the final CAGE-NAS growth result and the candidate selected by minimum validation MSE from the saved strict-budget grid results for each seed. Lower test RMSE is better.
Figure 4: Distribution of the test score 1−RMSE for the screened fixed MLPs satisfying p(A)≤600 . Red stars indicate the corresponding CAGE-NAS train-and-grow runs; higher is better.
Seed
Architectures evaluated
Growth percentile
0
17,339
99.95
1
17,339
99.88
2
17,339
100.00
3
17,339
100.00
Appendix
Table 3: Position of the CAGE-NAS growth result within the distribution of architectures evaluated by grid search.
Seed
Architecture
LR
WD
Scheduler
Val RMSE
Fixed RMSE
Growth RMSE
0
16-19-9
0.05
0.0
cosineannealing
0.0719
0.0705
0.0642
1
9-18-18
0.04
0.001
cosineannealing
0.0660
0.0626
0.0680
2
13-16-15
0.08
0.0
cosineannealing
0.0719
0.0749
0.0585
3
11-17-17
0.08
0.001
cosineannealing
0.0687
0.0721
0.0544
Appendix
Table 4: Comparison between the final CAGE-NAS train-and-grow result and retraining the discovered fixed architecture using a hyperparameter grid search.
(X,Y)←\textscBuildCertificationSet(Dtr)
Appendix
Algorithm 1 Epoch-level certified adaptive training
Zero-cost proxies rank architectures cheaply, but their reliability varies across search spaces. We introduce CoRA-NAS (COarse Ranking + Anchor-residual), a two-stage framework combining a static ranking prior with low-cost learning-curve refinement. CoRA-Rank aggregates capacity and structure-at-initialization proxies through an equal-weight log-rank consensus and a target-free consensus gate. CoRA-Refine samples anchors across this prior, extrapolates their early validation curves, and propagates a learned residual correction with an ExtraTrees model. The refinement uses approximately 1% of the cost of fully training the candidate set. Fully trained architecture-accuracy labels are not used to fit the ranker. One configuration is used across spaces, with space-specific architecture encodings. Across NAS-Bench-201, NAS-Bench-101, TransNAS-Bench-101, and NATS-SSS, CoRA-Refine achieves mean Spearman correlations of 0.946, 0.715, 0.786, and 0.894, respectively. Its worst-space correlation of 0.715 is the highest among the compared methods. On NAS-Bench-201/CIFAR-100, its selected architecture reaches 73.32% accuracy, near the reported ground-truth best of 73.37%. On the pure size space, refinement recovers the static prior's shortfall relative to parameter count, while remaining tied with the strongest capacity proxies within noise. The resulting framework combines cross-space ranking robustness with low-cost architecture selection.
Yifan Yang, Zhaoyan Wang, Zheng Gao +2
University of New South Wales · Korea Advanced Institute of Science and Technology
Standard deep-learning pipelines usually choose the network architecture before training and keep it fixed throughout optimization. In contrast, a model can also be adapted by editing its structure during training, for example by pruning existing hidden-neuron units or growing new ones. Although growth is appealing for adaptive and continual systems, we show that it is not simply the inverse of pruning. Pruning selects among units that have participated in training from the start, whereas growth inserts new units into an already specialized optimization trajectory. We isolate this insertion problem and show that newborn units are often forward-active but backward-starved: they participate in the forward computation, yet receive much weaker gradient signal than incumbent units. This disadvantage is minor in small MLP benchmarks, but becomes clear in harder image-classification settings with a convolutional trunk. In these settings, \textsc{Grow} can achieve high final accuracy during the structural-editing procedure, while \textsc{Prune} is stronger when performance is averaged over the training trajectory or when the final sparse network is retrained from scratch. Interventions targeting optimizer state, insertion, selection, and trainability show that improving the integration of newborn units can improve adaptive performance, but does not automatically produce better final subnetworks. In continual-learning benchmarks stressing plasticity loss, \textsc{Grow} becomes competitive mainly when new units have enough time to integrate. Together, these results suggest that \textsc{Grow} should be evaluated not only as an architecture-search operator, but as a time-sensitive optimization process whose success depends on insertion stability.
Prediction-based approaches are widely used in neural architecture search (NAS), where a predictor estimates the performance of candidate architectures to guide selection. However, existing predictors are typically trained via supervised regression on limited samples, leading to overfitting and poor generalization to unseen architectures. In this work, we propose a fundamentally different formulation that models performance prediction as a conditional function inference problem using a Convolutional Neural Process (ConvNP) with meta-learning capabilities. Instead of fitting a fixed mapping to limited samples, our approach meta-learns to infer performance from partial observations by training with context-target splits across a group of synthesized tasks, explicitly optimizing for generalization under data scarcity and aligning the training procedure with the deployment setting in NAS. We further design simple yet effective meta-features for cell-based architectures and evaluate our method on NAS-Bench-101 and NAS-Bench-201. Extensive experiments show that our approach consistently improves top-K ranking quality and achieves the state-of-the-art architecture selection using limited samples.
Liping Deng, MingQing Xiao
Department of Mathematics University of California, Riverside Riverside, CA 92507 USA · School of Mathematical and Statistical Sciences Southern Illinois University Carbondale Carbondale, IL 62901 USA