NGN: Learning Neural Network Size as a Differentiable Count
Organizations: Cornell University Ithaca, NY 14853
Abstract
Neural network size is usually chosen before training, separating architecture selection from weight optimization. We introduce the Neurogenesis Network (NGN), a differentiable parameterization for learning how many ordered structural components a model should use. For each ordered component group, one learnable boundary selects an active prefix while the model parameters are trained. The boundary can grow from a compact initialization and can be deployed by discarding components beyond the learned boundary. Controlled experiments examine convergence of the learned boundary, the performance of deployed prefixes, and comparisons with fixed-size models and alternative approaches to learning capacity. We then apply the same mechanism to MLPs, convolutional and graph networks, Transformers, state-space models, LoRA, and adapters. Across these settings, deploying only the learned prefix usually changes performance little, and the selected architectures perform similarly to fixed models trained at the same size. These results show that structural capacity can be optimized directly as a count.
Figures & tables
| setting | integer | soft MSE | truncated MSE | |
|---|---|---|---|---|
| Schedule robustness | ||||
| baseline | 6.00 | 100% | 1.31e-05 | 1.31e-05 |
| const | 5.45 | 0% | 1.84e-05 | 1.84e-05 |
| const | 7.00 | 100% | 1.41e-05 | 1.41e-05 |
| Component ablations | ||||
| 4.49 | 0% | 5.43e-05 | 4.83e-04 | |
| true | learned | exact rate | soft NMSE | trunc. NMSE |
|---|---|---|---|---|
| 1 | 1.00 | 100% | 9.94e-07 | 9.94e-07 |
| 2 | 2.00 | 100% | 2.39e-06 | 2.39e-06 |
| 4 | 4.00 | 100% | 1.14e-05 | 1.14e-05 |
| 6 | 6.00 | 100% | 5.55e-05 | 5.55e-05 |
| 8 | 8.00 | 90% | 1.57e-04 | 1.57e-04 |
| task | learned | integer | soft | trunc | fixed | prunes back? | |
|---|---|---|---|---|---|---|---|
| MNIST | tanh MLP features | [8.00] | 100% | 96.93 | 96.93 | 96.81 | yes |
| MNIST | conv branches | [4.00] | 100% | 99.18 | 99.18 | 99.13 | partial |
| CIFAR-10 | conv branches | [7.02] | 50% | 71.77 | 71.77 | 72.03 | partial |
| Cora | 2-hop GCN branches | [1.00] | 100% | 79.05 | 79.05 | 79.45 | partial |
| TinyStories | attention heads ( layers) | [2.11, 3.03] | 0% | 1.3055 | 1.3234 | 1.3011 | no |
| TinyStories | FFN blocks ( layers) | [1.14, 8.00] | 50% | 1.3454 | 1.3775 | 1.3402 | partial |
| construction / setting | learned | depth | soft MSE | trunc. MSE | fixed MSE |
|---|---|---|---|---|---|
| per-layer: const- , const- | [3,4,4,3,3,1,3,2,2,6] | 10 | 2.28e-02 | 2.28e-02 | 1.72e-02 |
| per-layer: const- , inc- | [4,2,3,3,1,3,2,3,4,9] | 10 | 7.37e-04 | 7.74e-04 | 6.98e-04 |
| per-layer: dec- , const- | [0,0,0,4,5,5,5,8,9,9] | 7 | 1.57e-02 | 1.58e-02 | 3.28e-02 |
| per-layer: dec- , inc- | [0,0,0,0,2,3,7,9,9,9] | 6 | 2.51e-04 | 3.30e-04 | 3.78e-03 |
| recursive | 7.02 | 8 | 2.49e-05 | 3.56e-05 | 1.54e-05 |
| LoRA rank | ||||||
|---|---|---|---|---|---|---|
| task | base | NGN rank (e/m/l) | NGN soft/trunc. | fixed | SoRA | AdaLoRA |
| WikiSQL-easy | 3.5 | 19 (2/9/8) | 78.2 / 78.6 | 77.7 | 75.2 | 45.1 |
| WikiSQL-hard | 2.0 | 21 (4/11/6) | 70.8 / 70.4 | 68.7 | 67.4 | 45.4 |
| Spider | 1.1 | 17 (5/7/5) | 12.6 / 12.4 | 12.3 | 13.9 | 10.3 |
| GSM8K | 45.2 | 18 (7/8/3) | 53.6 / 55.2 | 50.8 | 35.6 | 56.8 |
| SQuAD | 65.2 | 13 (2/7/4) | 67.2 / 67.3 | 67.1 | 67.1 | 66.9 |
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| setting | soft MSE | truncated MSE | |
|---|---|---|---|
| Schedule robustness | |||
| baseline | 6.10 1.45 | 9.32e-04 2.3e-03 | 9.32e-04 2.3e-03 |
| const | 5.75 1.83 | 3.12e-04 4.9e-04 | 3.12e-04 4.9e-04 |
| const | 7.30 0.78 | 2.24e-04 3.2e-04 | 2.24e-04 3.2e-04 |
| Component ablations | |||
| 4.70 1.49 | 3.25e-04 4.9e-04 | 1.45e-03 2.6e-03 | |
| task | learned | soft acc. | trunc. acc. | fixed acc. | prunes back? | |
|---|---|---|---|---|---|---|
| MNIST | tanh MLP features | [8.00 0.00] | 96.94 0.07 | 96.94 0.07 | 96.81 0.09 | yes ( [8.20 0.40] ; Y/P/N=10/0/0) |
| MNIST | conv branches | [4.10 0.30] | 99.17 0.03 | 99.17 0.03 | 99.14 0.04 | partial ( [5.60 0.49] ; Y/P/N=5/4/1) |
| CIFAR-10 | conv branches | [7.41 0.49] | 71.74 0.26 | 71.74 0.26 | 72.08 0.29 | partial ( [10.00 0.00] ; Y/P/N=0/10/0) |
| Cora | 2-hop GCN branches | [1.20 0.40] | 78.68 1.66 | 78.68 1.66 | 79.43 1.18 | partial ( [2.60 0.49] ; Y/P/N=5/5/0) |
| true | learned | exact rate | soft NMSE | trunc. NMSE |
|---|---|---|---|---|
| 1 | 1.00 0.00 | 100% | 9.79e-07 5.7e-08 | 9.79e-07 5.7e-08 |
| 2 | 2.00 0.00 | 100% | 1.39e-05 1.5e-05 | 1.36e-03 2.5e-03 |
| 4 | 4.00 0.00 | 100% | 2.86e-05 3.8e-05 | 1.18e-04 3.0e-04 |
| 6 | 6.00 0.00 | 100% | 6.79e-05 4.3e-05 | 6.79e-05 4.3e-05 |
| 8 | 7.90 0.30 | 90% | 1.25e-03 2.9e-03 | 1.25e-03 2.9e-03 |
| study | sched. | sched. | |||||
| 1D standard | 32 | power-decay | delayed linear | ||||
| 1D sweep | 32 | power-decay | delayed linear | ||||
| initialization sweep | 32 | power-decay | delayed linear | ||||
| schedule ablations | 32 | Table 1 | Table 1 | Table 1 | |||
| capacity baselines | 32 | selected | method-specific | method-specific | method-specific | method-specific | method-specific |
| manifold rank | 10 | linear decay | delayed linear |
| task | sched. | sched. | |||||
|---|---|---|---|---|---|---|---|
| MNIST MLP | 32 | power-decay | delayed linear | ||||
| MNIST / CIFAR-10 CNN | 8 / 16 | power-decay | delayed linear | ||||
| Cora | 8 | power-decay | delayed linear | ||||
| TinyStories attention | 8/layer | power-decay | delayed linear | ||||
| TinyStories FFN | 16/layer | power-decay | delayed linear | ||||
| TinyStories joint | 8/16/layer | power-decay | delayed linear |
| setting | soft MSE | hard MSE | fixed MSE | ||
|---|---|---|---|---|---|
| standard ( , learn ) | 0.50 | 7.00 | 6.27e-06 | 6.28e-06 | 6.28e-06 |
| learn and | 0.81 0.04 | 5.77 1.42 | 4.84e-05 5.7e-05 | 3.01e-02 6.3e-02 | 8.59e-05 2.1e-04 |
| learn , norm., fixed | 0.50 | 3.42 1.32 | 6.36e-02 4.7e-02 | 9.03e-02 1.1e-01 | 2.30e-05 2.6e-05 |
| learn only | 0.81 0.02 | – | 1.26e-06 9.8e-07 | 8.23e+00 2.3e+00 | 4.12e-05 8.5e-05 |
| learn , normalized | 0.89 0.04 | – | 2.75e-02 1.4e-02 | 2.75e-01 1.5e-01 | 4.05e-05 7.8e-05 |
| task | base | size | e/m/l | integer | NGN soft/trunc. | fixed | NGN ppl/base | fixed ppl/base |
|---|---|---|---|---|---|---|---|---|
| LoRA rank | ||||||||
| WikiSQL-easy | 3.5 | 19 | 2/9/8 | 20% | 78.2 / 78.6 | 77.7 | 1.016 | 0.996 |
| WikiSQL-hard | 2.0 | 21 | 4/11/6 | 28% | 70.8 / 70.4 | 68.7 | 1.022 | 0.999 |
| Spider | 1.1 | 17 | 5/7/5 | 58% | 12.6 / 12.4 | 12.3 | 0.997 | 0.984 |
| GSM8K | 45.2 | 18 | 7/8/3 | 62% | 53.6 / 55.2 | 50.8 | 0.968 | 0.973 |
| SQuAD | 65.2 | 13 | 2/7/4 | 34% | 67.2 / 67.3 | 67.1 | 0.993 | 0.980 |
| task | method | learned count | e/m/l | score | ppl base |
|---|---|---|---|---|---|
| WikiSQL-easy | NGN | 19 | 2/9/8 | 78.6 | 1.016 |
| WikiSQL-easy | SoRA | 19 | 7/12/0 | 75.2 | 1.009 |
| WikiSQL-easy | AdaLoRA | 19 | 3/13/3 | 45.1 | 1.087 |
| WikiSQL-hard | NGN | 21 | 4/11/6 | 70.4 | 1.022 |
| WikiSQL-hard | SoRA | 21 | 9/12/0 | 67.4 | 1.000 |
| WikiSQL-hard | AdaLoRA | 21 | 5/14/2 | 45.4 | 1.063 |
| dataset | use | stated license / source |
|---|---|---|
| MNIST | image classification | not specified on the original distribution page |
| CIFAR-10 | image classification | not specified on the original distribution page |
| Cora | graph node classification | not specified by the Planetoid distribution documentation |
| TinyStories | language modeling | CDLA-Sharing-1.0 on the dataset card |
| WikiSQL | text-to-SQL | BSD-3-Clause in the official repository |
| Spider | text-to-SQL | CC BY-SA 4.0 on the dataset card |
| task-only | mixed | |||||||
|---|---|---|---|---|---|---|---|---|
| method | kept | e/m/l (%) | score | ppl | kept | e/m/l (%) | score | ppl |
| NGN prune | 14.1% | 19.9/10.8/11.2 | 62.5 | 639 | 34.2% | 36.4/35.8/30.0 | 70.0 | 42 |
| random drop | 14.1% | 19.9/10.8/11.2 | 17.0 | 7.0e4 | 34.2% | 36.4/35.8/30.0 | 17.5 | 1.6e3 |
| Wanda (per-layer) | 14.1% | 19.9/10.8/11.2 | 69.5 | 356 | 34.2% | 36.4/35.8/30.0 | 73.0 | 45 |
| target | backbone | kept | e/m/l (%) | EM (trunc) | Pile base | |
|---|---|---|---|---|---|---|
| attention heads | frozen | 78.0% | 76.3/79.7/77.9 | 41.5 | 9.61 | |
| attention heads | frozen | 91.2% | 91.4/91.9/90.4 | 50.5 | 1.80 | |
| attention heads | frozen | 94.9% | 94.0/94.3/96.4 | 57.0 | 1.50 | |
| attention heads | updated | 72.9% | 69.8/76.0/72.9 | 37.0 | 561.01 | |
| attention heads | updated | 84.7% | 85.2/86.7/82.3 | 54.5 | 219.20 | |
| attention heads | updated | 89.3% | 89.1/92.2/86.7 | 63.0 | 220.37 |