Few-shot meta-learning traditionally formulates task adaptation either as analytical gradient descent through unrolled computational graphs or as metric-based distance comparisons over flattened 1D fea- ture vectors, which either incur costly test-time backpropagation or discard native 2D spatial geometry. In this work, we propose METALEARNNCA, a decentralized framework that achieves few-shot adapta- tion through the dynamical interaction of coupled Neural Cellular Automata (NCAs) without computing analytical gradients during inference. MetaLearnNCA decomposes task adaptation into an Active- NCA, which executes task inference conditioned on a continuous 2D spatial memory grid termed the spatial program, and a learned Meta-NCA, which acts as a decentralized cellular optimizer by diffusing spatial error residuals across local neighborhoods to dynamically update this program. METALEARN- NCA is competitive against canonical meta-learners in-distribution (96.12% on Omniglot) with Out-Of- Distribution transfer gains on MNIST, KMNIST, and Fashion-MNIST transfer across 10 independent testing seeds across 1-, 5-, and 10-shot regimes (e.g., surpassing Prototypical Networks by +10.54% on 10-shot MNIST and a +3.87% gain on 10-shot Fashion-MNIST over FOMAML). Our results establish that robust, gradient-free learning-to-learn can emerge from decentralized cellular dynamics on non-von Neumann substrates.
Figures & tables
Figure 1 : The MetaLearnNCA interaction loop: Active-NCA computes forward task predictions on the support set, while Meta-NCA translates the residual error maps into an updated spatial program Stask .
Gradient-Free Baselines
Gradient-Based
Ours
Dataset
Shot
MatchingNet
RelationNet
ProtoNet
FOMAML
MetaLearnNCA
Omniglot (In-Dist)
1
63.47±7.79
84.84±3.74
92.26±3.15
87.84±3.67
87.95±4.16
5
75.47±4.54
92.73±3.09
96.40±3.09
96.93±1.41
95.20±2.13
10
77.40±6.37
93.20±2.65
97.40±1.27
98.80±0.94
96.12±0.70
MNIST
1
44.79±4.93
51.17±2.66
55.32±1.87
48.16±2.98
65.02±3.07
5
58.93±1.34
63.74±1.39
75.17±1.56
70.80±1.72
84.12±1.57
Table 1 : Few-shot adaptation performance (Accuracy % ± 95% CI) across 1-shot, 5-shot, and 10-shot regimes on in-domain Omniglot and cross-domain out-of-distribution transfer benchmarks. All models (except ours) employ the canonical Conv4 backbone on 28×28 inputs.
Figure 2 : Few-shot offline 10-shot adaptation trajectories of MetaLearnNCA evaluated on out-of-distribution character datasets across offline cellular optimization repeats ( k∈{0,…,5} ). Solid curves depict mean accuracy across 42 independent meta-training runs and 42 independent testing seeds, with shaded regions denoting 95% confidence intervals.
Figure 3 : Resilience of MetaLearnNCA to physical spatial damage: test accuracy as a function of circular noise patches ( r=4 px ) injected into the static program grid Stask .
Figure 4 : Representative spatial and program channels during MNIST inference. (see Appendix C for the complete 46-channel atlas).
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5 : Theoretical overview of the conditioning mechanism: the static spatial program S acts as a localized pre-activation bias, dynamically modulating information flow through the Active-NCA transition operator.
Figure 6 : Complete atlas of Active-NCA dynamic reasoning channels (Channels 10 through 31) during MNIST inference.
Figure 7 : Complete atlas of static spatial program channels ( Stask , Channels 0 through 23) written by the Meta-NCA following support adaptation.
Figure 8 : Ablation analysis on Omniglot: (a) test accuracy under Gaussian noise injection across dynamic reasoning channels ( 10…31 ), and (b) test accuracy when individual static program channels ( 0…23 ) are ablated to zero.
Figure 9 : Difference in Stask across 2 independent testing runs on the same training seed
Figure 10 : MetaLearnNCA performance on MNIST under different levels of additive random noise to Stask .
Figure 11 : MetaLearnNCA Stask under different levels of noise.
Figure 12 : Cross dataset transfer, each model is trained on one dataset and tested on all the other ones. 50 repeats per dataset with 95% CI intervals.
Figure 13 : Dataset surfaces for seed 0, trained on OMNIGLOT across a combination of K-repeats and K-shots.
Dataset
Shift = 1 px
Shift = 2 px
Shift = 3 px
Shift = 4 px
Omniglot (In-Dist)
93.47±4.51%
90.93±4.56%
83.33±4.86%
70.13±8.81%
MNIST
84.43±1.41%
82.93±2.62%
77.13±2.86%
59.01±3.37%
USPS
57.89±7.27%
51.67±6.85%
40.71±5.31%
30.83±2.51%
Fashion-MNIST
48.72±4.56%
42.85±4.01%
34.03±5.16%
24.72±4.66%
KMNIST
41.70±3.88%
37.47±4.02%
30.50±3.55%
22.73±3.75%
MedMNIST
18.55±8.05%
17.23±7.92%
14.98±6.98%
13.00±5.90%
Appendix
Table 2 : Robustness of MetaLearnNCA to discrete spatial coordinate shifts ( Δ∈{1,2,3,4} pixels) in the 5-shot, 10-way transfer regime. Results report mean classification accuracy ± 95% confidence intervals evaluated over 5 independent testing runs on full test splits.
Hyperparameter
ProtoNet
MatchingNet
RelationNet
FOMAML
Backbone
Conv4-64
Conv4-64
Conv4-64
Conv4-64
Parameters
≈115k
≈115k
≈225k
≈115k
Output Head
Metric Centroid
Cosine Attention
Relation MLP
Linear ( 64→10 )
BatchNorm Stats
Frozen Running
Frozen Running
Frozen Running
Shuffled Mini-Batch
Meta-Loss
Cross-Entropy
Neg. Log-Likelihood
Mean Squared Error
Cross-Entropy
Inner Steps ( Ktrain/Ktest )
–
–
–
3 / 5
Appendix
Table 3 : Architectural and optimization hyperparameters for canonical baselines.
Self-organization is an emergent property of life, driven by the collective behavior of individual components acting on local information. Biological neurons, through local interactions transmitted through synapses, are able to learn efficiently and can adapt their connections over an organism's lifespan. Motivated by these desirable properties of adaptability and local interaction, neural cellular automata (NCA) models have been successful at learning morphogenesis solely through local update rules, demonstrating stability over many updates and robustness to perturbations. In this work, we introduce Meta Neural Cellular Automata (MetaNCA), a framework that learns local rules which self-organize the weights of artificial neural networks. A learned rule network iteratively updates the weights of a task network using only local interactions on the computation graph. We propose a novel Weight Transformer architecture for the local rule network, which uses linear attention to aggregate signals from neighboring weights and hidden states. Once trained, the rule network generates task networks of diverse architectures without backpropagation. We show that MetaNCA generates weights for feedforward MLPs, CNNs, and ResNets on MNIST and CIFAR-100, scaling to networks of 2 million parameters. We further show that MetaNCA generalizes to architectures not seen during meta-training, and that architectural diversity in the training phase strengthens this generalization.
Meet Barot, Daniel Berenberg, Sina Khajehabdollahi
Mythos Scientific, New York, USA · Independent Scholar
Neural networks are typically adapted by computing gradients and updating model parameters. We investigate whether task-specific adaptation can instead emerge from a meta-learned self-organising process that requires no gradients at adaptation time. We instantiate this idea with a Neural Cellular Automaton in which locally interacting recurrent cells maintain both a recurrent state and a fast associative memory. During meta-training, backpropagation is used to learn the recurrent dynamics together with how the memory is read and written. Once training is complete, the slow model parameters remain fixed, and online adaptation occurs only through cellwise memory updates driven by local prediction errors and a delta rule. We evaluate whether the learned mechanism can adapt to semantically distinct held-out classification tasks. A single pass over the support data produces substantial improvements in held-out performance without gradient computation or parameter updates during adaptation, and the mechanism remains effective across large changes in the number of examples processed jointly. These results show that task-specific adaptation can be achieved through explicit fast-memory updates while keeping the slow model parameters fixed.
Many modern learning approaches are still struggling with spatial reasoning tasks, i.e. they lack the ability to utilize geometric information of perceived entities and their spatial relation to each other to solve problems. We introduce a novel Adaptive Neural Cellular Automata (aNCA) architecture which uses deformable convolutions to dynamically adapt the perceptive field and iteratively reason over 2D spatial relations on grid-like data structures (e.g. images). Empirical results on public benchmarks show state of the art comprehensible results with high generalization abilities for solving image based puzzles like Sudoku or finding the shortest path in a maze.
Martin Spitznagel, Janis Keuper
Institute for Machine Learning and Analytics (IMLA), Offenburg University, Germany · University of Mannheim, Germany