Authors: Mayalen Etcheverry, Pietro Miotti, Aidan Sirbu, Konstantin Schürholt, Mariia Drozdova, Arna Ghosh, Blaise Agüera y Arcas, James Manyika, +2 more
Organizations: Google Paradigms of Intelligence Team · School of Computer Science, McGill University · Mila - Quebec AI Institute · University of Geneva · Department of Neurology and Neurosurgery, McGill University · Montreal Neurological Institute, McGill University · Learning in Machines and Brains Program, CIFAR
Modern AI architectures used to solve visual reasoning tasks typically rely heavily on global connectivity and synchronization. As biological systems demonstrate, though, sophisticated computation can be performed in a more decentralized fashion. In this work, we test the reasoning capabilities of Neural Cellular Automata (NCAs), networks of recurrent cells that use strictly local connectivity and asynchronous updates. NCAs have been extensively studied in artificial life experiments, but it is unclear whether they can perform complex multi-step reasoning. We show that NCAs produce spatio-temporal dynamics capable of solving challenging visual reasoning tasks, including large mazes, Sudoku, and ARC-AGI-1. Furthermore, we provide evidence that NCAs generalize out-of-distribution when running with larger grids, longer rollouts, or parallel trials; and that the latter can be made more efficient via pruning of redundant trajectories. We find that these generalization capabilities depend on training with sample replay and stochastic perturbations, and that stochasticity remains beneficial at test time. Finally, we show that NCAs are robust reasoners capable of dynamically modulating compute to recover efficiently from damage, and that they can scale to solve reasoning in raw pixel space.
Figures & tables
Figure 1: Overview of reasoning with Neural Cellular Automata (NCA). Cells maintain structured states ( Cin , Cout , Chid ) on an H×W grid and update iteratively using local 3×3 perception and a shared update module, applied residually under a stochastic firing mask. Training unrolls grids from the dataset or a sample replay buffer for N steps, optimizing the NCA to solve the task.
Model (test config)
Params
FLOPs
Acc (%)
Maze-OOD (K=1)
DeepThink (S=13x13, D=100)
784K
160.5B
99.9
NCA (S=13x13, D=300)
10K
1.5B
100.0
DeepThink (S=59x59, D=1000)
784K
24.1T
97.3
NCA (S=59x59, D=2000)
10K
163.4B
100.0
DeepThink (S=201x201, D=2000)
784K
521.7T
74.0
Table 1: Overview of NCA performances and efficiency across reasoning benchmarks (see section B.5 for FLOPs estimation).
Figure 2: Emergent reasoning dynamics in NCAs. (A) Backtracking in mazes: cell activations explore paths in parallel, then locally back-propagate “waves” to prune dead-end paths, until converging to the final solution path. (B) Iterative trajectory refinement in Maze-Hard: cells rapidly resolve unambiguous segments ( t=15 ) until settling first on a valid, sub-optimal solution (1, blue), before discovering a shorter valid alternative (2), and ultimately stabilizing into an optimal solution (3, green). (C) Local spatial propagation of objects and colors in ARC-AGI-1: solving the task of coloring gray shapes (target) according to the top-left reference pattern (source) shows a continuous, step-by-step spatial diffusion of color features from source to target. (D) Backtracking in Sudoku-Extreme: on the hardest Sudoku boards ( tdoku difficulty 10; Figure S 16 B), cells demonstrate collective trial-and-error: they first reach a near-valid grid but with a few conflicting digits ( t=60 ), they escape that local minimum by temporarily increasing constraint violations ( t=116 ), and finally converging to a correct global consensus ( t=250 ).
Figure 3: Multi-task execution in a single NCA. When evaluating the pre-trained NCA on an unseen input ( “CA” in blue), conditioned under different task embeddings, the model applies the specific transformation corresponding to each task (filling enclosed and/or open regions with target color), demonstrating context-dependent execution rather than input memorization.
Figure 4: Test-time scaling improves NCA generalization on out-of-distribution (OOD) boards. (A) Parallel state-space exploration: running parallel, stochastic rollouts significantly improves OOD generalization on Sudoku. (B) Spatial substrate scaling: NCAs trained on 9×9 mazes generalize to up to 500× larger mazes given additional substrate and iterations. Mean curves over 3 test seeds.
Figure 5: Test-time compute scaling on ARC-AGI public evaluation set. Pass@ K scaling under the offline-first-augmentation policy (see Appendix B.3 and B.4 ), reaching 60.3% Pass@64 2 2 2 Using an NCA ensemble (NCA E) of 3 models independently fine-tuned further boosts performances to 63.0%. .
Figure 6: Scaling laws and iso-compute allocation of test-time pruning on Sudoku-Extreme. (A) Exact accuracy versus total FLOPs across rollout depths D for unpruned baselines and pruning divisors m (shaded bands indicate ±1 SD). (B) Allocation of fixed compute budgets to breadth (unpruned Dbase ) versus depth ( 2⋅Dbase with m=4 ; 4⋅Dbase with m=8 ). Reallocating compute into depth yields accuracy gains up to Dbase=128 .
Figure 7: Stochasticity and perturbations are beneficial both at training and inference for generalization. (A) Training ingredients ablation study: the use of asynchronous updates (in particular for Sudoku), stochastic perturbations (noise, damage and target swap) as well as sample replay during training are critical ingredients for generalization to out-of-distribution instances. Mean-std curves over 3 train seeds are displayed. (B) Noise-injection at test time: when adding noise of varying magnitude at test time, not only do the learned NCA rules show perfect robustness but noise even boosts performances across all benchmarks and noise scales. Metrics are averaged over 3 test seeds.
Figure 8: Robust adaptive compute on extra-large mazes. (A) Compute savings: lowering the fire rate of confident cells ( pfire =0.4 if ci,j >0.95, else 0.8) requires more steps to solve the mazes (reach y=0.95 with 1.29× steps), but reduces total cell updates to 0.71× . (B) Localized activity: 10 steps after adding damage mid-rollout (cyan circles), cells automatically activate either on damaged zones or on the yet-unsolved backtracking front (pink), while stable path segments remain largely dormant. (C) Enhanced damage recovery: When injecting damage at t=6000 , the adaptive strategy recovers faster ( 0.94× steps) and cuts cumulative cell operations nearly in half ( 0.52× ). Metrics are averaged over 3 test seeds.
Figure 9: 256×256 NCA solving an out-of-distribution hard sample from Visual Sudoku.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 10: Maze-Hard task ambiguity. While trained on A* targets ( exact , left), there are multiple other valid (middle, blue) and optimal (right, green) solutions.
Figure 11: Sudoku-OOD dataset: the test set (47–64 cells to fill) presents a severe OOD shift relative to training (39–50 cells to fill)
Figure 12: Sudoku Extreme Difficulty. (A) We group test samples in bins of increasing difficulties to plot results across difficulties in Figure 16 . (B) Statistics of difficulties.
Method
Maze-Hard / ARC (Att, S=30×30 )
Sudoku-Extreme (MLP, S=9×9 )
PyTorch (FC)
238.9 B
25.6 B
PyTorch (Prof)
187.5 B
25.7 B
JAX (CPU)
245.2 B
26.2 B
JAX (A100 GPU)
243.5 B
25.8 B
Relative ( ΔFC, CPU )
0.026
0.023
Appendix
Table 2: TRM FLOPs per supervision step.
Hyperparameter
Maze-OOD
Maze-Hard
Sudoku-OOD
Sudoku-Extreme
Grid size ( H×W )
9×9
30×30
9×9
9×9
Vocabulary size ( V )
4
5
10
10
Channels ( C )
16
32
128
256
I/O emb dim ( Cin/Cout )
16 / 2
32 / 3
32 / 32
32 / 32
Sensing module
2D conv
2D conv
Fixed Attn
Fixed Attn
Sensing heads ( Kheads )
4
4
16
16
Appendix
Table 3: Training hyperparameters for the Maze and Sudoku benchmarks.
Hyperparameter
Value
Grid size ( H×W )
32×32
Vocabulary size ( V )
12
Channels ( C )
512
I/O emb dim ( Cin/Cout )
32 / 32
Task emb dim ( Ctask )
128
Hidden state dim ( Chid )
320
Appendix
Table 4: NCA architecture hyperparameters for ARC-AGI-1.
Hyperparameter
Value
Grid size ( H×W )
32×32
Vocabulary size ( V )
12
Channels ( C )
512
I/O emb dim ( Cin/Cout )
32 / 32
Task emb dim ( Ctask )
128
Hidden state dim ( Chid )
320
Appendix
Table 4: NCA architecture hyperparameters for ARC-AGI-1.
Hyperparameter
Pre-training
TTT
Rollout length ( N )
64
64
Loss window ( w )
16
16
Training steps
650,000
10,000 / task
Batch size ( B )
128
8
Optimizer
AdamW
Adam
Gradient clipping
1.0
1.0
Appendix
Table 5: Pre-training and TTT optimization hyperparameters for ARC-AGI-1.
Hyperparameter
Visual Sudoku
Grid size ( H×W )
256×256 (4 px pad)
Channels ( C )
256
Visual channels ( Cin/Cout )
1 / 1
Sensing module
2D Conv (Sobel + Identity)
Sensing heads ( Kheads )
3 (fixed)
Boundary padding
wrap
Appendix
Table 6: Training hyperparameters for Visual Sudoku.
Figure 13: Fixed-attention ablation study. Training loss ( A ) and board accuracy ( B ) show that fixed attention consistently trains faster and achieves higher solve rates than the standard self-attention on Sudoku-OOD and Sudoku-Extreme. Mean-std curves over 3 training seeds are shown.
Figure 14: Self-attention converges to static spatial patterns. Attention weights across the 16 perception heads for fixed attention (left) and self-attention at rollout steps t∈{50,100,150} (right). Each panel shows the 9×9 Sudoku grid, with each cell displaying 9 dots showing attention weight to its 3×3 Moore neighborhood (size indicates weight magnitude; green is positive, pink is negative). Self-attention weights, derived from query-key dot product at rollout step t , have converged to static spatial patterns that do not depend on cell states nor time step. Fixed attention parameterizes this inter-cell coupling directly, leading to more efficient learning.
Model
Update Steps ( D )
Time-to-Solve ( tsolve )
TRM
48
3.3 ± 2.3
NCA
200
42.3 ± 21.1
Appendix
Table 7: The locality bottleneck enforces iterative reasoning. Time-to-Solve (mean ± std) on the mutually solved subset of Maze-Hard.
Figure 15: TRM converges in very few steps.
Figure 16: Extended test-time scaling results on Sudoku benchmarks. Evaluation on (A) Sudoku-OOD and (B) Sudoku-Extreme across, rollout iterations (left sub-panel) and parallel trials (right sub-panel). Colored curves show test accuracy for different difficulties without noise (solid lines) and with noise injection (dashed lines). Metrics are averaged over 3 test seeds.
Figure 17: Extended test-time scaling results on Maze benchmarks. Evaluation on Maze-Hard (A) with A∗ targets and (B) with BFS targets, across rollout iterations (left sub-panel) and parallel trials (right sub-panel). Colored curves show test accuracy for different difficulties without noise (solid lines) and with noise injection (dashed lines). Metrics are averaged over 3 test seeds.
Figure 18: Test-time compute scaling on ARC-AGI public evaluation set. (A) Pass@1 majority-voting accuracy across increasing numbers of offline and online augmentations. (B) Pass@ K scaling under the offline-first-augmentation policy, reaching 60.3% Pass@64.
Figure 19: Test-time scaling via extended rollouts on Visual Sudoku hard puzzles , rendered both with unseen and seen MNIST digits. The accuracy improves from approximately 10% to 23% with additional test time compute with seen, and from approximately 8% to 19% with unseen.
Figure 20: End-to-end wall-clock inference runtime and speedup of test-time pruning on Sudoku-Extreme. (A) Total inference runtime (hours on NVIDIA H100 GPUs across 110,000 test puzzles) for the unpruned baseline ( K=512 ) and pruning divisors m∈{2,4,8} across rollout horizons D∈{64,…,2048} (error bars indicate ±1 SD over 3 seeds). (B) Empirical wall-clock speedup multiplier ( Timebase/Timepruned ) versus rollout horizon D . Horizontal dashed lines indicate the theoretical FLOP-reduction ceilings 2−2−(m−1)m ( 1.33× , 2.13× , and 4.01× for m=2,4,8 ); measured speedups closely track the theoretical limits across all horizons ( <1% sorting and compaction overhead).
Technical University of Darmstadt, Darmstadt, Germany · ImFusion GmbH, Munich, Germany · Inria Center at University Cˆote d’Azur, Sophia Antipolis, France +9