cs.CLAug 3, 2026

TextNCA: Neural Cellular Automata for Language Modeling via Hierarchical Local Attention

Authors: Avni MittalAvinash AnandAshutosh KumarDikshant KukrejaKritarth PrasadSushane DullooErik CambriaTimothy Liu+2 more

Abstract

Can a strictly local, iterated, weight-shared computation primitive support language modelling, and which of those three properties actually drives the model's behaviour? We define \textsc{TextNCA}, a 1D causal windowed-attention realisation of the Neural Cellular Automaton primitive, and study a hierarchical variant that cascades three stages with windows w{8,32,128}w \in \{8, 32, 128\} and TsT_s shared-weight iterations per stage, all on WikiText-103 at roughly 30M parameters and 60k training steps. The model does not match a parameter-matched Transformer at this scale (Hier-TextNCA 60.360.3 vs.\ Transformer-6L 52.852.8 and Transformer-12L 44.744.7 PPL), so we treat it as an analytical probe rather than a proposed alternative. The behaviour we observe is largely explained by the staged narrow-to-wide schedule: a non-iterating sliding-window Transformer that reuses the same schedule comes within +4.1+4.1 PPL of the iterated model, while reversing, flattening, or breaking the monotonic ordering of the schedule costs between +16.7+16.7 and +70.8+70.8 PPL. Iteration adds a smaller bounded benefit on top of the schedule, with a clear optimum at Ts=4T_s{=}4 and a U-shaped degradation beyond it. The GRU gate and learned per-step embeddings are required for that benefit to appear, and training with random TsT_s yields an inference-time iteration-count knob at the cost of substantially higher absolute PPL. We position the work as a controlled reading of which parts of NCA-style computation carry the weight in language modelling.

Explore similar work

CardsList
  1. Express Language Modeling

    Jun 9, 2026Albert Gong, Annabelle Michael Carrell, Raaz Dwivedi +1Graph Language ModelsCausal Attention