3 Spatial locality: keep neighbours together With one-dimensional rotary positions, a transformer executes the Game of Life in 39.1% of rollouts, and 95% of its errors fall on the 28 border cells. A cell’s next state depends on its 3×3 neighbourhood. A convolution sees this neighbourhood directly, and the standard CNN executes the Game of Life exactly (L1 in Figure 3 ). A raster transformer instead reads each frame as 64 consecutive tokens: a cell’s previous-frame neighbourhood becomes three runs of three tokens about 64 positions earlier, and at the edge of the torus each run wraps to the other side of its row. With learned absolute positions, the standard transformer learns this layout and is exact. With rotary position encoding (RoPE) along the raster sequence, as in most language models, the d64 transformer is not (Table 6 , Appendix E ): at the wrap-around edge a fixed raster offset no longer points to a fixed neighbour, and that is where the errors sit. Two-dimensional positions, with the same parameters, give 100%. Rotating by row and column separately (axial 2D RoPE), with identical parameters and initialization, gives 100% on every seed. Toroidal and deliberately mis-periodic variants also reach 99.4–100%: what matters is two-dimensional position, not the exact period. On unseen rules with one token per cell, 2D rotary positions lift the d128 transformer from 82.6% to 99.3%, higher on every seed, and at d64 neither scheme exceeds 18%; with the previous-frame neighbourhood added to each token, the preregistered 2D-over-1D gains missed their per-seed thresholds (Figure 10 , Appendix E ). Attention can learn the layout; what it lacks is reliability. With learned positions the same stack is exact on the known rule and reaches 85.6% on unseen rules at d128—yet no width is exact within our budget (Appendix E.3 ). The previous-frame binding of Section 4 makes the same decoders near-exact with a few hundred added parameters. The lesson is not that attention cannot represent locality but that it does not reliably learn it. Structure buys reliability, not capability—and reliability is the reason to build locality in rather than stack width. Spatial locality is necessary but not sufficient. The standard CNN, local by construction, executes the Game of Life exactly yet fails most held-out rollouts; what it lacks is temporal locality. 6 What the three properties have in common Each property is one link in Figure 2 : spatial locality brings a cell’s neighbours together (step 1), temporal locality (momentum induction) pairs each observed situation with its outcome (steps 2–3), and temporal stability settles each prediction before later ones use it (the rollout loop). Reachability can be checked before training; learnability only by training, and that is where every standard model within our budget fell short. When the evidence does not contain the answer (Section 4.3 ), the missing link is a prior about the world’s rule family, and the same two tests apply: does a route exist, and does training use it? This suggests a short checklist for a new world. (i) Write down which cells and frames must meet, and which parts of the future depend on which. (ii) Check that some part of the model can see both members of each pair. (iii) Supply the links that training is unlikely to find: a position scheme, a pairing, a commitment order. It is not a new method, and is close in spirit to algorithmic alignment ( Xu et al., 2020 ) and to Canon layers ( Allen-Zhu, 2025 ) . Settling intermediate states resembles the role of chain of thought in transformers ( Feng et al., 2023 ; Merrill and Sabharwal, 2023 ) .
7 Future Work Our evidence comes chiefly from a deterministic family of binary automata on an 8×8 grid; the physical and video directions below remain tests of transfer. Small exploratory experiments on continuous dynamics (Appendix G.2 , Figures 16 – 17 ) already suggest that the answer depends on the dynamics: with shared weights, causal freezing helps Burgers but not linear advection–diffusion, while the locality and pairing results are mixed. Exact physical generation. We will test the three properties in richer systems—interacting particles and continuous dynamics—where a known simulator makes fidelity checkable over a specified horizon: exact agreement for discrete states, declared tolerances for continuous ones. This separates executing the dynamics from the visual plausibility of a rendered rollout; the experiments of Appendix G.2 are a first step. From latent dynamics to rendered video. Instead of showing CA states directly to the model, a rendered automaton keeps the simple rules while adding textures, lighting, and camera motion. A video model could encode the observed frames, infer transitions with the two-frame pairing in a learned latent space, and decode the predicted states into frames; we will ask whether this improves state-level consistency without sacrificing visual quality, relative to an otherwise matched predictor (Appendix ). The question is whether the same information flow survives perception and rendering. Ethics Statement This work studies small synthetic cellular-automaton systems generated by a documented simulator. It involves no human subjects, personal data, or deployed systems, and we foresee no direct ethical risk from the models or the data.
Reproducibility Statement Every number in this paper is regenerated from per-run logs: each registered run records its command, seeds, hardware, a SHA-256 fingerprint of the exact training code, and the checksums of the pinned evaluation corpora used, and every reported number is transcribed from these logs without manual editing. One supplementary probe—the attention-depth ladder at L4—was registered from a retrospectively preserved source snapshot rather than a prelaunch code fingerprint; its predictions and checkpoints passed the same audits. Training data are generated deterministically from (seed, index) by a version-frozen simulator, so no training corpus needs to be shipped; evaluation corpora are pinned files with SHA-256 sidecars that regenerate identically on any machine. Protocol, architecture, and recipe details are in Appendices B – D . The full codebase—simulator, single-entry training script, evaluation and analysis tooling, and the pinned evaluation corpora—has been staged locally for release. Public availability and a stable archive link will be confirmed in the arXiv version after the release checks are complete.
Statement on the Use of AI Assistance In this work, we used generative AI tools (large-language-model assistants) for: providing feedback on research methodology and experiment design, with several of the paper’s techniques—including the constructor-swap and sampler-swap controls—originating in human–AI collaborative design sessions; helping develop the conceptual framework (the information-flow requirements and the audit protocol); implementing methods (experiment scripts, model code, and analysis tooling); interpreting results (drafting diagnostic summaries from run logs); assisting with translation between the authors’ working language and English; and drafting and editing parts of the paper for readability. Additionally, we used them for identifying and verifying literature, formatting references, creating figure-generation scripts, and suggesting experimental parameters. We did not use generative AI tools to generate synthetic datasets—all data are produced by a deterministic, version-frozen simulator from (seed, index)—and this work contains no mathematical proofs, so proof-related assistance is not applicable. We have reviewed all AI-assisted work: every experimental number was checked by the authors against the registered run logs; citation keys were checked against the bibliography, and bibliographic metadata was checked against available arXiv, CrossRef, and publisher records where applicable; AI-assisted code is covered by a test suite, including bit-identity tests for every compatibility flag; and the authors directed and adjudicated all design proposals. We take full responsibility for the final content of this work, including text, claims, and artifacts produced with the aid of generative AI.
附录 A Related work Cellular automata and rule learning. Cellular automata make global dynamics from local rules ( Wolfram, 1983 ; Gardner, 1970 ) . Convolutional networks can implement cellular-automaton updates by construction ( Gilpin, 2019 ) , although learning even the Game of Life from examples can be difficult ( Springer and Kenyon, 2021 ) . Neural cellular automata learn a parameterized local update ( Mordvintsev et al., 2020 ) ; related sequence-model studies include shortcuts on finite automata ( Liu et al., 2023 ) , elementary one-dimensional rules with rule-disjoint free runs ( Burtsev, 2024 ) , recursive prediction under the fixed Game of Life rule ( Berkovich and Buehler, 2025 ) , and two-dimensional rule-supplied forecasting alongside inverse rule inference ( Berkovich et al., 2026 ) . Our setting isolates prefix-conditioned continuation on held-out rules: queried situations are covered by observed transitions, and success requires every cell in every predicted frame to be correct (§ 2 ). We diagnose known devices under this exact metric without claiming priority for learning across rules. Spatial locality and trainability. Convolutions expose neighbouring cells directly, whereas a raster sequence requires the model to recover two-dimensional adjacency, including boundary wrap-around. Rotary position embeddings originated in sequence models ( Su et al., 2024 ) ; two-dimensional variants are established for image tokens ( Heo et al., 2024 ) . The choice of position scheme exemplifies how architecture and target computation should align ( Battaglia et al., 2018 ; Xu et al., 2020 ; Veličković et al., 2020 ) ; input parameterization can change which algorithm a transformer learns ( McLeish et al., 2024 ) . In our controlled spatial comparisons, the same architecture can represent the rule but its reliability changes when positions or local inputs expose the grid structure (§ 3 ). Thus a locality benefit in this testbed does not imply a general inability of attention to compute local updates. Induction and context-dependent retrieval. Induction-head accounts describe a previous-token operation followed by matching and copying ( Elhage et al., 2021 ; Olsson et al., 2022 ) ; controlled work studies how such behaviour forms and depends on data ( Singh et al., 2024 ; Edelman et al., 2024 ; Chan et al., 2022 ; Reddy et al., 2024 ) . Theory characterizes representational and training conditions for induction on restricted tasks ( Sanford et al., 2024 ; Ekbote et al., 2025 ; Nichani et al., 2024 ) . Our coverage-filtered cellular-automaton task calls for matching a local situation in the observed frames and retrieving its successor. The attention patterns we report are consistent with this two-step account, but observational attention and approximate matches do not identify a unique causal circuit (§ 4.2 , Appendix F ). Putting situations beside outcomes. Temporal shifts in video and sequence models ( Lin et al., 2019 ; Peng et al., 2023 ) , short convolutions for recall ( Fu et al., 2023 ; Arora et al., 2023 ) , and shifted keys and values for induction ( Xu et al., 2024 ) provide precedents for placing information from adjacent times in one representation. Our convolution over adjacent frames, index-aligned key–value shift, and previous-frame neighbourhood input adapt such existing devices to observed situation–successor pairs. They differ in where the pair is stored, but each makes the transition available before retrieval. Matched temporal-window controls test the pairing requirement without assigning special status to one direction (§ 4 ). The observed gain concerns learnability under this task and training budget, not a new pairing operation. Diffusion and commitment order. Denoising diffusion underlies video generation ( Ho et al., 2020 ; Ho et al., 2022 ) ; per-frame noise levels and ordered or streaming generation already appear in Diffusion Forcing, rolling diffusion, FIFO-Diffusion, and causal video generation ( Chen et al., 2024a ; Ruhe et al., 2024 ; Kim et al., 2024 ; Yin et al., 2025 ) . Decoding order also matters in masked discrete generation ( Kim et al., 2025 ) . In our controlled comparison, one denoiser’s weights are held fixed while joint denoising and causal freezing use different schedules with 49 and 48 network calls, respectively (§ 5 ). The result tests whether settling each predicted frame before later frames use it helps exact continuation in these worlds; it neither introduces frame ordering nor establishes a universal preference for it. World models and exact evaluation. Othello-GPT probes a learned board-state representation ( Li et al., 2023 ) . Other work on learned world models asks whether plausible trajectories reflect the underlying dynamics ( Ha and Schmidhuber, 2018 ; Vafa et al., 2024 ; Vafa et al., 2025 ) ; video and physical-reasoning evaluations likewise probe the gap between appearance and physical prediction ( Bear et al., 2021 ; Bansal et al., 2024 ; Motamed et al., 2026 ) . Measurement choices can change apparent capability ( Schaeffer et al., 2023 ) , while video quality measures such as FVD ( Unterthiner et al., 2018 ) answer a different question from whether every predicted cell in every frame follows the rule. We therefore report strict sequence accuracy on rule-disjoint, coverage-filtered continuations and check that the observed evidence determines their targets (§ 2 ). The local pairing and frame-order results suggest hypotheses for broader generators, rather than verified transfer beyond this controlled family (§ 7 ).
附录 B Protocol and reporting This appendix fixes the rule split, the test corpora, and the reporting rules. Rule indexing and split. A rule is an 18-bit integer whose bit 9s+n is the next state of situation (s,n) . The split is a fixed-seed permutation of the 262,144 indices, because table entries are index bits and a parity or modulo split would tie some entries to one half. Coverage and the self-consistency filter. Without the filter, 33% of L3 trajectories query a situation that the observed frames never show. If u counts such situations, a uniform prior on the unseen rule entries gives the realized continuation posterior mass 2−u , with mean 0.8099. This is not an upper bound on SeqAcc: averaged over the 2048 observed prefixes, the most probable continuation carries mass 0.8159, or 0.8732 under the prior restricted to the held-out half—conditional references, not measured accuracies. The analytic 3D-conv CNN (Section 4.1 ) scores 0.8032 on the unfiltered corpus and 1.0000 on its covered subset. Rejection sampling keeps 67% of draws at L2 and L3. Masking the loss on unobserved situations instead of filtering the stream is neutral ( Δ=0.0001 , paired). Level 4 corpora. Level 4 worlds come from held-out rules with the prior of Section 2 built in: the two situations with opposite centre cells and n∈{3,4} black neighbours have opposite outcomes, r(1,n)=1−r(0,n) . L4B, the Level 4 test of the main text, holds 768 trajectories; in each, a queried situation never appears among the observed transitions while its centre-flipped partner does. They were selected from 3.86M draws, so the event’s natural rate is about 2×10−4 . Level 4 training therefore replaces 10% of each batch (the 10% dose ) with such trajectories, drawn from a pool of 65,536 training-half trajectories generated with a seed of its own; natural rate denotes training without the dose. L4A, a companion test of 2048 trajectories, instead fixes four rule rows across the family and queries one of them without showing it, so the model must recall the family’s fixed value. Pixel-level versus sequence-level accuracy. At 0.999 per pixel, independent errors would still spoil about two in five eight-frame rollouts on the 8×8 torus ( 1−0.999512≈0.4 ), and the metric cannot separate a model that has the rule from one that is approximately right. Reporting conventions. Numbers are final-step values on exponential-moving-average (EMA) weights, without checkpoint selection, over three preregistered seeds (42/43/44). Main-text numbers are three-seed means in percent, truncated, not rounded, to one decimal: 0.9995 shows as 99.9, and only an exact 1.0 shows as 100. Appendix tables give the per-seed values behind them. Figure cells with fewer registered seeds are marked where they appear. Two choices used validation data, never test data: the checkpoint of one continuous-dynamics experiment (Appendix G.2 ) and the number of attention reads in Figure 14 . Data are generated deterministically from (seed, index); evaluation corpora carry SHA-256 sidecars and hold 512 trajectories at L1, 2048 at L2 and L3, and 768 at L4 (L4B). Models train on the filtered stream with bf16 autocast and are evaluated in fp32 unless a table says otherwise; runs on the unfiltered stream appear only where labelled. B.1 Preregistered predictions and corrections Two predictions registered for the standard transformer failed on the filtered stream (Table 7 ). The first was that it would stay near zero on held-out rules ( ≈ 0–5%); d128 reaches 0.9263/0.7646/0.8784 on L3. The second was that accuracy would rise monotonically with width; d128 is the best rung, while d192 and d256 each leave one seed at the marginal predictor. Its other half, that no rung would be exact, held.