Autoregressive models are trained to predict a system's behavior one step at a time, and recursive generation allows the learned dynamics to unfold over long horizons. To what extent can such dynamics learned from local observations recover broader organization of an underlying system that was only partially observed during training? Here we study small autoregressive transformers trained from scratch on trajectories sampled from restricted parameter regimes of several non-linear dynamical systems, including logistic and sine maps, the Lorenz system, and the generalized Hopf system, with control parameters and state trajectories represented as sequences of continuous tokens. Under closed-loop evaluation at parameters far outside the training distribution, the models can recover self-similar period-doubling cascades, chaotic dynamics, and attractor structures with remarkable visual and numerical fidelity. For the logistic map, a transformer reproduces successive period doublings up to period 128, yielding a finite-order scaling ratio of 4.6687, matching the Feigenbaum constant to within 5×10−4. We further investigate how these structures emerge over the course of training, and reveal with causal interventions how control-parameter information is processed through attention into state prediction and shapes the resulting closed-loop dynamics. These results suggest that a surprisingly narrow window into a system's local behavior may suffice for autoregressive transformers to generalize to its unseen global dynamical organization.
Figures & tables
Figure 1: Autoregressive Transformer trained only on pre-bifurcation trajectories of the logistic map recovers self-similar period-doubling structure beyond the domain it was trained on. The shaded areas mark the regime [2,3] used for training. Closed-loop results are generated by transformer of seed 42 trained after 10,000 steps with one state token, where the boxes mark zoomed-in areas in the successive panels. Quantitative comparisons with the reference are detailed in Section 4.1 and Appendices C.1 and D .
Figure 2: Emergence of period-doubling cascades beyond the observed regime can be numerically verified via the reconstructed scaling patterns of period-doubling cascades. Results for the Logistic map, averaged across seeds with shaded min–max ranges; solid segments are evidence supported by all seeds, dashed segments only partially supported by those that had reached that level. rn is verified as detailed in Section B.3 .
Figure 3: Emergence of period-doubling cascades beyond the observed regime can be visually examined via the similar patterns reconstructed. The shaded areas mark the regime [2,3] used for training. Results for the Logistic map, with seed 42, C=1 , on the same 513-point grid used in Figure 1 .
Figure 4
Appendix figures & tables48 assets
Supplementary material from the paper’s appendix.
Appendix
Map
Levels
Max. burn
Tail
Logistic reference
4
65,536
2,048
Sine reference
4
65,536
2,048
MLP
5
131,072
2,048
Transformer
4
32,768
1,024
Appendix
Table 1: Final adaptive budgets. Each level compares two burn-in budgets; “maximum” is the longer final budget. Neural maps use logistic seed 42 and one current state.
Index
Logistic reference
Sine reference
Logistic MLP
Logistic Transformer
1
3.0000000000
0.7199616830
2.9879977074
2.9692234640
2
3.4494897428
0.8332663537
3.4523234903
3.3927447143
3
3.5440903596
0.8586090599
3.5573408693
3.4848703460
4
3.5644072661
0.8640841737
3.5802157502
3.5048663661
5
3.5687594195
0.8652589607
3.5851355903
3.5091579915
6
3.5696916098
0.8655106638
3.5861901963
3.5100776679
Appendix
Table 2: All numerically verified flip parameters. Learned maps use logistic seed 42. The Transformer receives one state token and a parameter token.
Figure 6: Dense local views resolve the cascade shape. Reference and Transformer panels share the same physical parameter window within each system. All 1,025×3×512 retained samples enter the histograms. The parameter positions are not aligned or rescaled between reference and learned maps.
Figure 7: Local geometry at doubling 3. The verified parent period is 4 and the emerging doubled period is 8 . Each panel uses its own physical root and explicitly labeled local state range. Dense clouds near a root can include slow transients; root verification additionally uses cycle and multiplier checks.
Figure 8: Local geometry at doubling 4. The verified parent period is 8 and the emerging doubled period is 16 . Each panel uses its own physical root and explicitly labeled local state range. Dense clouds near a root can include slow transients; root verification additionally uses cycle and multiplier checks.
Figure 9: Local geometry at doubling 5. The verified parent period is 16 and the emerging doubled period is 32 . Each panel uses its own physical root and explicitly labeled local state range. Dense clouds near a root can include slow transients; root verification additionally uses cycle and multiplier checks.
Figure 10: Local geometry at doubling 6. The verified parent period is 32 and the emerging doubled period is 64 . Each panel uses its own physical root and explicitly labeled local state range. Dense clouds near a root can include slow transients; root verification additionally uses cycle and multiplier checks.
Figure 11: Local geometry at doubling 7. The verified parent period is 64 and the emerging doubled period is 128 . Each panel uses its own physical root and explicitly labeled local state range. Dense clouds near a root can include slow transients; root verification additionally uses cycle and multiplier checks.
Figure 12: Logistic learning curves. Training and probe measurements are recorded every 250 updates; autonomous diagnostics use the 23 fixed stages. One-step probes use 2,048 samples per region. Period matching uses confirmed periodic reference trajectories, with unresolved model outputs counted as mismatches. All curves concern seed 42 and one state token.
Figure 13: Sine learning curves. Training and probe measurements are recorded every 250 updates; autonomous diagnostics use the 23 fixed stages. One-step probes use 2,048 samples per region. Period matching uses confirmed periodic reference trajectories, with unresolved model outputs counted as mismatches. All curves concern seed 42 and one state token.
Figure 14: Coarse structural detection across training. Each point counts consecutive confirmed period doublings on the fixed 513-point grid. Zero denotes no cascade confirmed at this budget. Adaptive high-order verification uses the separate local protocol.
Figure 15: Flip parameters move during training. Points show the first three detected flip parameters and vertical bars show their trajectory brackets. Dashed horizontal lines mark the reference flip parameters, solved independently of the trajectory detector. Unresolved detections appear as gaps. These curves track geometric displacement separately from the number of detected doublings.
Figure 16: Logistic attractor evolution. All 23 closed-loop stages plus the ground-truth reference, arranged as four columns and six rows. Panels use the same compact style as Figure 3 : no colorbars or axis titles, ticks only at the range endpoints and training boundary, and gray shading for the training interval. Orange denotes the Transformer and blue the reference.
Figure 17: Sine attractor evolution. All 23 closed-loop stages plus the ground-truth reference, arranged as four columns and six rows. Panels use the same compact style as Figure 3 : no colorbars or axis titles, ticks only at the range endpoints and training boundary, and gray shading for the training interval. Orange denotes the Transformer and blue the reference.
Figure 18: Matched trajectories across regimes. Logistic ground truth and the one-state Transformer after 10,000 updates share seed 42, identical parameters, and the same initial state. Columns progress from a fixed point through period-2, period-4, and period-8 reference regimes to a chaotic reference. Each panel shows 48 autoregressive updates including transients. Blue marks ground truth and orange the Transformer. The chaotic illustration uses r=3.8 , where both maps have positive finite-orbit Lyapunov estimates; the appendix also shows the final 20k model’s period-3 window at r=3.9 . Appendix C.2.4 repeats these views across both systems and training stages.
Figure 19: Reference and learned cobwebs for the same regimes. Each column stacks the ground-truth cobweb above the Transformer cobweb at the matching parameter after 10,000 updates. Solid curves show the corresponding scalar map and dashed lines mark the identity. Blue marks ground truth and orange the Transformer. Repeated iteration turns the learned transition into stable, periodic, or irregular motion. The chaotic illustration uses r=3.8 , where both maps have positive finite-orbit Lyapunov estimates; the appendix also shows the final 20k model’s period-3 window at r=3.9 .
Figure 20: Logistic trajectories across training. Each row is one checkpoint; columns are the same five reference regimes. Blue denotes ground truth and orange the Transformer, with the legend only in the top-left panel. Regime names describe the reference, not a classification inferred for the checkpoint.
Figure 21: Logistic cobwebs across training. Each row shows the Transformer at the labeled stage; the final row is ground truth (GT). Columns keep the same reference regimes. Blue marks ground truth and orange the Transformer.
Figure 22: Sine trajectories across training. Each row is one checkpoint; columns are the same five reference regimes. Blue denotes ground truth and orange the Transformer, with the legend only in the top-left panel. Regime names describe the reference, not a classification inferred for the checkpoint.
Figure 23: Sine cobwebs across training. Each row shows the Transformer at the labeled stage; the final row is ground truth. Columns keep the same reference regimes. Blue marks ground truth and orange the Transformer.
Training window
Seed
ID W1
OOD W1
Agreement
Coarse flips
[2.0,3.0]
42
0.00179
0.03456
70.2%
3
43
0.00050
0.02772
84.0%
3
44
0.00045
0.02108
81.3%
3
[2.5,3.0]
42
0.00077
0.03172
90.0%
3
43
0.00073
0.03940
88.2%
3
44
0.00389
0.06319
76.9%
0
Appendix
Table 3: Per-seed window ablation. Same protocol for every row. W1 is the uniform-grid OOD average; agreement is over reference-periodic parameter–initial-state pairs; coarse flips count consecutive detected doublings on the 513-point grid.
Training window
OOD W1
Periodic agreement
First verified flip
Verified flips
[2.0,3.0]
27.8±6.7
78.5%
2.9692
7
[2.5,3.0]
44.8±16.4
85.0%
3.0170
6
[2.9,3.0]
36.1±5.0
88.2%
2.9884
6
[2.0,2.5]
96.8±55.5
69.1%
2.9304
6
Reference
—
—
3.0000
7
Appendix
Table 4: Observation-window ablation. All four windows use the locked protocol of Appendix A and are evaluated on the same 513-point grid. OOD W1 is in units of 10−3 , mean ± sample SD over three training seeds. “Verified flips” counts consecutive flips passing the local root and multiplier checks for seed 42 (three refinement rounds for every window).
Training window
Verified flips
First three flip parameters
Largest δn
[2.0,3.0] (published)
7
2.9692235, 3.3927447, 3.4848703
4.6687
[2.5,3.0]
6
3.0169678, 3.4462607, 3.5237060
4.6731
[2.9,3.0]
6
2.9884133, 3.3770377, 3.4491587
4.6745
[2.0,2.5]
6
2.9303844, 3.2847299, 3.3538837
4.6739
Logistic reference
7
3.0000000, 3.4494897, 3.5440904
4.6691
Appendix
Table 5: Verified flips per training window (seed 42). Parameters are the flip parameters verified by orbit closure and multiplier −1 with two-sided continuation; δn follows Equation ( 5 ). The last column shows the largest finite-order ratio reached in the available levels.
Figure 24: Logistic Transformer seeds and contexts. Columns share ground truth with seeds 42–44. Rows mark context length C . Gray shading marks the training interval.
Figure 25: Sine Transformer seeds and contexts. Columns share ground truth with seeds 42–44. Rows mark context length C . Gray shading marks the training interval.
Figure 26: Parameter-resolved fidelity for Transformer C=1 . Lines show each seed separately, using every original uniform-grid Wasserstein-1 measurement. The dashed vertical line marks the training boundary. Vertical ranges adapt to the observed error and are labeled; these curves localize differences summarized by the main OOD table.
Figure 27: Parameter-resolved fidelity for Transformer C=4 . Lines show each seed separately, using every original uniform-grid Wasserstein-1 measurement. The dashed vertical line marks the training boundary. Vertical ranges adapt to the observed error and are labeled; these curves localize differences summarized by the main OOD table.
Figure 28: Parameter-resolved fidelity for Transformer C=255 . Lines show each seed separately, using every original uniform-grid Wasserstein-1 measurement. The dashed vertical line marks the training boundary. Vertical ranges adapt to the observed error and are labeled; these curves localize differences summarized by the main OOD table.
Map
Fitted z
Reference z
Schwarzian negative
Logistic Transformer, C=1
2.000
2.000
97.3% (reference 100%)
Logistic MLP
2.021
2.000
100% (reference 100%)
Sine Transformer, C=1
2.100
2.000
96.6%
Unimodal p=4 Transformer
2.005
3.932
100%
Appendix
Table 6: Class diagnostics at an out-of-distribution parameter. z is the fitted order of the maximum ( z=2 quadratic, z=4 quartic). Derivatives are autograd first derivatives of the float64 map; second and third derivatives are finite differences of that exact first derivative.
Quantity
Reference
Transformer
Verified flips
6
1
First flip
0.4882812
0.4987569
Finite-order ratios ( n=2..4 )
5.580, 6.963, 7.283
—
Fitted critical exponent z
3.932
2.005
Coarse flips in μ∈[0.4,1.7]
3
1
OOD W1 (three seeds)
—
0.0982, 0.3668, 0.0962
Appendix
Table 7: Quartic family. Reference flips and finite-order ratios obtained with the manuscript’s pipeline; the literature value for the quartic class is δ≈7.2847 , clearly separated from the quadratic 4.6692 . The learned map is verified at one flip only, and its fitted critical exponent is quadratic.
Map
supx∣Fθ−F∣
mean ∣Δ∂xF∣
mean ∣Δ∂rF∣
mean ∣∂r2Fθ∣
Logistic Transformer, C=1
0.041
0.166
0.036
0.073
Logistic MLP
—
0.048
0.024
0.041
Sine Transformer, C=1
—
0.079
0.121
0.638
Window [2.9,3]
0.102
—
0.143
0.138
Window [2,2.5]
0.103
—
0.143
0.138
Parameter token permuted
0.255
—
0.164
0.012
Appendix
Table 8: Map-level diagnostics (OOD slice, unclipped float64 map except where noted). supx∣Fθ−F∣ is in physical state units; ∂r2Fθ is spurious parameter curvature, identically zero for the reference.
Model
Logistic
Sine
TF, C=1
27.86±6.76
34.42±6.20
TF, C=4
34.02±16.32
33.29±16.39
TF, C=255
129.77±81.24
82.32±36.12
MLP
22.97±1.53
46.23±6.66
Polynomial NVAR
2.15±0.13
2.24±0.16
Appendix
Table 9: OOD distributional fidelity. Uniform-grid W1 in units of 10−3 , mean ± sample SD over three training seeds. C counts state tokens, excluding the parameter token. The polynomial NVAR model uses a separate, smaller fitting budget.
Figure 29: Logistic MLP comparisons. Columns are ground truth and seeds 42–44. The MLP directly represents the scalar next-state map.
Figure 30: Logistic NVAR comparisons. Columns are ground truth and seeds 42–44. NVAR is a parameter-conditioned polynomial model with its separate training protocol.
Figure 31: Sine MLP comparisons. Columns are ground truth and seeds 42–44. The MLP directly represents the scalar next-state map.
Figure 32: Sine NVAR comparisons. Columns are ground truth and seeds 42–44. NVAR is a parameter-conditioned polynomial model with its separate training protocol.
Figure 33: Parameter-resolved fidelity for MLP. Lines show each seed separately, using every original uniform-grid Wasserstein-1 measurement. The dashed vertical line marks the training boundary. Vertical ranges adapt to the observed error and are labeled; these curves localize differences summarized in Table 10 .
Figure 34: Parameter-resolved fidelity for NVAR. Lines show each seed separately, using every original uniform-grid Wasserstein-1 measurement. The dashed vertical line marks the training boundary. Vertical ranges adapt to the observed error and are labeled; these curves localize differences summarized in Table 10 .
Family
Model / context
Seed
ID W1
OOD W1
Period agreement
Coarse flips
Logistic
TF, C=255
42
0.087319
0.036792
73.6%
1
Logistic
TF, C=1
42
0.001788
0.034595
70.2%
3
Logistic
TF, C=4
42
0.002201
0.030639
66.4%
3
Logistic
MLP
42
0.000503
0.023623
90.0%
3
Logistic
Polynomial NVAR
42
0.000000
0.002137
99.8%
3
Sine
TF, C=255
42
0.004700
0.054363
78.1%
3
Appendix
Table 10: All fixed-grid evaluations. ID/OOD columns are mean W1 over the corresponding parameter subset. Period agreement is over accepted reference pairs on the full grid.
Figure 35: Token-level parameter interventions alter closed-loop dynamics. (a) OOD donor-error reduction after patching final-query attention outputs, for all six Transformers; negative values mean increased donor error. (b,c) OOD W1 after blocking parameter-to-state attention at every step. Lines connect conditions for the same seed; solid/dashed lines indicate one/four state tokens. All 36 conditions share the 129-point grid and trajectory budget.
Family
Seed
L0 donor
L1 donor
L0 random
L1 random
Logistic
42
0.000151
1.000000
0.000022
-40.688593
Logistic
43
0.042383
0.999443
-0.005526
-0.444243
Logistic
44
0.522322
0.129718
-2.708184
-0.225740
Sine
42
-0.003656
0.999995
-0.000583
-0.853852
Sine
43
-0.184315
0.969431
-0.161228
-0.314544
Sine
44
-0.945527
0.591043
-0.155559
-0.433707
Appendix
Table 11: OOD donor-error reduction for complete attention-output patches and equal-norm random controls. Sham reduction is zero throughout. Each entry uses 384 OOD histories from one trained model.
Family
Seed
C
Native W1
Block L0
Block L1
Flip counts
Logistic
42
1
0.034967
0.042183
0.606053
2/2/0
Logistic
42
4
0.030323
0.030537
0.565538
2/0/0
Logistic
43
1
0.027007
0.140915
0.187707
2/0/1
Logistic
43
4
0.018944
0.160885
0.190042
2/0/1
Logistic
44
1
0.023330
0.623765
0.210141
2/1/1
Logistic
44
4
0.052811
0.636572
0.190833
2/0/2
Appendix
Table 12: All closed-loop intervention conditions. Each row contains the three W1 conditions for one model/context. Counts give native/L0-blocked/L1-blocked consecutive flips on the shared 129-point grid.
Seed
ID W1
OOD W1
Periodic agreement
Coarse flips
42
0.0417
0.1914
0.0%
0
43
0.0394
0.1912
0.0%
0
44
0.0421
0.1921
0.0%
0
Aligned pairs (mean of 3)
0.0009
0.0278
78.5%
3
Appendix
Table 13: Trained with permuted parameter labels. Three seeds, identical protocol and evaluation grid as the main comparison. Periodic agreement is over reference-periodic parameter–initial-state pairs; the parameter derivative is measured on the out-of-distribution slice, where the reference mean ∣∂rF∣ is 0.173 .
(ρ,σ)
Reference λ
Model λ
28, 10
0.9046
1.0404
32, 10
0.9431
1.0257
30, 10
0.9919
0.8687
26, 10
0.9001
0.8766
28, 9.375
0.8916
0.8496
28, 11.25
0.8111
0.9837
Appendix
Table 14: Largest Lyapunov exponent of the reference Lorenz flow and of the frozen 2,048-trajectory, width-128 model of Appendix F.1 . Both columns use the same estimator. Positive values indicate exponential divergence on the attractor; negative values correspond to fixed-point or periodic destinations.
N
Width
ID MSE
Fixed/268
P2/358
P4/139
P8/22
Finite/2304
2048
64
1.56e-06
249
1
2
0
2155
2048
128
1.05e-06
266
9
3
0
2225
8192
64
1.74e-06
227
149
6
0
1945
8192
128
5.96e-07
259
2
5
0
1173
Appendix
Table 15: Joint Lorenz evaluation on a 48×48 parameter grid after training on numerically certified fixed-point parameters. Entries count matching labels against the saved reference; denominators exclude the critical ρ=1 line. Finite counts include all cells. All arms use seed 42 and 20,000 updates.
Training support
W1(28)
W1(32)
Repeated switching
I0
2.3415
2.7005
0/8
I1
1.7427
1.9778
0/8
I2
0.1973
0.2857
8/8
Mixed
0.1022
0.2507
8/8
Appendix
Table 16: Training-region controls at a fixed final checkpoint. One model seed per condition; 512 training trajectories each. Mixed uses balanced counts over I0,I1,I2 . Switching counts trajectories across the two OOD parameters and four initial states, not independent model seeds.
Figure 36: Lorenz atlas for mixed-region training. Twenty parameters are shown in a 5×4 layout. Gray denotes the training support 0<ρ<24.0579 , while white denotes unseen parameters. Both panels use the same parameter values, initial states, time windows, and axes. The Transformer result is from the frozen seed-42 model after 10k updates.
Figure 37: Lorenz phase portraits across trained and unseen parameter regimes. Each 5×4 atlas overlays four common initial states at each displayed ρ , retaining t∈[0,5.12] without burn-in. Curves show the x – z projection; arrows indicate motion. Blue: GT; orange: frozen parameter-token Transformer, seed 42, final 10k checkpoint. Gray denotes the training support 0<ρ<24.0579 ; white denotes unseen parameters. All cells and both atlases share the same physical scale, with no trajectory normalization. These are test trajectories at new parameter/initial-state combinations. The equal-duration portraits show transients; the separate 100-unit tests quantify sustained double-wing motion.
For autoregressive modeling of chaotic dynamical systems over long time horizons, the stability of both training and inference is a major challenge in building scientific foundation models. We present a hybrid technique in which an autoregressive transformer is embedded within a novel shooting-based mixed finite element scheme, exposing topological structure that enables provable stability. For forward problems, we prove preservation of discrete energies, while for training we prove uniform bounds on gradients, provably avoiding the exploding gradient problem. Combined with a vision transformer, this yields latent tokens admitting structure-preserving dynamics. We outperform modern foundation models with a 65× reduction in model parameters and long-horizon forecasting of chaotic systems. A "mini-foundation" model of a fusion component shows that 12 simulations suffice to train a real-time surrogate, achieving a 9,000× speedup over particle-in-cell simulation.
Brooks Kinch, Xiaozhe Hu, Yilong Huang +6
Mechanical Engineering and Applied Mechanics, University of Pennsylvania, Philadelphia, PA, USA · Department of Mathematics, Tufts University, Medford, MA, USA · Department of Mathematics and Cybernetics, SINTEF Digital, Oslo, Norway +3
We study the inference-time behavior of deep linear encoder-only transformers through the lens of interacting particle systems. In this perspective, tokens are modeled as particles that interact dynamically through successive linear self-attention layers. We show that in embedding dimension two, for any key, query, and value matrices, the dynamics can be reformulated as a generalized Kuramoto-type model with pure second-harmonic coupling. This formulation is amenable to Watanabe--Strogatz theory which reveals the dynamics are intrinsically low-dimensional regardless of the parameter matrices. For a class of token initializations associated with the Ott--Antonsen (OA) manifold, we show that the parameter matrices induce a diverse variety of long-time behaviors in linear transformers, including clustering, oscillations, and bifurcations. The oscillations and bifurcations are characterized by uncovering a hidden Hamiltonian structure in the dynamics. By establishing a structural stability result, we further show that dynamics initialized near the OA manifold exhibit the same long-time behavior as those initialized exactly on the manifold. Motivated by our theory in dimension two, we conduct numerical experiments for analogous parameter regimes in higher-dimensional transformers. Our numerical experiments suggest that the long-time behaviors characterized in our theoretical results persist in higher dimensions.
Sixu Li, Thomas Jacob Maranzatto, Jan Peszek +5
University of Wisconsin–Madison, Department of Statistics · University of Maryland, Department of Electrical and Computer Engineering · University of Warsaw, Institute of Applied Mathematics and Mechanics +2
Chaotic systems pose fundamental challenges for data-driven dynamics discovery, as small modeling errors lead to exponentially growing trajectory discrepancies. Since exact long-term prediction is unattainable, it is natural to ask what a good surrogate model for chaotic dynamics is. Prior work has largely focused either on reproducing the Jacobian of the underlying dynamics, which governs local expansion and contraction rates, or on training surrogate models that reproduce the ground-truth dynamics' long-term statistical behavior. In this work, we propose a new framework that aims to bridge these two paradigms by training surrogate dynamics models with accurate Jacobians and long-term statistical properties. Our method constructs a local covering of a chaotic attractor in phase space and analyzes the expansion and contraction of these coverings under the dynamics. The surrogate model is trained by minimizing the maximum mean discrepancy between the pushforward distributions of the coverings under the surrogate and ground-truth dynamics. Experiments show that our method significantly improves Jacobian accuracy while remaining competitive with state-of-the-art statistically accurate dynamics learning methods. Our code is fully available at https://anonymous.4open.science/r/neighborwatch.
Joon-Hyuk Ko, Andrus Giraldo, Deok-Sun Lee
Center for AI and Natural Sciences Korea Institute for Advanced Study Seoul, Korea 02455 · School of Computational Sciences Korea Institute for Advanced Study Seoul, Korea 02455