cs.LGSep 30, 2026

Learning Chaos Without Seeing Chaos: Extrapolation of Global Dynamics in Autoregressive Transformers

Authors: Yilun Liu, Yi Zhang, Ganyu Wu, Sikuan Yan, Mengyue Wang, Alois Knoll, Volker Tresp, Yunpu Ma

Organizations: Ludwig Maximilian University of Munich · Munich Center for Machine Learning · Technical University of Munich

Abstract

Autoregressive models are trained to predict a system's behavior one step at a time, and recursive generation allows the learned dynamics to unfold over long horizons. To what extent can such dynamics learned from local observations recover broader organization of an underlying system that was only partially observed during training? Here we study small autoregressive transformers trained from scratch on trajectories sampled from restricted parameter regimes of several non-linear dynamical systems, including logistic and sine maps, the Lorenz system, and the generalized Hopf system, with control parameters and state trajectories represented as sequences of continuous tokens. Under closed-loop evaluation at parameters far outside the training distribution, the models can recover self-similar period-doubling cascades, chaotic dynamics, and attractor structures with remarkable visual and numerical fidelity. For the logistic map, a transformer reproduces successive period doublings up to period 128, yielding a finite-order scaling ratio of 4.6687, matching the Feigenbaum constant to within 5×10−45\times10^{-4}. We further investigate how these structures emerge over the course of training, and reveal with causal interventions how control-parameter information is processed through attention into state prediction and shapes the resulting closed-loop dynamics. These results suggest that a surprisingly narrow window into a system's local behavior may suffice for autoregressive transformers to generalize to its unseen global dynamical organization.

Figures & tables

Appendix figures & tables48 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. A Hybridizable Neural Time Integrator for Stable Autoregressive Forecasting

    Apr 22, 2026Brooks Kinch, Xiaozhe Hu, Yilong Huang +6Autoregressive TransformersAutoregressive Model

  2. On the Diverse Dynamical Behaviors Arising in Deep Linear Transformers

    Jul 20, 2026Sixu Li, Thomas Jacob Maranzatto, Jan Peszek +5Transformer ArchitecturesLinear Attention

  3. Watch your neighbors: Training statistically accurate chaotic systems with local phase space information

    May 14, 2026Joon-Hyuk Ko, Andrus Giraldo, Deok-Sun LeeChaosSurrogate Models