Understanding Latent-Dimension Scaling in Dynamical-System Learning through Spectral Reliability
Authors: Itsushi Sakata, Yuta Miyauchi, Yoshinobu Kawahara
Organizations: RIKEN Center for Advanced Intelligence Project, Tokyo, Japan. · Graduate School of Information Science and Technology, The University of Osaka, Suita, Osaka, Japan.
In deep learning, approximation theory motivates increasing representation size. We ask whether this benefit extends to dynamics learning through autoregressive prediction. We analyze the learned time evolution through the eigenstructure of Koopman operators, using relative residuals to detect spurious eigenpairs arising even as one-step error falls. For bounded Koopman operators, we show that minimal residuals over learned dictionary spaces converge pointwise to their full-space counterparts as these spaces approximate the observable space in L2. Our hypothesis is that Koopman spectral reliability helps explain how consistently rollout error decreases with increasing dimension. We compare two models of a shared Koopman autoencoder trained alternately for reconstruction and latent evolution, using latent-prediction loss (one-step prediction errors in latent coordinates) or spectral-residual loss (relative residuals of candidate eigenpairs). Across six chaotic systems, both models reduced median windowed rollout error from smallest to largest dimension. The spectral-residual model achieved lower medians than the latent-prediction model for all systems and dimensions, and its median fell by a larger factor in every system. Its median decreased monotonically with dimension in four systems, against one for latent prediction. Against four baseline families, its mean valid prediction times were nearly always longer. At the largest dimension under two-stage training, we compared eigenvalue positions with each learned dictionary's residual contours. Spectral-residual eigenvalues concentrated in low-residual regions, whereas latent-prediction eigenvalues also appeared in high-residual regions, consistent with the hypothesis.
Figures & tables
System
d
N0
Δt
Steps
Burn-in
Steps per LT
Rössler
3
30
0.01
442,000
2,000
1,365.75
Lorenz-63
3
30
0.01
202,000
2,000
110.75
Duffing
2
30
0.075
180,000
14,000
133
Mackey–Glass
1
30
1.0
112,000
12,000
175.2
Lorenz-96
40
40
0.01
230,500
30,000
60.375
Kuramoto–Sivashinsky
64
64
0.25
88,000
8,000
70.5
Table 1: Measured state dimension d , input dimension N0 after delay coordinates, and simulation settings for the six systems.
Figure 1: Windowed autoregressive rollout VRMSE versus latent dimension for the two Stage 2 objectives in two-stage training (lower is better). Prediction and Residual denote the latent-prediction and spectral-residual models, and +L2 adds the operator-norm penalty wop∥A∥2 . Panels a–f show Rössler, Lorenz-63, Duffing, Mackey–Glass, Lorenz-96, and Kuramoto–Sivashinsky. The evaluation window is (0.5,0.7] LT, and the latent dimensions are 1, 2, 4, 8, and 16 times the input dimension N0 of each system. For each seed, we average VRMSE over the test starting points and prediction steps within the evaluation window. Lines show the median of these averages across five seeds, and bands show the interquartile range. Each panel uses a linear vertical scale adjusted to its interquartile range
System
N
Two-stage
Joint
Baselines
Prediction
Prediction +L2
Residual
Residual +L2
Prediction
Prediction +L2
Residual
Residual +L2
ESN
Kernel DMD
Neural ODE
Consistent KAE
Rössler
30
0.16 ± 0.01
0.19 ± 0.01
0.26 ± 0.01
0.24 ± 0.02
0.28 ± 0.01
0.26 ± 0.01
0.24 ± 0.00
0.24 ± 0.01
0.01 ± 0.00
0.05 ± 0.02
0.12 ± 0.02
0.01 ± 0.01
60
0.16 ± 0.00
0.25 ± 0.05
0.27 ± 0.00
0.33 ± 0.03
0.16 ± 0.00
0.16 ± 0.00
0.28 ± 0.01
0.30 ± 0.02
0.01 ± 0.01
0.04 ± 0.01
0.09 ± 0.01
0.05 ± 0.03
120
0.16 ± 0.01
0.24 ± 0.03
0.33 ± 0.01
0.37 ± 0.01
0.19 ± 0.02
0.20 ± 0.01
0.31 ± 0.01
0.31 ± 0.01
0.01 ± 0.01
0.05 ± 0.01
0.09 ± 0.01
0.03 ± 0.00
240
0.27 ± 0.04
0.34 ± 0.01
0.50 ± 0.02
0.53 ± 0.02
0.28 ± 0.01
0.25 ± 0.01
0.39 ± 0.01
0.39 ± 0.01
0.01 ± 0.00
0.27 ± 0.09
0.09 ± 0.01
0.03 ± 0.00
480
0.42 ± 0.05
0.33 ± 0.04
0.68 ± 0.03
0.78 ± 0.06
0.27 ± 0.05
0.24 ± 0.01
0.42 ± 0.01
0.42 ± 0.01
0.01 ± 0.00
0.72 ± 0.11
0.08 ± 0.02
0.03 ± 0.00
Table 2: VPT for 12 conditions in LT (longer is better). Rows correspond to systems and representation sizes N ; columns correspond to the four two-stage and four joint model variants and the four baselines. Prediction and Residual denote the latent-prediction and spectral-residual models, and +L2 adds the operator-norm penalty (Fig. 1 ). Section 3.4 defines what N counts for each model family. For each seed, VPT at a threshold of 0.5 is averaged over 3,000 test starting points; the table reports the mean ± sample standard deviation of these seed-level means across five seeds. Bold indicates the cell with the largest unrounded mean in each row, and a displayed value of 0.00 indicates an unrounded mean below 0.005 LT.
Figure 2: Autoregressive rollout VRMSE versus representation size for the baseline families and the spectral-residual model (lower is better). “Residual (two stage)” denotes the spectral-residual model under two-stage training. Panels a–f show Rössler, Lorenz-63, Duffing, Mackey–Glass, Lorenz-96, and Kuramoto–Sivashinsky. The evaluation window is the same (0.5,0.7] LT window as in Fig. 1 . Markers show medians over five random seeds, and shaded bands show interquartile ranges. The vertical axis is logarithmic
Figure 3: Minimal residuals τM,Nnum and eigenvalues for eight model variants in Rössler at latent dimension 16N0 . Colors show, on a logarithmic scale, contours of the minimal residual computed over the retained numerical subspace, with the same color range across all panels. Low-residual regions appear near λ=1 in all eight panels. White curves mark the unit circle, and red points mark the estimated Koopman eigenvalues within the plotted range. The weighted latent snapshot matrix used for the contours is defined in Sect. 3.3 , and the matrix whose eigenvalues are plotted is defined in Appendix 8 . The top row shows two-stage training, and the bottom row shows joint training. From left to right, the columns show Prediction (the latent-prediction model), Prediction+L2, Residual (the spectral-residual model), and Residual+L2; +L2 adds the operator-norm penalty wop∥A∥2
Figure 4: Predictions for Rössler at latent dimension 16N0 (top: two-stage training; bottom: joint training). From left to right, the columns show the reference trajectory, Prediction (the latent-prediction model), Prediction+L2, Residual (the spectral-residual model), and Residual+L2; +L2 adds the operator-norm penalty wop∥A∥2 . Each row corresponds to one prediction horizon (0.2, 0.6, or 1.0 LT), and predictions are made that far ahead from different observation vectors in the reference trajectory. The plots show the three components of the measured state extracted from predictions whose target times fall within the first 5 LT of the test trajectory. Axis limits follow the reference trajectory and are shared across all panels
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter
Value
Hidden layers per network
2
Hidden width
N
Activation
ReLU
Optimizer
AdamW
Learning rate
10−4
Batch size
2,048
Appendix
Table 3: Training parameters for the Koopman autoencoders.
Fig. S1: Minimal residuals τM,Nnum and eigenvalues for eight model variants in Lorenz-63 at latent dimension 16N0 . Colors show, on a logarithmic scale, contours of the minimal residual computed over the retained numerical subspace, with the same color range across all panels. White curves mark the unit circle, and red points mark the estimated Koopman eigenvalues within the plotted range. The top row shows two-stage training, and the bottom row shows joint training. From left to right, the columns show Prediction (the latent-prediction model), Prediction+L2, Residual (the spectral-residual model), and Residual+L2; +L2 adds the operator-norm penalty
Fig. S2: Predictions for Lorenz-63 at latent dimension 16N0 (top: two-stage training; bottom: joint training). The predictions are plotted in (x,y,z) coordinates. From left to right, the columns show the reference trajectory, Prediction (the latent-prediction model), Prediction+L2, Residual (the spectral-residual model), and Residual+L2; +L2 adds the operator-norm penalty. Each row corresponds to one prediction horizon (0.2, 0.6, or 1.0 LT), and predictions are made that far ahead from different observation vectors in the reference trajectory. The prediction target times fall within the first 5 LT of the test trajectory. Axis limits follow the reference trajectory and are shared across all panels
Fig. S3: Minimal residuals τM,Nnum and eigenvalues for eight model variants in Duffing at latent dimension 16N0 . Colors show, on a logarithmic scale, contours of the minimal residual computed over the retained numerical subspace, with the same color range across all panels. White curves mark the unit circle, and red points mark the estimated Koopman eigenvalues within the plotted range. The top row shows two-stage training, and the bottom row shows joint training. From left to right, the columns show Prediction (the latent-prediction model), Prediction+L2, Residual (the spectral-residual model), and Residual+L2; +L2 adds the operator-norm penalty
Fig. S4: Predictions for Duffing at latent dimension 16N0 (top: two-stage training; bottom: joint training). The predictions are plotted in the (x,x˙) plane. From left to right, the columns show the reference trajectory, Prediction (the latent-prediction model), Prediction+L2, Residual (the spectral-residual model), and Residual+L2; +L2 adds the operator-norm penalty. Each row corresponds to one prediction horizon (0.2, 0.6, or 1.0 LT), and predictions are made that far ahead from different observation vectors in the reference trajectory. The prediction target times fall within the first 5 LT of the test trajectory. Axis limits follow the reference trajectory and are shared across all panels
Fig. S5: Minimal residuals τM,Nnum and eigenvalues for eight model variants in Mackey–Glass at latent dimension 16N0 . Colors show, on a logarithmic scale, contours of the minimal residual computed over the retained numerical subspace, with the same color range across all panels. White curves mark the unit circle, and red points mark the estimated Koopman eigenvalues within the plotted range. The top row shows two-stage training, and the bottom row shows joint training. From left to right, the columns show Prediction (the latent-prediction model), Prediction+L2, Residual (the spectral-residual model), and Residual+L2; +L2 adds the operator-norm penalty
Fig. S6: Predictions for Mackey–Glass at latent dimension 16N0 (top: two-stage training; bottom: joint training). The predictions are plotted in the delay plane (x(t),x(t−τ)) , with τ=17 in the time units of the governing equation. From left to right, the columns show the reference trajectory, Prediction (the latent-prediction model), Prediction+L2, Residual (the spectral-residual model), and Residual+L2; +L2 adds the operator-norm penalty. Each row corresponds to one prediction horizon (0.2, 0.6, or 1.0 LT), and predictions are made that far ahead from different observation vectors in the reference trajectory. The prediction target times fall within the first 5 LT of the test trajectory. Axis limits follow the reference trajectory and are shared across all panels
Fig. S7: Minimal residuals τM,Nnum and eigenvalues for eight model variants in Lorenz-96 at latent dimension 16N0 . Colors show, on a logarithmic scale, contours of the minimal residual computed over the retained numerical subspace, with the same color range across all panels. White curves mark the unit circle, and red points mark the estimated Koopman eigenvalues within the plotted range. The top row shows two-stage training, and the bottom row shows joint training. From left to right, the columns show Prediction (the latent-prediction model), Prediction+L2, Residual (the spectral-residual model), and Residual+L2; +L2 adds the operator-norm penalty
Fig. S8: Predictions for Lorenz-96 at latent dimension 16N0 (top: two-stage training; bottom: joint training). From left to right, the columns show the reference field, Prediction (the latent-prediction model), Prediction+L2, Residual (the spectral-residual model), and Residual+L2; +L2 adds the operator-norm penalty. Each row corresponds to one prediction horizon (0.2, 0.6, or 1.0 LT), and predictions are made that far ahead from different observation vectors in the reference trajectory. Horizontal axis: prediction target time in LT; vertical axis: component index 0–39; color: component value, with the same color range across all panels. Component indices 0–39 correspond, in order, to x1,…,x40 in the main text
Fig. S9: Minimal residuals τM,Nnum and eigenvalues for eight model variants in Kuramoto–Sivashinsky at latent dimension 16N0 . Colors show, on a logarithmic scale, contours of the minimal residual computed over the retained numerical subspace, with the same color range across all panels. White curves mark the unit circle, and red points mark the estimated Koopman eigenvalues within the plotted range. The top row shows two-stage training, and the bottom row shows joint training. From left to right, the columns show Prediction (the latent-prediction model), Prediction+L2, Residual (the spectral-residual model), and Residual+L2; +L2 adds the operator-norm penalty
Fig. S10: Predictions for Kuramoto–Sivashinsky at latent dimension 16N0 (top: two-stage training; bottom: joint training). From left to right, the columns show the reference field, Prediction (the latent-prediction model), Prediction+L2, Residual (the spectral-residual model), and Residual+L2; +L2 adds the operator-norm penalty. Each row corresponds to one prediction horizon (0.2, 0.6, or 1.0 LT), and predictions are made that far ahead from different observation vectors in the reference trajectory. Horizontal axis: prediction target time in LT; vertical axis: spatial sample index 0–63 on a domain of length 22; color: field value, with the same color range across all panels
Koopman theory promises linear structure in nonlinear dynamics, but numerical Koopman spectra are easy to compute and hard to trust. A finite EDMD matrix always has eigenvalues; the problem is that many of them may have nothing to do with the infinite-dimensional operator. In this paper we make spectral reliability the objective of dictionary learning. We train neural-network dictionaries not merely to predict the next snapshot, but to minimize Residual Dynamic Mode Decomposition residuals: operator-level a posteriori errors that test whether computed eigenvalues and modes are genuine Koopman spectral objects. To keep the learned observables from collapsing into an unstable coordinate system, the loss also penalizes the condition number of the lifted data matrix. Thus the method couples two requirements that should not be separated: small Koopman residuals and a well-conditioned representation. The result is a learned dictionary that is expressive, numerically stable, and spectrally disciplined. Across conservative and dissipative benchmark systems, the method sharply reduces spectral pollution, improves residual pseudospectral inclusion, and lowers forecast error relative to standard fixed dictionaries. On sea-surface temperature data, it gives cleaner Koopman diagnostics and substantially better one-step forecasts from noisy observations with no governing equations. The message is simple: neural Koopman learning should be judged not by prediction alone, but by whether its spectral claims can be certified. Residuals provide the certificate; conditioning makes it computable.
George Coote, Matthew J. Colbrook
Department of Applied Mathematics and Theoretical Physics, University of Cambridge, Wilberforce Road, CB3 0WA, United Kingdom
Koopman theory turns nonlinear dynamics into a linear spectral problem. In computation, however, everything depends on a hard finite-dimensional choice: the observables must be expressive, nearly invariant under the dynamics, and, ideally, compatible with composition. Deep Koopman methods learn flexible coordinates, whereas structure-preserving methods enforce operator identities on fixed dictionaries. We combine these ideas by introducing Deep Embedded Multiplicative Dynamic Mode Decomposition (DeepMDMD), a method that learns a latent space and a partition of it, while enforcing the Koopman product rule as an exact algebraic constraint. Training alternates between an exact multiplicative operator update and a differentiable latent-clustering step that promotes Koopman closure. The result is a finite transition map on learned latent cells. Its nonzero spectrum lies on the unit circle, its dictionary is shaped by the dynamics rather than by ambient geometry, and forecasts are made in latent coordinates before being decoded to physical space. Across Hamiltonian, chaotic, and fluid examples, DeepMDMD learns dictionaries that are far more compact and dynamically coherent than those produced by geometric MDMD partitions. It reduces spectral pollution, reveals richer continuous-spectrum structure, and gives stable forecasts under severe noise. In high-dimensional flows, including a 158,624-dimensional cylinder wake and a noisy Re=20,000 lid-driven cavity, it preserves coherent structures and long-time spectral statistics where state-space MDMD fails. These results suggest a practical rule for Koopman learning: learn the coordinates, constrain the algebra.
Kelan Gray, Finlay Brown, Nicolas Boullé +1
Department of Mathematics, Imperial College London, London, SW7 2AZ, UK. · Department of Applied Mathematics and Theoretical Physics, University of Cambridge, Cambridge, CB3 0WA, UK.
Learning dynamical systems through operator-theoretic representations provides a powerful framework for analyzing complex dynamics, as spectral quantities such as eigenvalues and invariant structures encode characteristic time scales and long-term behavior. However, dynamical operators are typically estimated independently for each system, preventing the discovery of shared structure across related dynamics. To address this limitation, we posit that related dynamical systems lie near a low-dimensional manifold in spectral operator space. Based on this hypothesis, we introduce DOODL (Dynamical OperatOr Dictionary Learning), a framework that learns a dictionary of characteristic spectral dynamics whose combinations approximate this manifold and yield compact, interpretable embeddings of individual systems. Beyond representation learning, DOODL enables fast and interpretable operator estimation from short and partially observed trajectories by constraining the estimation to the learned operator manifold. Experiments on metastable Langevin dynamics and turbulent plasma simulations demonstrate that DOODL scales to highly complex multiscale regimes while capturing characteristic spectral structure governing the dynamics rather than merely fitting trajectories, achieving errors one to two orders of magnitude lower than independent operator estimation methods in challenging low-data regimes.
Thibaut Germain, Sami Chemlal, Rémi Flamary +2
CMAP, Ecole Polytechnique · Istituto Italiano di Tecnologia & University of Novi Sad