Bilevel optimization for data-driven learning of Koopman embeddings using kernel-based autoencoders
Authors: Joel-Pascal Ntwali N'konzi, Feliks Nüske, Stefan Klus
Organizations: Maxwell Institute for Mathematical Sciences, The University of Edinburgh & Heriot–Watt University, Edinburgh, United Kingdom · Max Planck Institute for Dynamics of Complex Technical Systems, Magdeburg, Germany · LAAS-CNRS, Université de Toulouse, Toulouse, France · School of Mathematical & Computer Sciences, Heriot–Watt University, Edinburgh, United Kingdom
Koopman operator theory provides a linear framework for analyzing nonlinear dynamical systems and has become a major tool for data-driven modeling. A central challenge, however, is that finite-dimensional approximations computed by methods such as extended dynamic mode decomposition (EDMD) require the dictionary to be specified a priori. Recent machine-learning approaches address this limitation by learning the dictionary from data, predominantly using artificial neural network (ANN) autoencoder architectures. Although kernel methods offer an alternative with greater interpretability and tractability for theoretical analysis, they have received little attention in this setting. We introduce extended dynamic mode decomposition with kernel-based dictionary learning (EDMD-kDL), a kernel-based method for learning finite-dimensional Koopman embeddings directly from data. The method combines ideas from collocation methods and bilevel optimization to simultaneously learn a kernel dictionary and the corresponding Koopman approximation. We evaluate EDMD-kDL against state-of-the-art ANN-based approaches on a range of numerical experiments, including global sea-surface-temperature forecasting and learning directly from video data. Across all tested settings, EDMD-kDL achieves performance comparable to or better than the ANN-based methods. Moreover, in contrast to standard kernel methods, the proposed approach is scalable to large datasets by design since the size of the required kernel matrices depends on the number of collocation points rather than the size of the training dataset.
Figures & tables
Figure 1 : Nonlinear pendulum in the clean data setting. Phase portraits and relative prediction errors over 1000 time steps of testing using learned Koopman models. The top and bottom rows correspond to the almost-linear and truly nonlinear regimes, respectively.
Figure 2 : Pendulum with measurement noise in the almost-linear regime. Across noise levels, EDMD-kDL outperforms the two Koopman autoencoder methods, which produce non-smooth predictions. On the other hand, EDMD-RFF suffers from phase drift (see the left column of Figure 4 ) which leads to worse performance in terms of prediction error compared to EDMD-kDL when the amount of noise is large.
Figure 3 : Pendulum with measurement noise in the truly nonlinear regime. Across noise levels, EDMD-kDL outperforms the two Koopman autoencoder methods, which produce non-smooth predictions. On the other hand, EDMD-RFF suffers from phase drift (see the right column of Figure 4 ), which leads to worse performance in terms of prediction error compared to EDMD-kDL when the amount of noise is substantial.
Figure 4 : Pendulum with measurement noise. Comparison of predicted trajectories obtained by EDMD-kDL and three baseline methods over the final quarter of the prediction horizon. EDMD-RFF suffers from phase drift when the noise level is substantial whereas KAE and cKAE produce non-smooth predictions.
Figure 5 : Kármán vortex shedding . Relative prediction errors over the testing horizon measured in the SVD-reduced space.
Figure 6 : Kármán vortex shedding . Full-space predictions obtained by various Koopman methods. “Ground truth” refers to data obtained by projecting the original test data onto the retained principal components and reconstructing.
Figure 7 : Pendulum video data prediction . Predictions obtained by various Koopman methods in the pixel space. “Ground truth” refers to data obtained by projecting the original test data onto the retained SVD modes and reconstructing.
Figure 8 : Sea surface temperature forecasting . Scaled predictions at randomly sampled SST locations.
Figure 9 : Sea surface temperature forecasting . Relative prediction errors over time.
Figure 10 : Sea surface temperature forecasting . (a) Predicted global SST maps at specified time steps. (b) Associated absolute errors.
Studying nonlinear dynamical systems through their state space behavior can be challenging, and one possible alternative is to analyze them via their associated Koopman operator. This turns the nonlinear problem into a linear, infinite-dimensional one. To approximate the operator in finite dimensions, extended dynamic mode decomposition (EDMD) is a commonly used algorithm. It requires a finite list of functionals and a set of snapshots from the system to compute an approximation of the operator and its corresponding spectrum. Instead of choosing the list of functionals directly, it can be implicitly defined via kernels, a method known as kernel extended dynamic mode decomposition (kEDMD). However, one still needs to define the kernel and choose its parameter values. In this paper, we aim to streamline this process by extending dictionary learning for EDMD to kernel learning in kEDMD. By simplifying kEDMD we show how to perform gradient-based optimization over the learnable kernel parameters, and demonstrate that this method leads to useful kernels for the original kEDMD. The focus of our work is a method that takes a weighted list of kernels with randomly initialized values as input and outputs a list of kernels and parameter values suitable for approximating the Koopman operator of the underlying system. We demonstrate that unimportant kernels can be removed from the list by analyzing the weights in the weighted sum. We evaluate the method across several experiments, including the Duffing oscillator and the Kuramoto-Sivashinsky PDE, showcasing the method's different strengths.
Erik Lien Bolager, Boumediene Hamzi, Houman Owhadi +2
School of Computation, Information and Technology, Munich Data Science Institute, Munich Center for Machine Learning, Technical University of Munich, Germany · Department of Computing and Mathematical Sciences, Caltech, USA · The Alan Turing Institute, UK +1
Koopman theory turns nonlinear dynamics into a linear spectral problem. In computation, however, everything depends on a hard finite-dimensional choice: the observables must be expressive, nearly invariant under the dynamics, and, ideally, compatible with composition. Deep Koopman methods learn flexible coordinates, whereas structure-preserving methods enforce operator identities on fixed dictionaries. We combine these ideas by introducing Deep Embedded Multiplicative Dynamic Mode Decomposition (DeepMDMD), a method that learns a latent space and a partition of it, while enforcing the Koopman product rule as an exact algebraic constraint. Training alternates between an exact multiplicative operator update and a differentiable latent-clustering step that promotes Koopman closure. The result is a finite transition map on learned latent cells. Its nonzero spectrum lies on the unit circle, its dictionary is shaped by the dynamics rather than by ambient geometry, and forecasts are made in latent coordinates before being decoded to physical space. Across Hamiltonian, chaotic, and fluid examples, DeepMDMD learns dictionaries that are far more compact and dynamically coherent than those produced by geometric MDMD partitions. It reduces spectral pollution, reveals richer continuous-spectrum structure, and gives stable forecasts under severe noise. In high-dimensional flows, including a 158,624-dimensional cylinder wake and a noisy Re=20,000 lid-driven cavity, it preserves coherent structures and long-time spectral statistics where state-space MDMD fails. These results suggest a practical rule for Koopman learning: learn the coordinates, constrain the algebra.
Kelan Gray, Finlay Brown, Nicolas Boullé +1
Department of Mathematics, Imperial College London, London, SW7 2AZ, UK. · Department of Applied Mathematics and Theoretical Physics, University of Cambridge, Cambridge, CB3 0WA, UK.
Koopman theory promises linear structure in nonlinear dynamics, but numerical Koopman spectra are easy to compute and hard to trust. A finite EDMD matrix always has eigenvalues; the problem is that many of them may have nothing to do with the infinite-dimensional operator. In this paper we make spectral reliability the objective of dictionary learning. We train neural-network dictionaries not merely to predict the next snapshot, but to minimize Residual Dynamic Mode Decomposition residuals: operator-level a posteriori errors that test whether computed eigenvalues and modes are genuine Koopman spectral objects. To keep the learned observables from collapsing into an unstable coordinate system, the loss also penalizes the condition number of the lifted data matrix. Thus the method couples two requirements that should not be separated: small Koopman residuals and a well-conditioned representation. The result is a learned dictionary that is expressive, numerically stable, and spectrally disciplined. Across conservative and dissipative benchmark systems, the method sharply reduces spectral pollution, improves residual pseudospectral inclusion, and lowers forecast error relative to standard fixed dictionaries. On sea-surface temperature data, it gives cleaner Koopman diagnostics and substantially better one-step forecasts from noisy observations with no governing equations. The message is simple: neural Koopman learning should be judged not by prediction alone, but by whether its spectral claims can be certified. Residuals provide the certificate; conditioning makes it computable.
George Coote, Matthew J. Colbrook
Department of Applied Mathematics and Theoretical Physics, University of Cambridge, Wilberforce Road, CB3 0WA, United Kingdom