We study the problem of recovering the governing ODE of a dynamical system from unstructured, high-dimensional observations such as images. Existing methods for ODE discovery typically assume direct measurements of the variables, or do not provide theoretical guarantees on the learned variables and equations. While Causal Representation Learning (CRL) methods provide guarantees on identifying variables from high-dimensional observations up to component-wise diffeomorphisms, we show that in general these variables cannot be used directly as input to equation discovery methods, which typically assume that the variables will lead to sparse equations. So we introduce SParse Equivalent Equation Discovery AutoEncoder (SPEED-AE), a framework that combines a pretrained CRL method with a component-wise autoencoder that learns transformations of variables that are amenable to sparse ODE discovery. We show that for polynomial ODEs, this additional step allows us to restrict the identifiability of each variable from polynomial to monomial diffeomorphisms. Experiments on Lotka-Volterra, Lorenz, and a two-pendulum system show that SPEED-AE improves on the disentanglement of the CRL methods and that it recovers ODEs that are closest to the ground truth, while achieving state-of-the-art forecasting performance.
Figures & tables
Figure 1: SPEED-AE pipeline: we observe high-dimensional data from an invertible mapping x(t)=g(z(t)) of the true latents z(t) . First, SPEED-AE employs a pretrained CRL model to obtain low-dimensional latents z^(t) ( red block). We assume that the recovered latents correspond to the ground-truth z up to a permutation and component-wise diffeomorphism zi=h^i(z^π(i)) . For simplicity, in the figure we assume that the permutation is the identity function. SPEED-AE learns a set of component-wise autoencoders (ϕienc,ϕidec) and an ODE by combining the functions in a library Θ(z~) , imposing sparsity of this representation ( blue block). The output latents z~(t) identify the ground-truth up to permutation and monomial and, in most cases, linear transformation.
Figure 2: Comparison between SINDyAE, SINDy applied to CRL latents (using CITRIS), and SPEED-AE (with CITRIS as CRL method and SINDy as equation discovery method) on the Lotka-Volterra dataset ( z1˙=0.66z1−1.33z1z2 , z2˙=−z2+z1z2 ). For each model, the left box reports scatter plots of the true versus learned latents. We then compare a test trajectory with the one predicted from only the initial condition and the discovered equation, after converting both to the true latent space, following the evaluation protocol in Section 6 . SPEED-AE is the only method that recovers a sparse equation and achieves strong long-term prediction performance.
Figure 3: Learned coefficients for Lotka-Volterra for each polynomial of the state variables (each row) and different methods (each column, where the last four are versions of SPEED-AE). The first column shows the ground truth coefficients and the black boxes highlight the non-zero coefficients.
Figure 4: Results for the Lotka-Volterra and Lorenz experiments with mean and standard deviation over 10 seeds. For all metrics, lower is better. We limit the height of the boxes for better readability, reporting the actual scales of the outlier results inside their bars.
Table 1: Results for the Pendulum experiment for the model.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter
Lotka-Volterra
Lorenz
Pendulum
ntrain
20
1024
100
nval
5
256
25
ntest
5
256
25
T
50
5
10
Δt
0.01
0.02
0.02
num. steps
5000
250
500
Appendix
Table 2: Data generation parameters.
Model
Training Time (mins)
LSTM
≈45
Transformer
≈62
SINDyAE
≈100
MNNAE
≈191
Autoencoder (CITRIS-NF)
≈22
CITRIS-NF
≈52
Appendix
Table 5: Training time for each model on the Lotka-Volterra experiment.
Figure 5: Evaluation pipeline of our experiments.
Figure 6: Experimental results for the ablation on the independence from CRL method. Lotka-Volterra dataset with an additional version of SPEED-AE with an oracle CRL model that achieves perfect disentanglement and SINDy-based loss.
Figure 7: Results on the synthetic experiments with perfect disentanglement, polynomial identifiability in the CRL latents, and SPEED-AE encoder and decoders implemented as polynomials. (a) Some of the runs show strong instabilities, while others (b) remain stable but with flat dynamics. In general, the polynomial functions are hard to optimize simultaneously to the dynamics.
Figure 8: Experimental results for the Lotka-Volterra and the Lorenz experiments.
Model
Lotka-Volterra
Lotka-Volterra
Lorenz (Stable)
Lorenz (Chaotic)
α=32,β=34
α=56,β=23
ρ=14
ρ=28
LSTM
13.860 ±2.609
29.313 ±14.719
0.157 ±0.056
1.624 ±0.210
Transformer
16.001 ±2.076
25.999 ±6.919
0.422 ±0.076
1.902 ±0.210
SINDyAE
23.075 ±11.611
17.610 ±2.898
1.988 ±0.323
3.137 ±0.120
MNNAE
1.257 ±0.240
0.930 ±0.290
0.448 ±0.102
≈105 (failed)
SPEED-AE (C+S)
1.279 ±0.483
3.999 ±2.390
0.851 ±0.384
2.411 ±0.379
Appendix
Table 6: x Err: mean and standard deviation across 10 seeds.
Figure 9: Pearson cross correlation matrices between z~ and z for all experiments and models. The rightmost matrix in each subfigure corresponds to the true correlation matrix (i.e., between z and z ).
Figure 10: Lotka-Volterra experiment with α=32,β=34 . Scatter plots of the found latents z~ against the true ones z . If their function is well defined, the plots well represent the maps zi=hi(z~i) . In the second row, we include the same plots for the CRL latents z^ from CITRIS and DMSVAE. In most cases, SPEED-AE can go from a general diffeomorphism h to a linear one.
Figure 11: Lotka-Volterra experiment woth α=56,β=23 . Scatter plots of the found latents z~ against the true ones z . If their function is well defined, the plots well represent the maps zi=hi(z~i) . In the second row, we include the same plots for the CRL latents z^ from CITRIS and DMSVAE. In most cases, SPEED-AE can go from a general diffeomorphism h to a linear one.
Figure 12: Identified coefficients for the Lotka-Volterra experiment with transfer of CRL models and α=56,β=23 . The true ones are in the leftmost column and highlighted by the black borders. For comparison, we include four baselines that consist of equation discovery directly applied to the CRL latents z^ .
Figure 13: Example test trajectory of the Lotka-Volterra experiment (first 10 dimensions of x ). Ground truth (blue) against the predicted trajectory from the LSTM model (red). During rollout, the model collapses after around 500 steps.
Figure 14: Example test trajectory of the Lotka-Volterra experiment (first 10 dimensions of x ). Ground truth (blue) against the predicted trajectory from the Transformer model (red). During rollout, the model collapses after a few iterations.
Figure 15: Lorenz experiment with stable trajectories ρ=14 . Scatter plots of the found latents z~ against the true ones z . If their function is well defined, the plots well represent the maps zi=hi(z~i) . In the second row, we include the same plots for the CRL latents z^ from CITRIS and DMSVAE. SPEED-AE models with DMSVAE as disentangler suffer more due to the CRL latents being less disentangled, although the SINDy-based loss is actually able to recover a linear map h .
Figure 16: Scatter plots of the found latents z~ against the true ones z . If their function is well-defined, the plots well represent the maps zi=hi(z~i) . In the second row, we include the same plots for the CRL latents z^ from CITRIS and DMSVAE. SPEED-AE models with DMSVAE as disentangler suffer more due to the CRL latents being less disentangled, although the SINDy-based loss is actually able to recover a linear map h .
Figure 17: Identified coefficients for the Lorenz experiments in the stable ( ρ=14 , left) and chaotic case ( ρ=28 , right). The true ones are in the leftmost column and highlighted by the black borders. For comparison, we include four baselines that consist of equation discovery directly applied to the CRL latents z^ .
Figure 18: Example images from the experiment with two pendulums. Pendulums are colored differently to ensure invertibility of the mixing function.
Figure 19: Pendulums experiment. Scatter plots of the found latents z~ against the true ones z . If their function is well defined, the plots well represent the maps zi=hi(z~i) . In the second plot, we include the same for the CRL latents z^ from CITRIS. SPEED-AE is able to achieve linear identifiability from the general diffeomorphism of CITRIS. SINDyAE only identifies one variable, while the other likely represents a mapping between the two pendulums.
Department of Computer Science and Applied Mathematics, Weizmann Institute of Science, Rehovot, Israel · Department of Molecular Cell Biology, Weizmann Institute of Science, Rehovot, Israel · Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, United Arab Emirates +2