Organizations: School of Mathematical Sciences, Fudan University, China. · Research Institute of Intelligent Complex Systems, Fudan University, China. · Shanghai Center for Mathematical Sciences, Fudan University, China. · Shanghai Artificial Intelligence Laboratory, China. · State Key Laboratory of Medical Neurobiology and MOE Frontiers Center for Brain Science, Institutes of Brain Science, Fudan University, China.
Many complex systems are observed only through temporally unpaired distribution snapshots, making trajectory-based dynamical learning difficult without additional assumptions. We therefore formulate the problem directly in distribution space, treating the distribution itself as the dynamical state. The challenge is that distribution space is infinite-dimensional, making compact and approximately closed representations difficult to learn from finite snapshots. We introduce DisKO, which extends deep Koopman learning to distribution dynamics by jointly learning predictive distributional observables, a finite-dimensional Koopman representation, and a generative map back to the full distribution. Across seven diverse benchmarks, DisKO achieves state-of-the-art extrapolation performance, with substantially slower error accumulation on long-horizon prediction tasks. DisKO further recovers leading Koopman eigenvalues and eigenfunctions on systems with analytic spectra, revealing meaningful dynamical structure in the learned representation.
Figures & tables
Method
Distribution snapshots
Learned observables
Distribution space
Full-distribution prediction
Spectral access
Deep Koopman ( Lusch et al., 2018 )
×
✓
×
×
✓
DeepRUOT ( Zhang et al., 2025b )
✓
×
×
✓
×
WLM ( Guan et al., 2026 )
✓
×
✓
✓
×
Distributional DMD ( Oprea et al., 2025 )
✓
×
✓
×
✓
DisKO
✓
✓
✓
✓
✓
Table 1: Comparison of methods for learning distribution dynamics.
Figure 1: Overview of DisKO. Distribution snapshots are encoded into learned observables, evolved under Koopman dynamics, and decoded back to distributions for reconstruction and extrapolation.
Method
Metric
OU
Duffing
Boids
Ocean
CAMELS
Pancreas
WOT iPSC
TrajectoryNet
W1
1.369 ± 0.071
9.222 ± 0.783
11.843 ± 0.560
2.198 ± 0.900
143.767 ± 2.170
47.376 ± 2.950
118.524 ± 2.509
SW1
0.857 ± 0.047
5.815 ± 0.712
7.267 ± 0.365
1.426 ± 0.556
35.364 ± 1.128
6.765 ± 0.826
14.502 ± 0.647
DeepRUOT
W1
0.846 ± 0.044
1.032 ± 0.287
5.962 ± 0.283
0.214 ± 0.020
79.989 ± 1.252
17.286 ± 1.572
23.635 ± 1.623
SW1
0.554 ± 0.030
0.650 ± 0.186
3.796 ± 0.172
0.149 ± 0.014
14.362 ± 0.466
2.113 ± 0.221
2.108 ± 0.220
VGFM
W1
1.290 ± 0.486
12.834 ± 0.976
31.008 ± 1.709
0.195 ± 0.047
123.071 ± 2.075
27.497 ± 1.337
187.154 ± 2.894
SW1
0.807 ± 0.315
9.358 ± 0.352
19.604 ± 1.428
0.128 ± 0.034
24.995 ± 17.602
2.978 ± 2.126
27.389 ± 1.762
Table 2: Results over the extrapolation window for test initial distributions in datasets with multiple initial distributions, and for single sequences. Red and blue indicate the best and second best results, respectively.
Figure 2: W1 prediction errors on four benchmarks. Vertical dashed lines mark the end of training, with the extrapolation window shaded. All error axes use a logarithmic scale.
Figure 3: Distributional eigenfunctions on wrapped Gaussian probes ( σ=0.15 ), indexed by center. (a) Torus. (b) Circle. Plots show real components and complex errors after amplitude and phase calibration on training snapshots.
Circle
Torus
DisKO
k=1
k=2
k=3
k=(1,0)
k=(0,1)
Discrete
1.315
3.191
5.928
2.404
2.892
Continuous
1.314
2.410
5.868
1.781
1.377
Table 3: Absolute eigenvalue errors ∣ρk−ρk∣ ( ×10−3 ) for K0.1 . Full eigenvalues and higher Circle harmonics appear in Appendix C.3 .
Figure 4: Reference encodings and forecasts for one OU test sequence, with one standard deviation across five seeds. PCA and rigid alignment use all times after forecasting.
Figure 5: OU errors averaged over 64 test initial distributions and five seeds. Shading shows one seed standard deviation; the dotted line marks the training boundary. Forecasts start at t=0 .
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
p
Sequences
Snapshots
Samples/snapshot
Training
Future
OU
2
192/64
101
1,024
0 – 50
51 – 100
Duffing
2
192/64
161
1,024
0 – 80
81 – 160
Boids
2
1
101
1,000
0 – 50
51 – 100
Ocean
2
1
11
400
0 – 9
10
CAMELS
10
21/6
9
4,096
0 – 5
6 – 8
Pancreas
30
1
8
3,562–12,477
0 – 5
6 – 7
Appendix
Table A.1: Distribution prediction datasets and temporal partitions. The observation dimension is p . Sequence counts are training/test counts for OU, Duffing, and CAMELS; the other datasets contain one sequence. Snapshot indices start at zero. Sample counts refer to the available data before evaluation subsampling.
Dataset
Evaluation samples
W1/W2 sample limit
OU, Duffing, Ocean
512
256
Boids
1,000
256
CAMELS, Pancreas, WOT iPSC
1,024
512
Appendix
Table A.2: Standard evaluation sample counts per distribution. SW1 uses the full evaluation sample and 128 projection directions; MMD 2 uses at most 512 samples per distribution. The DeepRUOT evaluations on Pancreas and CAMELS use the weighted configuration described below.
Dataset
State dim.
Enc. depth
Width
Latent dim.
Dec. depth
OU
2
3
128
32
4
Duffing
2
3
256
64
8
Boids
2
4
512
32
4
Ocean
2
3
128
32
4
CAMELS
10
3
256
64
8
Pancreas
30
3
256
96
8
Appendix
Table B.1: DisKO model configurations for the distribution extrapolation benchmarks. Enc. and Dec. denote encoder and decoder, respectively.
Dataset
Stage updates ( ×103 )
Stage learning rates ( ×10−4 )
Batch
Samples/ snapshot
OU
10/1/2.5
3/3/3
4
512
Duffing
30/3/15
3/3/3
4
512
Boids
30/30/75
3/1/1
4
512
Ocean
10/10/20
3/1/1
4
400
CAMELS
40/3/15
3/3/3
4
1024
Pancreas
60/3/24
3/3/3
8
512
Appendix
Table B.2: DisKO training and inference settings. Stage updates and learning rates follow the order pre/dyn/joint. The two panels report complementary settings for the same seven datasets. Loss weights follow the order specified in the text.
Figure B.1: W1 prediction error across the training and extrapolation intervals for (a) OU, (b) Duffing, (c) CAMELS, (d) Boids, (e) Ocean, (f) Pancreas, and (g) WOT iPSC. Curves compare DisKO, DeepRUOT, MIOFlow, and WLM. Vertical dashed lines mark the training boundary, and shaded regions indicate extrapolation.
Method
Metric
OU
Duffing
Boids
Ocean
CAMELS
Pancreas
WOT iPSC
TrajectoryNet
W1
1.536 ± 0.026
0.788 ± 0.067
4.715 ± 0.131
1.508 ± 0.098
46.276 ± 0.186
20.423 ± 0.497
35.166 ± 1.357
SW1
0.956 ± 0.029
0.493 ± 0.053
2.758 ± 0.308
0.979 ± 0.076
11.758 ± 0.333
2.488 ± 0.201
2.847 ± 0.291
W2
1.592 ± 0.027
0.843 ± 0.057
5.595 ± 0.135
1.513 ± 0.097
52.473 ± 0.268
20.984 ± 0.539
54.721 ± 2.050
MMD
0.813 ± 0.015
0.621 ± 0.102
0.264 ± 0.044
1.002 ± 0.038
0.175 ± 0.008
0.329 ± 0.054
0.666 ± 0.099
DeepRUOT
W1
1.221 ± 0.020
0.773 ± 0.145
4.342 ± 0.546
0.102 ± 0.031
20.784 ± 0.084
13.669 ± 1.096
18.648 ± 1.053
SW1
0.768 ± 0.024
0.484 ± 0.091
2.715 ± 0.358
0.064 ± 0.020
3.683 ± 0.103
1.340 ± 0.319
1.441 ± 0.178
Appendix
Table B.3: Prediction within the training time window for training initial distributions and single sequences. Boids, Ocean, Pancreas, and WOT iPSC each contain a single population sequence and use only a temporal partition. Means and standard deviations are rounded to three decimal places. Red and blue mark the lowest and second lowest means before rounding.
Method
Metric
OU
Duffing
Boids
Ocean
CAMELS
Pancreas
WOT iPSC
TrajectoryNet
W1
1.585 ± 0.044
0.751 ± 0.064
-
-
45.835 ± 0.506
-
-
SW1
0.985 ± 0.034
0.469 ± 0.050
-
-
11.567 ± 0.342
-
-
W2
1.648 ± 0.031
0.820 ± 0.055
-
-
53.090 ± 0.311
-
-
MMD
0.706 ± 0.022
0.499 ± 0.098
-
-
0.186 ± 0.015
-
-
DeepRUOT
W1
1.263 ± 0.034
0.772 ± 0.146
-
-
21.261 ± 0.725
-
-
SW1
0.797 ± 0.027
0.484 ± 0.092
-
-
3.622 ± 0.106
-
-
Appendix
Table B.4: Prediction within the training time window for test initial distributions. A hyphen (-) denotes the absence of a separate set of test initial distributions for a single sequence. Means and standard deviations are rounded to three decimal places. Red and blue mark the lowest and second lowest means before rounding.
Method
Metric
OU
Duffing
Boids
Ocean
CAMELS
Pancreas
WOT iPSC
TrajectoryNet
W1
1.279 ± 0.055
12.544 ± 1.069
-
-
145.995 ± 1.209
-
-
SW1
0.800 ± 0.033
7.876 ± 0.975
-
-
36.488 ± 1.768
-
-
W2
1.339 ± 0.052
26.279 ± 1.501
-
-
167.079 ± 1.926
-
-
MMD
0.656 ± 0.019
0.523 ± 0.138
-
-
0.212 ± 0.020
-
-
DeepRUOT
W1
0.765 ± 0.033
1.033 ± 0.290
-
-
78.561 ± 0.676
-
-
SW1
0.492 ± 0.020
0.651 ± 0.187
-
-
14.119 ± 0.687
-
-
Appendix
Table B.5: Temporal extrapolation from training initial distributions in datasets with multiple sequences. Hyphens (-) mark single sequences, whose temporal extrapolation results are reported in Table B.6 . Means and standard deviations are rounded to three decimal places. Red and blue mark the lowest and second lowest means before rounding.
Method
Metric
OU
Duffing
Boids
Ocean
CAMELS
Pancreas
WOT iPSC
TrajectoryNet
W1
1.369 ± 0.071
9.222 ± 0.783
11.843 ± 0.560
2.198 ± 0.900
143.767 ± 2.170
47.376 ± 2.950
118.524 ± 2.509
SW1
0.857 ± 0.047
5.815 ± 0.712
7.267 ± 0.365
1.426 ± 0.556
35.364 ± 1.128
6.765 ± 0.826
14.502 ± 0.647
W2
1.443 ± 0.067
17.953 ± 1.433
15.612 ± 0.973
2.204 ± 0.893
166.789 ± 2.871
53.487 ± 1.372
321.936 ± 2.588
MMD
0.666 ± 0.052
0.523 ± 0.140
0.348 ± 0.078
0.957 ± 0.074
0.233 ± 0.024
0.372 ± 0.101
0.628 ± 0.086
DeepRUOT
W1
0.846 ± 0.044
1.032 ± 0.287
5.962 ± 0.283
0.214 ± 0.020
79.989 ± 1.252
17.286 ± 1.572
23.635 ± 1.623
SW1
0.554 ± 0.030
0.650 ± 0.186
3.796 ± 0.172
0.149 ± 0.014
14.362 ± 0.466
2.113 ± 0.221
2.108 ± 0.220
Appendix
Table B.6: Temporal extrapolation from test initial distributions and single sequences. Boids, Ocean, Pancreas, and WOT iPSC each contain a single population sequence and use only a temporal partition. Means and standard deviations are rounded to three decimal places. Red and blue mark the lowest and second lowest means before rounding.
Setting
DisKO (discrete)
DisKO (continuous)
Encoder/decoder pretraining
40,000 shared steps
40,000 shared steps
Dynamics initialization
Full ridge fit
6,000 gradient steps
Joint optimization
10,000 steps
10,000 steps
Dynamics update
Full ridge refit at each step
AdamW, learning rate 3×10−4
Dynamics regularization
Ridge 10−3 on K,b
Zero weight decay on A,c
Encoder/decoder optimizer
AdamW, learning rate 3×10−4 , weight decay 10−5
Same
Appendix
Table C.1: Training configurations for the spectral experiments. Pretraining weights are shared within each dataset. The initialization procedures and dynamics update budgets differ between the two parameterizations.
Circle
k=1
k=2
k=3
k=4
k=5
k=6
Ground truth
0.9940+0.0997i
0.9762+0.1979i
0.9468+0.2929i
0.9064+0.3832i
0.8559+0.4676i
0.7962+0.5447i
DisKO (discrete)
0.9927+0.0996i
0.9731+0.1969i
0.9419+0.2895i
0.8974+0.3678i
0.8301+0.3715i
0.7937+0.3706i
DisKO (continuous)
0.9927+0.0997i
0.9738+0.1981i
0.9410+0.2919i
0.9125+0.1688i
0.8901+0.0941i
0.8627+0.0956i
Abs. error (discrete)
0.001315
0.003191
0.005928
0.017900
0.099453
0.174081
Abs. error (continuous)
0.001314
0.002410
0.005868
0.214545
0.375084
0.454035
Appendix
Table C.2: Eigenvalues of K0.1 and their approximations. Each column corresponds to one prescribed Fourier index; conjugate partners are omitted. For the continuous model, ρk=exp(0.1λk) . Complex values are rounded to four decimals, with real and imaginary parts on separate lines. Absolute errors are computed before rounding. All estimates use the final model from seed zero.
Figure C.1: Circle eigenvalue approximation at Δt=0.1 . Discrete and Continuous denote the two DisKO parameterizations. Circles show the analytical eigenvalues, stars the matched learned values, and gray points the remaining learned eigenvalues. Upper panels include conjugate partners; lower panels enlarge the selected representatives. Colors index k=1,…,6 . Both axes have equal scale. A matched star denotes the assigned estimate for a reference value.
Figure C.2: Torus eigenvalue approximation at Δt=0.1 , with the same conventions as Figure C.1 . The six Fourier indices identify the reference eigenpairs selected for evaluation. The gray points show the additional eigenvalues of the learned finite-dimensional operators.
Figure C.3: Eigenfunction values on observed future snapshots from test sequences. The top row shows Circle indices k=1,2,3 ; the remaining rows show all six Torus indices. Each panel overlays DisKO (discrete) and DisKO (continuous), with the equality line dashed. Coordinates display real parts after the training calibration; annotations give centered complex correlations. Each point is obtained by encoding the observed snapshot at that time, so the figure assesses the learned functions on future distributions.
Figure C.4: Circle eigenfunctions evaluated on translated wrapped Gaussian probes. Panels (a) and (b) display real and imaginary components using the same color scale. Each angle is the center of a distribution of width σ=0.15 ; ring radius and thickness are fixed. Columns compare the analytical reference with the discrete and continuous DisKO functions. NRMSE uses the full complex error and is consequently the same in both component panels.
Figure C.5: Torus eigenfunctions on wrapped Gaussian probes. Rows are the six Fourier indices; the first three columns show the reference and the two DisKO approximations. The final two columns show ∣hk−hk∣ for the discrete and continuous models. These complex error fields are identical in panels (a) and (b). All color ranges are shared, and surface coordinates index probe centers.
Figure D.1: OU latent evolution. Solid curves (Ground truth) show reference encodings; dashed curves (Koopman latent) show forecasts. Blue and gold indicate training and future time ranges; all snapshots belong to the test population. Curves are means over five models after PCA and rigid alignment. Shading is one sample standard deviation, normal to each mean curve in (a) and coordinatewise in (b,c). Stars, squares, and circles mark t=0 , 2.5 , and 5 ; filled and open markers denote reference and forecast.
Figure E.1: OU distribution prediction errors for the continuous Koopman model and the Neural ODE. At each time, errors are averaged over 64 test initial distributions within each run, then over five training seeds. Shading denotes one sample standard deviation across seeds. The dotted vertical line marks the end of training at t=2.5 . All predictions originate at t=0 .
Figure F.1: Effect of the number of training initial distributions R on future OU prediction errors. Colored circles and lines show means over three training seeds; shading denotes one sample standard deviation. Gray points show individual seeds, slightly displaced horizontally for visibility. Inputs contain 1,024 particles. The horizontal axis is logarithmic with base two.
Figure F.2: Effect of the available samples per training snapshot N , with R=128 training initial distributions and 128 particles per input. Means, sample standard deviations, and individual seeds follow Figure F.1 . Errors summarize the future time window. The horizontal axis is logarithmic with base two; the MMD2 axis in both figures uses units of 10−3 .
Department of Mathematics, Imperial College London, London, SW7 2AZ, UK. · Department of Applied Mathematics and Theoretical Physics, University of Cambridge, Cambridge, CB3 0WA, UK.
School of Mathematical Sciences, SCMS, SCAM, and CCSB, Fudan University, China · Research Institute of Intelligent Complex Systems, Fudan University, China · State Key Laboratory of Medical Neurobiology and MOE Frontiers Center for Brain Science, Institutes of Brain Science, Fudan University, China +1