Estimating conditional statistics and learning representations of a population of conditional distributions are central problems in many data-driven applications, including uncertainty quantification and dynamical systems analysis. Conditional mean operators (CMOs), a class of linear operators between function spaces, resolve these objectives by providing access to a broad class of conditional statistics. However, existing methods typically estimate each CMO independently or constrain it to prespecified function spaces, thereby preventing the exploitation of shared structure across related distributions. In this work, we posit that related CMOs share finite-dimensional input and output function spaces, and are specialized for each task with a linear operator mapping these spaces. Based on this hypothesis, we introduce MTL-CMO, a multi-task framework that jointly learns shared function spaces and task-specific operators across multiple datasets. We further introduce T-CMO, a transfer learning method that reuses the shared spaces to estimate, in closed form, the operator of a new conditional distribution. We establish statistical guarantees quantifying the benefits of jointly learning the shared function spaces. Our experiments demonstrate that learning shared function spaces improves uncertainty quantification across a broad range of conditional distributions and, when applied to Langevin and plasma dynamics, yields compact representations of complex dynamics that retain physically meaningful information and enable parameter identification.
Figures & tables
Method
CD1
CD2
CD3
CD4
Multi-task
MTL-CMO (ours)
0.088 ± 0.004
0.102 ± 0.002
0.106 ± 0.003
0.070 ± 0.001
Pooled-NCP ( Kostic et al., 2024 )
0.091 ± 0.002
0.120 ± 0.009
0.139 ± 0.008
0.077 ± 0.008
MTL-MDN ( Bishop, 1994 )
0.110 ± 0.001
0.138 ± 0.005
0.216 ± 0.009
0.107 ± 0.005
DeepJMQR ( Rodrigues and Pereira, 2020 )
0.091 ± 0.002
0.114 ± 0.005
0.125 ± 0.008
0.066 ± 0.004
Single-task
NCP ( Kostic et al., 2024 )
0.096 ± 0.001
0.167 ± 0.037
0.185 ± 0.001
0.150 ± 0.002
CFM ( Lipman et al., 2023 )
0.112 ± 0.001
0.182 ± 0.021
0.170 ± 0.006
0.150 ± 0.005
Table 1: Mean 1-Wasserstein distance to the ground-truth conditional distribution averaged over 100 tasks. Scores in <mean>±<std> over 10 training seeds. Lower is better. First and second .
Figure 1: Comparison between MTL-CMO and the single-task NCP for representing CDFs on four conditional distribution families. Performances measured with the 1-Wasserstein to the ground truth.
Figure 2: Transfer learning evaluation on unseen conditional distributions. Comparison in 1 -Wasserstein between T-CMO and the single-task NCP depending on the number of observations.
Trajectory length N
Method
5 000
10 000
20 000
40 000
80 000
T-CMO (ours)
0.261±0.148
0.190±0.108
0.140±0.075
0.099±0.055
0.075±0.039
Single-task NCP ( Kostic et al., 2024 )
0.297±0.138
0.224±0.098
0.176±0.068
0.138±0.047
0.115±0.031
Pooled NCP ( Kostic et al., 2024 )
0.306±0.143
0.219±0.102
0.163±0.078
0.118±0.058
0.087±0.041
RFF/EDMD ( Li et al., 2017 )
0.295±0.140
0.233±0.123
0.177±0.069
0.135±0.045
0.109±0.030
RFF/Laplace ( Kostic et al., 2025 )
0.296±0.141
0.221±0.100
0.172±0.072
0.132±0.051
0.106±0.034
Table 2: Reconstruction error between estimated and ground truth operators. Hilbert-Schmidt error in <mean>±<std> over the 150 unseen systems. Lower is better. First and second .
Figure 3: Recovery of 3rd eigenfunction of an unseen Müller-Brown system with T-CMO/Laplace and its single-task counterpart NCP/Laplace. Global error in cosine distance.
Figure 4: Reconstruction error across eigenfunctions and potential values between T-CMO/Laplace and NCP/Laplace. Error averaged over the 150 unseen systems. Global error in cosine distance.
Figure 5: Left: t-SNE on the distance matrix between the 320 training operator spectra colored by g and κ . Right: parameter identification with kernel regression evaluated on the 80 held-out systems.
Figure 6: Identification of the parameters g and κ versus trajectory length for T-CMO and competing methods.
Appendix figures & tables25 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Representative conditional densities from CD1 , CD3 , and CD4 for different conditioning inputs and task parameters.
Hidden layers
d
rk
Activation
Dropout
CD1
3×64
64
20
GELU
0.15
CD2
2×64
64
7
Tanh
0.2
CD3
4×64
64
18
GELU
0.14
CD4
4×64
64
13
Tanh
0.16
Appendix
Table 3: Neural architecture and operator rank used for the four conditional-distribution families.
Epochs
Batch size
lrshared
lrspecific
wdshared
wdspecific
Grad. clip
Scheduler
CD1
5400
32
5.8×10−5
1.4×10−4
2.2×10−3
2.6×10−2
5.0
None
CD2
4200
64
2.7×10−4
6.4×10−5
3.9×10−3
2.6×10−4
2.0
None
CD3
3600
64
1.2×10−4
1.8×10−3
2.2×10−4
3.4×10−6
1.0
None
CD4
2600
64
3.0×10−3
2.9×10−3
9.1×10−6
4.4×10−3
2.0
Cosine
Appendix
Table 4: Optimization parameters for the uncertainty-quantification experiments.
ntr
50
100
150
200
250
300
350
400
CD1
NCP
0.2362 ± 0.0059
0.1739 ± 0.0025
0.1455 ± 0.0038
0.1258 ± 0.0021
0.1152 ± 0.0015
0.1062 ± 0.0021
0.1004 ± 0.0018
0.0948 ± 0.0024
T-CMO
0.2629 ± 0.0053
0.1869 ± 0.0015
0.1546 ± 0.0042
0.1315 ± 0.0012
0.1188 ± 0.0020
0.1072 ± 0.0022
0.1004 ± 0.0015
0.0943 ± 0.0015
NCP / T-CMO
0.90
0.93
0.94
0.96
0.97
0.99
1.00
1.01
CD2
NCP
0.2326 ± 0.0082
0.1826 ± 0.0041
0.1704 ± 0.0016
0.1588 ± 0.0015
0.1538 ± 0.0024
0.1500 ± 0.0007
0.1470 ± 0.0024
0.1458 ± 0.0028
Appendix
Table 5: Transfer to unseen conditional distributions. Wasserstein- 1 distance to the ground-truth conditional distribution (mean ± standard deviation over 10 target-data seeds), averaged over 100 unseen tasks. Lower is better. The last row of each block reports the ratio between single-task NCP and T-CMO.
Figure 8: Sensitivity of transfer performance on CD4 . Left: Wasserstein- 1 error as a function of the number of observations per source task n , with K=200 . Right: Wasserstein- 1 error as a function of the number of source tasks K , with n=200 .
T-CMO (ours)
NCP
Gain
Varying n (samples per task), K=200
n=50
0.281 ± 0.015
0.283 ± 0.013
1%
n=100
0.169 ± 0.005
0.229 ± 0.007
26%
n=200
0.099 ± 0.003
0.186 ± 0.005
47%
n=500
0.068 ± 0.002
0.143 ± 0.005
52%
n=1000
0.058 ± 0.001
0.128 ± 0.003
55%
Appendix
Table 6: Sensitivity of transfer performance on CD4 . Wasserstein- 1 distance to the ground-truth conditional distribution. Gain denotes the relative reduction in W1 of T-CMO with respect to single-task NCP. Lower is better.
Figure 9: Parametric Müller–Brown family. Top: varying r at fixed θ=0∘ modifies the relative depth of the third metastable basin while largely preserving its orientation. Bottom: varying θ at fixed r=0.95 rotates the local geometry of the same basin over a 60∘ range.
Figure 10: Spectral-rank diagnostic. Mean reduced-rank reconstruction error of the whitened lagged operator as a function of the retained rank. Most of the reduction occurs within the first three components, after which the error rapidly saturates.
Figure 11: Geometry of the learned Müller–Brown operators. t-SNE of pairwise S-GOT distances between operators learned during the multi-task stage, colored by r (left) and θ (right).
Figure 12: Multi-task versus single-task CMO learning. Eigenfunction errors of MTL-CMO versus independently trained single-task CMO models. Points below the diagonal favor MTL-CMO.
Configuration
Value
Architecture
Input dimension
2
Shared feature dimension d
64
Hidden layers
3×576
Activation
SiLU
Task-specific rank
20
Appendix
Table 7: Architecture, training, and spectral-estimation configuration for the Müller–Brown experiment.
Trajectory length N
Method
2 000
5 000
10 000
20 000
40 000
80 000
ψ1
T-CMO (ours)
0.0731±0.0966
0.0294±0.0414
0.0151±0.0206
0.0079±0.0095
0.0040±0.0053
0.0021±0.0026
RFF-EDMD
0.0982±0.1537
0.0333±0.0427
0.0248±0.0821
0.0098±0.0098
0.0058±0.0054
0.0037±0.0026
MetaKoopman
0.0837±0.1154
0.0327±0.0424
0.0174±0.0209
0.0099±0.0098
0.0058±0.0055
0.0040±0.0029
Single-task NCP
0.0839±0.1071
0.0332±0.0415
0.0177±0.0206
0.0101±0.0095
0.0059±0.0054
0.0041±0.0026
Appendix
Table 8: Eigenfunction errors for the first three non-trivial modes on the 150 unseen Müller–Brown systems. Values are mean ± standard deviation. Lower is better; best and second best per column.
Trajectory length N
Method
2 000
5 000
10 000
20 000
40 000
80 000
λ1
T-CMO (ours)
0.8288±2.8035
0.1919±0.2668
0.1241±0.1057
0.0916±0.0824
0.0643±0.0569
0.0443±0.0367
RFF-EDMD
0.9852±2.8284
0.3160±0.3197
0.2710±0.1682
0.1236±0.1025
0.2239±0.0935
0.1620±0.0706
MetaKoopman
1.0705±3.0177
0.2779±0.4465
0.1842±0.1614
0.1356±0.1260
0.1043±0.0853
0.0870±0.0689
Single-task NCP
1.0360±2.9714
0.2714±0.4275
0.1829±0.1489
0.1346±0.1168
0.1026±0.0793
0.0885±0.0590
Appendix
Table 9: Relative eigenvalue errors for the first three non-trivial modes on the 150 unseen Müller–Brown systems. Values are mean ± standard deviation. Lower is better; best and second best per column.
Figure 13: Transfer versus single-task spectral recovery. (a) Distribution of mode-wise 1−cosπ(ψj,ψj) errors over unseen systems. (b–d) Error distribution across potential-energy levels for ψ1 , ψ2 , and ψ3 .
Figure 14: Spatial recovery of the first three eigenfunctions on a representative unseen system. From left to right: finite-difference reference, T-CMO, single-task NCP, and the corresponding pointwise errors. Eigenfunctions are sign-aligned and normalized in L2(π) .
Figure 15: Representative electrostatic potential ϕ (top) and vorticity Ω=Δ⊥ϕ (bottom) across the (g,κ) domain. Each field is standardized independently for visualization.
Configuration
Value
Architecture
Input size
3×64×64
Shared feature dimension d
128
Task-specific rank
80
Encoder
Convolutional ResNet
Residual stages
(64,128,256)
Appendix
Table 10: Architecture and optimization configuration for the plasma experiment.
Figure 16: Evolution of the MTL-CMO training objective over 120 epochs.
Figure 17: Mean one-step RRR reconstruction error across the 320 source systems versus retained rank. The transfer rank is fixed to 32 .
Figure 18: MDS embeddings of the dynamical representations of the 320 source systems, colored by g (top) and κ (bottom).
Figure 19: Physical-parameter recovery on held-out systems versus trajectory length for T-CMO and competing representations.
Trajectory length T
256
474
878
1625
3008
4093
Interchange parameter g
T-CMO (ours)
0.225
0.516
0.778
0.878
0.897
0.932
MIDST
0.178
0.326
0.421
0.564
0.761
0.763
Pooled NCP
0.172
0.339
0.438
0.546
0.587
0.700
RFF (128)
−0.096
0.156
0.217
0.511
0.669
0.692
Density-gradient parameter κ
Appendix
Table 11: Held-out R2 for physical-parameter recovery versus available trajectory length. Lower-dimensional spectral methods use 62 features; MIDST uses its native 128 -dimensional representation.
Figure 20: Held-out R2 versus the number of slow generator modes retained from the T-CMO representation. Each mode contributes two spectral coordinates.
Features D
32
64
128
256
512
1024
Interchange g
0.388
0.662
0.714
0.641
0.639
0.639
Density gradient κ
0.968
0.979
0.986
0.985
0.983
0.982
Appendix
Table 12: Gaussian RFF sensitivity to the random-feature dimension D on complete trajectories.
Operations Research Center MASSachusetts Institute of Technology Cambridge, MA 02139 · Department of Computing and Mathematical Sciences California Institute of Technology Pasadena, CA 91125 · Laboratory for Information and Decision Systems Center for Computational Science and Engineering MASSachusetts Institute of Technology Cambridge, MA 02139 +2
State Key Lab of General AI, School of Intelligence Science and Technology, Peking University, China · School of Mathematical Sciences, Peking University, China · Institute for Artificial Intelligence, Peking University, China