Estimating conditional statistics and learning representations of a population of conditional distributions are central problems in many data-driven applications, including uncertainty quantification and dynamical systems analysis. Conditional mean operators (CMOs), a class of linear operators between function spaces, resolve these objectives by providing access to a broad class of conditional statistics. However, existing methods typically estimate each CMO independently or constrain it to prespecified function spaces, thereby preventing the exploitation of shared structure across related distributions. In this work, we posit that related CMOs share finite-dimensional input and output function spaces, and are specialized for each task with a linear operator mapping these spaces. Based on this hypothesis, we introduce MTL-CMO, a multi-task framework that jointly learns shared function spaces and task-specific operators across multiple datasets. We further introduce T-CMO, a transfer learning method that reuses the shared spaces to estimate, in closed form, the operator of a new conditional distribution. We establish statistical guarantees quantifying the benefits of jointly learning the shared function spaces. Our experiments demonstrate that learning shared function spaces improves uncertainty quantification across a broad range of conditional distributions and, when applied to Langevin and plasma dynamics, yields compact representations of complex dynamics that retain physically meaningful information and enable parameter identification.
Figures & tables
Method
CD1
CD2
CD3
CD4
Multi-task
MTL-CMO (ours)
0.088 ± 0.004
0.102 ± 0.002
0.106 ± 0.003
0.070 ± 0.001
Pooled-NCP ( Kostic et al., 2024 )
0.091 ± 0.002
0.120 ± 0.009
0.139 ± 0.008
0.077 ± 0.008
MTL-MDN ( Bishop, 1994 )
0.110 ± 0.001
0.138 ± 0.005
0.216 ± 0.009
0.107 ± 0.005
DeepJMQR ( Rodrigues and Pereira, 2020 )
0.091 ± 0.002
0.114 ± 0.005
0.125 ± 0.008
0.066 ± 0.004
Single-task
NCP ( Kostic et al., 2024 )
0.096 ± 0.001
0.167 ± 0.037
0.185 ± 0.001
0.150 ± 0.002
CFM ( Lipman et al., 2023 )
0.112 ± 0.001
0.182 ± 0.021
0.170 ± 0.006
0.150 ± 0.005
Table 1: Mean 1-Wasserstein distance to the ground-truth conditional distribution averaged over 100 tasks. Scores in <mean>±<std> over 10 training seeds. Lower is better. First and second .
Figure 1: Comparison between MTL-CMO and the single-task NCP for representing CDFs on four conditional distribution families. Performances measured with the 1-Wasserstein to the ground truth.
Figure 2: Transfer learning evaluation on unseen conditional distributions. Comparison in 1 -Wasserstein between T-CMO and the single-task NCP depending on the number of observations.
Trajectory length N
Method
5 000
10 000
20 000
40 000
80 000
T-CMO (ours)
0.261±0.148
0.190±0.108
0.140±0.075
0.099±0.055
0.075±0.039
Single-task NCP ( Kostic et al., 2024 )
0.297±0.138
0.224±0.098
0.176±0.068
0.138±0.047
0.115±0.031
Pooled NCP ( Kostic et al., 2024 )
0.306±0.143
0.219±0.102
0.163±0.078
0.118±0.058
0.087±0.041
RFF/EDMD ( Li et al., 2017 )
0.295±0.140
0.233±0.123
0.177±0.069
0.135±0.045
0.109±0.030
RFF/Laplace ( Kostic et al., 2025 )
0.296±0.141
0.221±0.100
0.172±0.072
0.132±0.051
0.106±0.034
Table 2: Reconstruction error between estimated and ground truth operators. Hilbert-Schmidt error in <mean>±<std> over the 150 unseen systems. Lower is better. First and second .
Figure 3: Recovery of 3rd eigenfunction of an unseen Müller-Brown system with T-CMO/Laplace and its single-task counterpart NCP/Laplace. Global error in cosine distance.
Figure 4: Reconstruction error across eigenfunctions and potential values between T-CMO/Laplace and NCP/Laplace. Error averaged over the 150 unseen systems. Global error in cosine distance.
Figure 5: Left: t-SNE on the distance matrix between the 320 training operator spectra colored by g and κ . Right: parameter identification with kernel regression evaluated on the 80 held-out systems.
Figure 6: Identification of the parameters g and κ versus trajectory length for T-CMO and competing methods.
Appendix figures & tables25 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Representative conditional densities from CD1 , CD3 , and CD4 for different conditioning inputs and task parameters.
Hidden layers
d
rk
Activation
Dropout
CD1
3×64
64
20
GELU
0.15
CD2
2×64
64
7
Tanh
0.2
CD3
4×64
64
18
GELU
0.14
CD4
4×64
64
13
Tanh
0.16
Appendix
Table 3: Neural architecture and operator rank used for the four conditional-distribution families.
Epochs
Batch size
lrshared
lrspecific
wdshared
wdspecific
Grad. clip
Scheduler
CD1
5400
32
5.8×10−5
1.4×10−4
2.2×10−3
2.6×10−2
5.0
None
CD2
4200
64
2.7×10−4
6.4×10−5
3.9×10−3
2.6×10−4
2.0
None
CD3
3600
64
1.2×10−4
1.8×10−3
2.2×10−4
3.4×10−6
1.0
None
CD4
2600
64
3.0×10−3
2.9×10−3
9.1×10−6
4.4×10−3
2.0
Cosine
Appendix
Table 4: Optimization parameters for the uncertainty-quantification experiments.
ntr
50
100
150
200
250
300
350
400
CD1
NCP
0.2362 ± 0.0059
0.1739 ± 0.0025
0.1455 ± 0.0038
0.1258 ± 0.0021
0.1152 ± 0.0015
0.1062 ± 0.0021
0.1004 ± 0.0018
0.0948 ± 0.0024
T-CMO
0.2629 ± 0.0053
0.1869 ± 0.0015
0.1546 ± 0.0042
0.1315 ± 0.0012
0.1188 ± 0.0020
0.1072 ± 0.0022
0.1004 ± 0.0015
0.0943 ± 0.0015
NCP / T-CMO
0.90
0.93
0.94
0.96
0.97
0.99
1.00
1.01
CD2
NCP
0.2326 ± 0.0082
0.1826 ± 0.0041
0.1704 ± 0.0016
0.1588 ± 0.0015
0.1538 ± 0.0024
0.1500 ± 0.0007
0.1470 ± 0.0024
0.1458 ± 0.0028
Appendix
Table 5: Transfer to unseen conditional distributions. Wasserstein- 1 distance to the ground-truth conditional distribution (mean ± standard deviation over 10 target-data seeds), averaged over 100 unseen tasks. Lower is better. The last row of each block reports the ratio between single-task NCP and T-CMO.
Figure 8: Sensitivity of transfer performance on CD4 . Left: Wasserstein- 1 error as a function of the number of observations per source task n , with K=200 . Right: Wasserstein- 1 error as a function of the number of source tasks K , with n=200 .
T-CMO (ours)
NCP
Gain
Varying n (samples per task), K=200
n=50
0.281 ± 0.015
0.283 ± 0.013
1%
n=100
0.169 ± 0.005
0.229 ± 0.007
26%
n=200
0.099 ± 0.003
0.186 ± 0.005
47%
n=500
0.068 ± 0.002
0.143 ± 0.005
52%
n=1000
0.058 ± 0.001
0.128 ± 0.003
55%
Appendix
Table 6: Sensitivity of transfer performance on CD4 . Wasserstein- 1 distance to the ground-truth conditional distribution. Gain denotes the relative reduction in W1 of T-CMO with respect to single-task NCP. Lower is better.
Figure 9: Parametric Müller–Brown family. Top: varying r at fixed θ=0∘ modifies the relative depth of the third metastable basin while largely preserving its orientation. Bottom: varying θ at fixed r=0.95 rotates the local geometry of the same basin over a 60∘ range.
Figure 10: Spectral-rank diagnostic. Mean reduced-rank reconstruction error of the whitened lagged operator as a function of the retained rank. Most of the reduction occurs within the first three components, after which the error rapidly saturates.
Figure 11: Geometry of the learned Müller–Brown operators. t-SNE of pairwise S-GOT distances between operators learned during the multi-task stage, colored by r (left) and θ (right).
Figure 12: Multi-task versus single-task CMO learning. Eigenfunction errors of MTL-CMO versus independently trained single-task CMO models. Points below the diagonal favor MTL-CMO.
Configuration
Value
Architecture
Input dimension
2
Shared feature dimension d
64
Hidden layers
3×576
Activation
SiLU
Task-specific rank
20
Appendix
Table 7: Architecture, training, and spectral-estimation configuration for the Müller–Brown experiment.
Trajectory length N
Method
2 000
5 000
10 000
20 000
40 000
80 000
ψ1
T-CMO (ours)
0.0731±0.0966
0.0294±0.0414
0.0151±0.0206
0.0079±0.0095
0.0040±0.0053
0.0021±0.0026
RFF-EDMD
0.0982±0.1537
0.0333±0.0427
0.0248±0.0821
0.0098±0.0098
0.0058±0.0054
0.0037±0.0026
MetaKoopman
0.0837±0.1154
0.0327±0.0424
0.0174±0.0209
0.0099±0.0098
0.0058±0.0055
0.0040±0.0029
Single-task NCP
0.0839±0.1071
0.0332±0.0415
0.0177±0.0206
0.0101±0.0095
0.0059±0.0054
0.0041±0.0026
Appendix
Table 8: Eigenfunction errors for the first three non-trivial modes on the 150 unseen Müller–Brown systems. Values are mean ± standard deviation. Lower is better; best and second best per column.
Trajectory length N
Method
2 000
5 000
10 000
20 000
40 000
80 000
λ1
T-CMO (ours)
0.8288±2.8035
0.1919±0.2668
0.1241±0.1057
0.0916±0.0824
0.0643±0.0569
0.0443±0.0367
RFF-EDMD
0.9852±2.8284
0.3160±0.3197
0.2710±0.1682
0.1236±0.1025
0.2239±0.0935
0.1620±0.0706
MetaKoopman
1.0705±3.0177
0.2779±0.4465
0.1842±0.1614
0.1356±0.1260
0.1043±0.0853
0.0870±0.0689
Single-task NCP
1.0360±2.9714
0.2714±0.4275
0.1829±0.1489
0.1346±0.1168
0.1026±0.0793
0.0885±0.0590
Appendix
Table 9: Relative eigenvalue errors for the first three non-trivial modes on the 150 unseen Müller–Brown systems. Values are mean ± standard deviation. Lower is better; best and second best per column.
Figure 13: Transfer versus single-task spectral recovery. (a) Distribution of mode-wise 1−cosπ(ψj,ψj) errors over unseen systems. (b–d) Error distribution across potential-energy levels for ψ1 , ψ2 , and ψ3 .
Figure 14: Spatial recovery of the first three eigenfunctions on a representative unseen system. From left to right: finite-difference reference, T-CMO, single-task NCP, and the corresponding pointwise errors. Eigenfunctions are sign-aligned and normalized in L2(π) .
Figure 15: Representative electrostatic potential ϕ (top) and vorticity Ω=Δ⊥ϕ (bottom) across the (g,κ) domain. Each field is standardized independently for visualization.
Configuration
Value
Architecture
Input size
3×64×64
Shared feature dimension d
128
Task-specific rank
80
Encoder
Convolutional ResNet
Residual stages
(64,128,256)
Appendix
Table 10: Architecture and optimization configuration for the plasma experiment.
Figure 16: Evolution of the MTL-CMO training objective over 120 epochs.
Figure 17: Mean one-step RRR reconstruction error across the 320 source systems versus retained rank. The transfer rank is fixed to 32 .
Figure 18: MDS embeddings of the dynamical representations of the 320 source systems, colored by g (top) and κ (bottom).
Figure 19: Physical-parameter recovery on held-out systems versus trajectory length for T-CMO and competing representations.
Trajectory length T
256
474
878
1625
3008
4093
Interchange parameter g
T-CMO (ours)
0.225
0.516
0.778
0.878
0.897
0.932
MIDST
0.178
0.326
0.421
0.564
0.761
0.763
Pooled NCP
0.172
0.339
0.438
0.546
0.587
0.700
RFF (128)
−0.096
0.156
0.217
0.511
0.669
0.692
Density-gradient parameter κ
Appendix
Table 11: Held-out R2 for physical-parameter recovery versus available trajectory length. Lower-dimensional spectral methods use 62 features; MIDST uses its native 128 -dimensional representation.
Figure 20: Held-out R2 versus the number of slow generator modes retained from the T-CMO representation. Each mode contributes two spectral coordinates.
Features D
32
64
128
256
512
1024
Interchange g
0.388
0.662
0.714
0.641
0.639
0.639
Density gradient κ
0.968
0.979
0.986
0.985
0.983
0.982
Appendix
Table 12: Gaussian RFF sensitivity to the random-feature dimension D on complete trajectories.
We study the approximation and statistical complexity of learning collections of operators in a shared multi-task setting, with a focus on the Multiple Neural Operators (MNO) architecture. For broad classes of Lipschitz multiple operator maps, we derive near-optimal upper bounds for approximation and statistical generalization. On the lower-bound side, we establish a curse of parametric complexity and prove corresponding minimax rates. Together, these results show that shared representations across tasks do not increase the overall cost: multi-task operator learning follows the same scaling laws as single operator learning. We also compare MNO with a multi-task extension of DeepONet based on concatenated task inputs and show that, from a worst-case approximation-complexity perspective, both architectures satisfy essentially the same asymptotic rates.
Adrien Weihs, Hayden Schaeffer
Department of Mathematics, University of California Los Angeles, Los Angeles, CA 90095, USA.
Probabilistic conditioning is concerned with the identification of a distribution of a random variable X given a random variable Y. It is a cornerstone of scientific and engineering applications where modeling uncertainty is key. This problem has traditionally been addressed in machine learning by directly learning the conditional distribution of a fixed joint distribution. This paper introduces a novel perspective: we propose to solve the conditioning problem by identifying a single operator that maps any joint density to its conditional, thus amortizing over joint-conditional pairs. We establish that the conditioning operator can be approximated to arbitrary accuracy by neural operators. Our proof relies on new results establishing continuity of the conditioning operator over suitable classes of densities. Finally, we learn the conditioning map for a class of Gaussian mixtures using neural operators, illustrating the promise of our framework. This work provides the theoretical underpinnings for general-purpose, amortized methods for probabilistic conditioning, such as foundation models for Bayesian inference.
Operations Research Center MASSachusetts Institute of Technology Cambridge, MA 02139 · Department of Computing and Mathematical Sciences California Institute of Technology Pasadena, CA 91125 · Laboratory for Information and Decision Systems Center for Computational Science and Engineering MASSachusetts Institute of Technology Cambridge, MA 02139 +2
Multi-task learning (MTL) has emerged as a pivotal paradigm in machine learning by leveraging shared structures across multiple related tasks. Despite its empirical success, the development of likelihood-based efficiently solvable algorithms--even for shared linear representations--remains largely underdeveloped, primarily due to the non-convex structure intrinsic to matrix factorization. This paper introduces a first-order algorithm that jointly learns a shared representation and task-specific parameters, with guaranteed efficiency. Notably, it converges in O(1) iterations and attains a \emph{near-optimal} estimation error of O(dk/(TN)), \emph{improving} over existing likelihood-based methods by a factor of k, where d, k, T, N denote input dimension, representation dimension, task count, and samples per task, respectively. Our results justify that likelihood-based first-order methods can efficiently solve the MTL problem.
Shihong Ding, Fangyu Du, Cong Fang
State Key Lab of General AI, School of Intelligence Science and Technology, Peking University, China · School of Mathematical Sciences, Peking University, China · Institute for Artificial Intelligence, Peking University, China