Organizations: Shanghai Center for Mathematical Sciences, Shanghai, China · Department of Mathematics, Fudan University, Shanghai, China · Center for Applied Mathematics, Fudan University, Shanghai, China · Alibaba Group, Hangzhou, China · Alibaba Group
Learning mappings between probability distributions arises naturally when inputs and outputs are represented by populations of samples rather than individual observations. We develop an approximation-theoretic framework for distribution-to-distribution learning and extend it to mappings between stochastic processes. For continuous operators on W2-compact families of finite-dimensional probability laws, we establish uniform neural approximation in the 2-Wasserstein metric using finite law statistics, a simplex-valued neural map, and a shared atomic output support that guarantees valid probability measures. We further extend this principle to probability laws on separable Hilbert spaces through finite-rank orthogonal projections. These results establish the representational feasibility of learning transformations between probability laws rather than deterministic vectors or functions. To demonstrate practical relevance, we study two problems naturally defined at the distribution level: prediction of first-passage-time distributions for an Ornstein--Uhlenbeck process and nonlinear response-path laws of a Duffing oscillator. Because the theory is model-agnostic and broader than any single practical architecture, the experiments use task-adapted neural models rather than reproducing the theoretical construction exactly. In both problems, the proposed models outperform a fixed-feature MLP baseline and distribution-space kernel regression. These experiments complement the theory by demonstrating the practical learnability of distribution-to-distribution transformations in random systems.
Figures & tables
Figure 1: A general example of the model pipeline. The model used for a specific task may differ in certain details, but broadly follows this design rationale. For task-specific model structures, see Figures 5 and 6 in the appendix.
Model
NLL
Hellinger
KL divergence
W2,fin
Tail error
Distributional Operator
3.11176±0.00085
0.18055±0.00048
0.10636±0.00085
0.16634±0.00286
0.00373±0.00018
Fixed-Feature MLP
3.11680±0.00011
0.18592±0.00013
0.11140±0.00011
0.20957±0.00026
0.00399±0.00003
Kernel regression
3.15241
0.21616
0.14701
0.34567
0.00639
Table 1: Test results for input process laws mapped to OU first-passage-time laws. Entries are the mean ± standard deviation over five runs with random initializations. Kernel regression is deterministic and therefore has zero standard deviation. Lower is better for every metric.
Figure 2: OU first-passage-time probability densities for three representative test laws. The continuous blue curve is the theoretical reference calculated from the Bessel-bridge representation. The step curves show the densities predicted by the Distributional Operator, kernel regression, and the Fixed-Feature MLP on the 48 finite time bins. Agreement with the reference curve compares how well the models recover the location, height, and tail of the first-passage-time density.
Model
Sinkhorn
MMD
SW2
Energy
Distributional Operator
8.2050±0.2786
0.0760±0.0010
0.1639±0.0036
0.3407±0.0104
Fixed-Feature MLP
9.9426±0.6795
0.0777±0.0017
0.1718±0.0036
0.3890±0.0173
Kernel regression
18.6291
0.1255
0.2498
0.8650
Table 2: Test performance on the Duffing dataset. Neural-model entries are the mean ± sample standard deviation on the 200-law test set. Kernel regression is a single fit with a fixed bandwidth. Sinkhorn denotes the Sinkhorn divergence, MMD the maximum mean discrepancy, SW2 the sliced 2-Wasserstein distance, and Energy the energy distance; lower values indicate better agreement.
Figure 3: Conditional distribution prediction for the Duffing oscillator. Target and predicted response distributions are shown for three randomly chosen test forcing measures. Solid blue lines and shaded regions denote the reference mean and 5th–95th percentile envelope, while dashed orange lines and hatched regions show the corresponding model predictions. Agreement in both statistics demonstrates the model’s accuracy in approximating the nonlinear pushforward from forcing to response distributions.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Predicted and empirical interspike-interval distributions. Ground-truth probability masses (left) and those predicted by the Process Distributional Operator (right) are shown for 50 test examples selected uniformly at random and ordered by increasing empirical mean interspike interval (ISI). Each heatmap row represents one input law, and the columns span the 48 finite ISI bins from 0 to 8. The narrow column beside each heatmap shows the probability of no threshold crossing before the end of the observation window. All panels use the same color scale.
Figure 5: Distributional Operator used for the OU first-passage task. An empirical law of PCA-compressed input paths is encoded by shared path and DeepSets networks, mapped to categorical output-law parameters, and sampled when realizations of the predicted first-passage distribution are required.
Figure 6: Process-valued Distributional Operator used for the Duffing task. Top: the input-process ensemble is projected to finite-dimensional coordinate laws and encoded into a condition through expectation features. Bottom: the condition and independent latent draws drive a conditional generator, whose projected output samples are reconstructed as response-process paths.
Model
NLL
W2
KL divergence
Hellinger distance
Distributional Operator
4.016462±0.002379
0.162049±0.004262
0.048246±0.001804
0.099313±0.002195
Fixed-Feature MLP
4.033581±0.002687
0.187644±0.003589
0.066297±0.002362
0.112535±0.002348
Kernel regression
4.776576±0.000000
0.741837±0.000000
0.806815±0.000000
0.396992±0.000000
Appendix
Table 3: Test results for the controlled Gaussian mapping. Entries are the mean ± standard deviation over five runs. Kernel regression is deterministic and therefore has zero standard deviation. Lower is better for every metric.
Figure 7: Predicted versus true output means on the test set for the trained Distributional Operator. Each panel shows one of the four mean components, with one point per test distribution; the dashed diagonal y=x denotes perfect prediction.
Figure 8: Elementwise mean absolute error of the predicted output covariance matrices on the test set for the trained Distributional Operator. For covariance entry (i,j) , the displayed value is Ntest−1∑k=1Ntest∣Σij(k)−Σij(k)∣ ; darker errors identify components that are systematically more difficult to predict.
Operations Research Center MASSachusetts Institute of Technology Cambridge, MA 02139 · Department of Computing and Mathematical Sciences California Institute of Technology Pasadena, CA 91125 · Laboratory for Information and Decision Systems Center for Computational Science and Engineering MASSachusetts Institute of Technology Cambridge, MA 02139 +2
Department of Mathematics, KTH Royal Institute of Technology, 100 44 Stockholm, Sweden. · Digital Futures, Sweden. · Grenoble INP – Ensimag, Université Grenoble Alpes, Grenoble, France. +1