Organizations: Shanghai Center for Mathematical Sciences, Shanghai, China · Department of Mathematics, Fudan University, Shanghai, China · Center for Applied Mathematics, Fudan University, Shanghai, China · Alibaba Group, Hangzhou, China · Alibaba Group
Learning mappings between probability distributions arises naturally when inputs and outputs are represented by populations of samples rather than individual observations. We develop an approximation-theoretic framework for distribution-to-distribution learning and extend it to mappings between stochastic processes. For continuous operators on W2-compact families of finite-dimensional probability laws, we establish uniform neural approximation in the 2-Wasserstein metric using finite law statistics, a simplex-valued neural map, and a shared atomic output support that guarantees valid probability measures. We further extend this principle to probability laws on separable Hilbert spaces through finite-rank orthogonal projections. These results establish the representational feasibility of learning transformations between probability laws rather than deterministic vectors or functions. To demonstrate practical relevance, we study two problems naturally defined at the distribution level: prediction of first-passage-time distributions for an Ornstein--Uhlenbeck process and nonlinear response-path laws of a Duffing oscillator. Because the theory is model-agnostic and broader than any single practical architecture, the experiments use task-adapted neural models rather than reproducing the theoretical construction exactly. In both problems, the proposed models outperform a fixed-feature MLP baseline and distribution-space kernel regression. These experiments complement the theory by demonstrating the practical learnability of distribution-to-distribution transformations in random systems.
Figures & tables
Figure 1: A general example of the model pipeline. The model used for a specific task may differ in certain details, but broadly follows this design rationale. For task-specific model structures, see Figures 5 and 6 in the appendix.
Model
NLL
Hellinger
KL divergence
W2,fin
Tail error
Distributional Operator
3.11176±0.00085
0.18055±0.00048
0.10636±0.00085
0.16634±0.00286
0.00373±0.00018
Fixed-Feature MLP
3.11680±0.00011
0.18592±0.00013
0.11140±0.00011
0.20957±0.00026
0.00399±0.00003
Kernel regression
3.15241
0.21616
0.14701
0.34567
0.00639
Table 1: Test results for input process laws mapped to OU first-passage-time laws. Entries are the mean ± standard deviation over five runs with random initializations. Kernel regression is deterministic and therefore has zero standard deviation. Lower is better for every metric.
Figure 2: OU first-passage-time probability densities for three representative test laws. The continuous blue curve is the theoretical reference calculated from the Bessel-bridge representation. The step curves show the densities predicted by the Distributional Operator, kernel regression, and the Fixed-Feature MLP on the 48 finite time bins. Agreement with the reference curve compares how well the models recover the location, height, and tail of the first-passage-time density.
Model
Sinkhorn
MMD
SW2
Energy
Distributional Operator
8.2050±0.2786
0.0760±0.0010
0.1639±0.0036
0.3407±0.0104
Fixed-Feature MLP
9.9426±0.6795
0.0777±0.0017
0.1718±0.0036
0.3890±0.0173
Kernel regression
18.6291
0.1255
0.2498
0.8650
Table 2: Test performance on the Duffing dataset. Neural-model entries are the mean ± sample standard deviation on the 200-law test set. Kernel regression is a single fit with a fixed bandwidth. Sinkhorn denotes the Sinkhorn divergence, MMD the maximum mean discrepancy, SW2 the sliced 2-Wasserstein distance, and Energy the energy distance; lower values indicate better agreement.
Figure 3: Conditional distribution prediction for the Duffing oscillator. Target and predicted response distributions are shown for three randomly chosen test forcing measures. Solid blue lines and shaded regions denote the reference mean and 5th–95th percentile envelope, while dashed orange lines and hatched regions show the corresponding model predictions. Agreement in both statistics demonstrates the model’s accuracy in approximating the nonlinear pushforward from forcing to response distributions.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Predicted and empirical interspike-interval distributions. Ground-truth probability masses (left) and those predicted by the Process Distributional Operator (right) are shown for 50 test examples selected uniformly at random and ordered by increasing empirical mean interspike interval (ISI). Each heatmap row represents one input law, and the columns span the 48 finite ISI bins from 0 to 8. The narrow column beside each heatmap shows the probability of no threshold crossing before the end of the observation window. All panels use the same color scale.
Figure 5: Distributional Operator used for the OU first-passage task. An empirical law of PCA-compressed input paths is encoded by shared path and DeepSets networks, mapped to categorical output-law parameters, and sampled when realizations of the predicted first-passage distribution are required.
Figure 6: Process-valued Distributional Operator used for the Duffing task. Top: the input-process ensemble is projected to finite-dimensional coordinate laws and encoded into a condition through expectation features. Bottom: the condition and independent latent draws drive a conditional generator, whose projected output samples are reconstructed as response-process paths.
Model
NLL
W2
KL divergence
Hellinger distance
Distributional Operator
4.016462±0.002379
0.162049±0.004262
0.048246±0.001804
0.099313±0.002195
Fixed-Feature MLP
4.033581±0.002687
0.187644±0.003589
0.066297±0.002362
0.112535±0.002348
Kernel regression
4.776576±0.000000
0.741837±0.000000
0.806815±0.000000
0.396992±0.000000
Appendix
Table 3: Test results for the controlled Gaussian mapping. Entries are the mean ± standard deviation over five runs. Kernel regression is deterministic and therefore has zero standard deviation. Lower is better for every metric.
Figure 7: Predicted versus true output means on the test set for the trained Distributional Operator. Each panel shows one of the four mean components, with one point per test distribution; the dashed diagonal y=x denotes perfect prediction.
Figure 8: Elementwise mean absolute error of the predicted output covariance matrices on the test set for the trained Distributional Operator. For covariance entry (i,j) , the displayed value is Ntest−1∑k=1Ntest∣Σij(k)−Σij(k)∣ ; darker errors identify components that are systematically more difficult to predict.
Many learning tasks map an input distribution to an output distribution. A natural way to model such an operator is to transform each input sample using a continuous function that may depend on the entire input distribution, and then take the distribution of the transformed samples. This defines a measure-dependent pushforward model and includes measure-theoretic formulations of transformers. We ask when such models can approximate arbitrary continuous operators between spaces of probability measures. We first show that universal approximation fails when atomic inputs are allowed: some continuous measure-to-measure operators that split or redistribute atomic mass cannot be approximated arbitrarily well by deterministic pushforward models. We then introduce the uniform level set condition, which requires a continuous measure-dependent scalarization whose shrinking level set neighborhoods carry uniformly vanishing mass over the input family. This condition is satisfied, in particular, by compact families of absolutely continuous measures. On every compact family satisfying this condition, we prove that any continuous measure-to-measure operator with outputs of finite p-th moment can be uniformly approximated, in the p-Wasserstein distance, by continuous measure-dependent pushforwards. Combining our theorem with existing approximation results for measure-dependent in-context maps yields universal approximation by measure-theoretic transformers. We also extend the framework to continuously-varying source measures, yielding a corresponding universality result for a class of pushforward models that are closely aligned with cross-attention architectures.
Probabilistic conditioning is concerned with the identification of a distribution of a random variable X given a random variable Y. It is a cornerstone of scientific and engineering applications where modeling uncertainty is key. This problem has traditionally been addressed in machine learning by directly learning the conditional distribution of a fixed joint distribution. This paper introduces a novel perspective: we propose to solve the conditioning problem by identifying a single operator that maps any joint density to its conditional, thus amortizing over joint-conditional pairs. We establish that the conditioning operator can be approximated to arbitrary accuracy by neural operators. Our proof relies on new results establishing continuity of the conditioning operator over suitable classes of densities. Finally, we learn the conditioning map for a class of Gaussian mixtures using neural operators, illustrating the promise of our framework. This work provides the theoretical underpinnings for general-purpose, amortized methods for probabilistic conditioning, such as foundation models for Bayesian inference.
Operations Research Center MASSachusetts Institute of Technology Cambridge, MA 02139 · Department of Computing and Mathematical Sciences California Institute of Technology Pasadena, CA 91125 · Laboratory for Information and Decision Systems Center for Computational Science and Engineering MASSachusetts Institute of Technology Cambridge, MA 02139 +2
We introduce conditional cylindrical neural networks for approximating functionals of conditional laws in McKean-Vlasov equations with common noise. Fourier moments of the initial law and truncated signatures of the time augmented common noise are mapped by a mixture density network to a Gaussian mixture approximation of the conditional law. A cylindrical neural network then evaluates the target functional through analytic integrals against this predicted measure. Rough path well posedness and stability provide a conditional law map that is continuous in the initial distribution and the rough driver and agrees almost surely with the classical conditional law at the Itô Brownian lift. Combining this continuity with Fourier separation, signature uniqueness, Wasserstein density of Gaussian mixtures, and neural universal approximation, we prove an L2 universal approximation theorem for continuous square integrable functionals. The numerical study implements the resulting two stage procedure on six examples, including non Gaussian initial laws, nonlinear drift, multiplicative common noise, and a two dimensional state. Independent particle references are used when no closed form law is available. The learned conditional law and functional approximations consistently improve on the empirical particle plug in, and additional experiments examine feature sensitivity, training from one terminal observation per common noise scenario, and Itô--Stratonovich consistency.
Nacira Agram, Reda Hmioui, Jan Rems
Department of Mathematics, KTH Royal Institute of Technology, 100 44 Stockholm, Sweden. · Digital Futures, Sweden. · Grenoble INP – Ensimag, Université Grenoble Alpes, Grenoble, France. +1