Learning Conditional Expectation Operators via Functional Newton Updates
Authors: Thiago Ramos, Alek Fröhlich, Daniel Perazzo, Massimiliano Pontil
Organizations: Federal University of São Carlos · CSML, Istituto Italiano di Tecnologia and University of Genoa · CSML, Istituto Italiano di Tecnologia and University College London
We introduce the Functional Spectral-Newton Method (FSNM) for learning the leading singular structure of a conditional expectation operator without fixing a basis or reproducing kernel Hilbert space. FSNM fits a low-rank representation of the centered joint-to-product density ratio kernel by alternating functional Newton updates. Each update reduces to a preconditioned regression, which we approximate with vector-valued regression trees in a stagewise boosting procedure. At the population level, we establish descent and an O(1/T) best-iterate block-stationarity rate under a relative weak-learner accuracy condition, and show that every nondegenerate local minimum over the full centered L2 spaces is a globally optimal rank-d approximation. Synthetic experiments show that FSNM recovers a low-rank density ratio and its leading spectral structure, and that the same learned kernel can answer multiple conditional queries without refitting.
Figures & tables
Figure 1: Rank-three synthetic experiment with close singular values. From left to right: the exact density ratio κ , the FSNM estimate κ , and the signed error κ−κ .
Figure 2: Empirical training and validation losses over 40 boosting iterations. The vertical line marks the iteration selected by validation.
Method
Kernel RMSE
Spectrum error
Subspace error
FSNM
0.1261±0.0075
0.0286±0.0058
0.3579±0.0242
ACE
0.1489±0.0046
0.0238±0.0048
0.4010±0.0322
uLSIF
0.1149±0.0044
–
–
Kernel CCA
0.1756±0.0198
0.0720±0.0196
0.3613±0.0663
Table 1: Baseline comparison on the rank-three synthetic experiment, over 10 independent data draws (mean ±95% CI). Bold entries are statistically tied for best in that column, i.e., their interval overlaps the best mean’s interval. uLSIF does not factorize κ , so it has no associated spectrum or subspace error.
Figure 3: Discontinuous regional experiment. From left to right: the exact piecewise-constant density ratio, the FSNM estimate using tree weak learners, and their signed difference.
Method
Kernel RMSE
Spectrum error
Subspace error
FSNM
0.1393±0.0198
0.0301±0.0113
0.2678±0.0846
ACE
0.1583±0.0152
0.0353±0.0066
0.2978±0.0742
uLSIF
0.2312±0.0065
–
–
Kernel CCA
0.2908±0.0148
0.0716±0.0287
0.4878±0.0317
Table 2: Baseline comparison on the discontinuous regional experiment (mean ±95% CI, bold = tied for best, as in Table 1 ).
Figure 4: Twenty-dimensional tabular experiment. Left: exact and estimated density ratios on held-out product pairs. Center: exact and estimated spectra. Right: post-fit permutation sensitivity; red bars are the three relevant inputs and gray bars are the 17 irrelevant inputs.
Method
Kernel RMSE
Spectrum error
Subspace error
FSNM
0.1880±0.0227
0.0363±0.0114
0.4689±0.1062
ACE
0.1818±0.0152
0.0375±0.0065
0.3754±0.0689
uLSIF
0.7102±0.0042
–
–
Kernel CCA
0.6651±0.0180
0.3618±0.0117
0.6994±0.0088
Table 3: Baseline comparison on the twenty-dimensional tabular experiment with seventeen irrelevant input coordinates (mean ±95% CI, bold = tied for best, as in Table 1 ).
Figure 5: Conditional queries obtained from the learned rank-three kernel. From left to right: conditional mean, conditional variance, and conditional tail probability. Solid black curves are exact and dashed red curves are computed from κ .
Figure 6: Projected direct conditional CDF estimates and pointwise bootstrap bands at three values of x using tree weak learners. Solid black curves are exact, dashed blue curves are the fitted estimates, and shaded regions are pointwise 95% bands.
Figure 7: Independent case. Left: distribution of κ over independent reference pairs, with the dashed line marking the population value one. Right: estimated singular values, on the same vertical scale as Figure 8 .
Figure 8: Nonlinear dependent case Y=X2+ϵ . From left to right: test observations, learned centered density ratio kernel, and estimated singular values. Pearson’s correlation is near zero, whereas the learned kernel and the held-out permutation test detect dependence.
Group
Oxides
Motivation
Network formers
SiO 2 , B 2 O 3 , P 2 O 5
Backbone-forming oxides
Alkali modifiers
Li 2 O, Na 2 O, K 2 O
Network modification and fluxing
Optical-property components
TiO 2 , Nb 2 O 5 , Ta 2 O 5 , La 2 O 3
High-index optical-glass systems
Table 4: Ten-oxide family retained for the real-data experiment.
Figure 9: Fit of the restricted glass composition–RI. Left: training and validation losses for the selected rank- 20 trajectory; early stopping chooses iteration 39 . Right: fitted singular values and cumulative share of ∑jσj2 .
Figure 10: Prediction of RI using g(y)=y . Left: the dashed line is perfect prediction and the dotted line is the training mean. Right: mean observed and predicted within prediction octiles, with pointwise 95% normal intervals for the observed means.
Figure 11: Conditional queries. Left: for a random subset of 150 held-out compositions, the shaded bands are validation-calibrated 70% and 90% intervals. Right: reliability of the direct indicator query P(RI>1.8∣x) in eight equal-count bins, with pointwise 95% normal intervals for observed frequencies.
Conditional expectation operators (CEOs) and their associated conditional mean embeddings (CMEs) play a central role across applied mathematics and machine learning, appearing in nonparametric regression, Bayesian inverse problems, and Koopman operator theory. A fundamental question is when a CEO maps a function space on Y into a prescribed function space on X, particularly a reproducing kernel Hilbert space (RKHS). We show that such mapping properties are characterized by the regularity of the Radon--Nikodym density of the conditional law, and establish a simple, verifiable sufficient condition under which the CEO is bounded and Hilbert--Schmidt. For RKHSs norm-equivalent to Sobolev spaces, this condition reduces to Sobolev regularity of the conditional density. The result yields a direct route to validate CME representations and error bounds for Galerkin-type and CME-based estimators. We verify the regularity condition in three settings: nonparametric regression, Bayesian inverse problems, and Koopman operator theory for stochastic dynamical systems. We show in each case that classical regularity results on the underlying probabilistic model imply the required mapping properties. The resulting framework offers a unified perspective on conditional expectation operators across probability, operator theory, kernel methods, and stochastic dynamics.
Maximiliano Hertel, Ilja Klebanov, Manuel Schaller +1
Optimization-based Control Group, Institute of Mathematics, Technische Universität Ilmenau, Germany · Department of Mathematics and Computer Science, Freie Universität Berlin, Germany · Faculty of Mathematics, Chemnitz University of Technology, Germany
Probabilistic conditioning is concerned with the identification of a distribution of a random variable X given a random variable Y. It is a cornerstone of scientific and engineering applications where modeling uncertainty is key. This problem has traditionally been addressed in machine learning by directly learning the conditional distribution of a fixed joint distribution. This paper introduces a novel perspective: we propose to solve the conditioning problem by identifying a single operator that maps any joint density to its conditional, thus amortizing over joint-conditional pairs. We establish that the conditioning operator can be approximated to arbitrary accuracy by neural operators. Our proof relies on new results establishing continuity of the conditioning operator over suitable classes of densities. Finally, we learn the conditioning map for a class of Gaussian mixtures using neural operators, illustrating the promise of our framework. This work provides the theoretical underpinnings for general-purpose, amortized methods for probabilistic conditioning, such as foundation models for Bayesian inference.
Operations Research Center MASSachusetts Institute of Technology Cambridge, MA 02139 · Department of Computing and Mathematical Sciences California Institute of Technology Pasadena, CA 91125 · Laboratory for Information and Decision Systems Center for Computational Science and Engineering MASSachusetts Institute of Technology Cambridge, MA 02139 +2
We develop a stochastic approximation framework for learning nonlinear operators between infinite-dimensional spaces utilizing general Mercer operator-valued kernels. Our framework encompasses two key classes: (i) operator-valued kernels whose associated integral operators are compact and hence admit discrete spectral decompositions, and (ii) separable kernels of the form K(x,x′)=k(x,x′)T, where k is a scalar-valued kernel and T is a positive operator on the output space. This broad setting induces expressive vector-valued reproducing kernel Hilbert spaces (RKHSs) that generalize the classical K=kI paradigm, thereby enabling rich structural modeling with rigorous theoretical guarantees. To address target operators lying outside the RKHS, we introduce vector-valued interpolation spaces to precisely quantify misspecification error. Within this framework, we establish non-asymptotic convergence rates for prediction, estimation, and misspecification errors in the online and finite-horizon settings. Importantly, the framework also accommodates a range of operator learning settings, from Fredholm integral operators to encoder--decoder architectures. Numerical experiments on the two-dimensional Navier--Stokes equations illustrate the proposed approach.
Jia-Qi Yang, Lei Shi
School of Mathematical Sciences and Shanghai Key Laboratory for Contemporary Applied Mathematics, Fudan University, Shanghai 200433, China.