We propose Sufficiently Reduced Distributional Regression (SRDR), a generative method that combines conditional distribution estimation with nonlinear sufficient dimension reduction (SDR). It builds on a characterization of sufficiency through strictly proper scoring rules: a dimension reduction is sufficient if and only if predicting the response from the reduced covariates incurs no loss in expected score relative to the full covariates. Sufficient dimension reduction thus becomes a risk minimization problem. SRDR jointly trains a dimension reduction map and a generative prediction model by minimizing the energy score, which can be estimated by sampling without density evaluation or adversarial training. The framework extends to multi-environment data and to classification. We prove that the estimated conditional distributions converge in energy distance to the true ones, which implies that the learned representation is asymptotically sufficient. In simulations and applications to CT slice localization, superconductivity, and digit classification, SRDR recovers low-dimensional sufficient structure and matches or outperforms state-of-the-art nonlinear SDR methods in representation quality and predictive performance.
Figure 1 : Test energy score in the sparse-linear and sparse-nonlinear settings of Section 6.1 . SRDR-linear and SRDR-nonlinear refer to SRDR with a linear and a nonlinear dimension reduction map, respectively. Engression does not incorporate dimension reduction, whence its test scores do not depend on q .
Figure 2 : Distance correlation of the learned representation with two reference sufficient reductions, U=b1⊤X and cos(U) , in the heavy-tailed setting ( 5 ) of Section 6.2 .
Figure 3 : CT localization results by representation dimension q . Panel 3(a) reports MSE, and panel 3(b) reports CRPS for generative methods and MAE for deterministic methods.
Figure 4 : Patient-level distributional comparison for the CT localization dataset at q=3 . Each panel shows the histograms of the observed slice location and model output (generated samples or point predictions) distributions, together with their squared energy distance D2 .
Figure 5 : Superconductivity results by representation dimension q . Panel 5(a) reports MSE, and panel 5(b) reports CRPS for generative methods and MAE for deterministic methods.
Figure 6 : Distributional comparison for the superconductivity dataset at q=1 . Each panel shows the histograms of the observed critical temperature and model output (generated samples or point predictions) distributions for the two subgroups defined by the range of thermal conductivity, together with their squared energy distance D2 .
Figure 7 : Two-dimensional embeddings of MNIST images obtained from different SDR methods.
Figure 8 : Classification accuracy on the MNIST dataset by representation dimension q .
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 9 : Additional results of the synthetic experiments.
Figure 10 : dCor((f1(X),f2(X)),e^(X)) for the heteroscedastic simulation model. Results are reported as mean with error bar over 10 replicates.
Method
q
90% cov.
95% cov.
90% width
95% width
SRDR
1
0.878 (0.007)
0.928 (0.005)
29.945 (0.483)
36.212 (0.755)
SRDR
2
0.873 (0.006)
0.923 (0.005)
29.080 (0.532)
35.015 (0.915)
SRDR
4
0.873 (0.008)
0.924 (0.005)
28.541 (0.693)
34.440 (0.995)
SRDR
8
0.863 (0.007)
0.914 (0.006)
28.069 (0.553)
33.847 (0.655)
SRDR
16
0.865 (0.010)
0.915 (0.007)
27.626 (0.790)
33.305 (0.910)
SRDR
32
0.849 (0.010)
0.902 (0.008)
26.165 (0.666)
31.393 (0.823)
Appendix
Table 1 : Superconductivity prediction-interval results for SRDR, GenSDR, and engression. Entries are reported as mean (sd) across replicates.
Figure 11 : More histograms of the model outputs for the superconductivity experiment.
Figure 12 : Further histograms of the model outputs for the superconductivity experiment, for larger q .
Method
q
90% cov.
95% cov.
90% PI width
95% PI width
SRDR
3
0.857 (0.020)
0.903 (0.016)
1.283 (0.054)
1.520 (0.064)
SRDR
6
0.849 (0.022)
0.897 (0.016)
1.217 (0.065)
1.441 (0.077)
SRDR
12
0.848 (0.018)
0.897 (0.014)
1.190 (0.067)
1.409 (0.080)
SRDR
24
0.825 (0.016)
0.877 (0.015)
1.167 (0.087)
1.381 (0.102)
SRDR
48
0.836 (0.020)
0.887 (0.016)
1.101 (0.053)
1.303 (0.062)
SRDR
96
0.838 (0.015)
0.889 (0.012)
1.104 (0.070)
1.306 (0.083)
Appendix
Table 2 : CT prediction-interval results for SRDR, GenSDR, and engression. Entries are reported as mean (sd) across replicates.
Figure 13 : Histograms for the CT localization experiment.
Figure 14 : More histograms for the CT localization experiment.
Method
Source val.
Transfer
Scratch
Gap
CIFAR-10 → art painting
SRDR
0.543 (0.002)
0.733 (0.017)
0.724 (0.020)
0.010 (0.017)
SRDR-Frozen
0.543 (0.002)
0.294 (0.011)
0.725 (0.008)
-0.431 (0.012)
TESR
0.481 (0.005)
0.572 (0.032)
0.779 (0.016)
-0.207 (0.039)
CIFAR-10 → cartoon
SRDR
0.543 (0.002)
0.835 (0.011)
0.824 (0.022)
0.011 (0.019)
Appendix
Table 3 : CIFAR-10-to-PACS transfer learning classification results. Entries are reported as mean (sd) across seeded PACS resplits. Transfer gap is target accuracy minus scratch accuracy on the same held-out PACS test split.
Modern conditional generative models face significant challenges when learning complex covariate dependencies. While sufficient dimension reduction (SDR) provides a principled approach to compress these dependencies, traditional SDR frameworks were not formulated for conditional generation. To bridge this gap, we propose Belted Engression, a unified and architecturally parameter-efficient framework for generative distributional regression. Our approach establishes an end-to-end compress-then-generate paradigm driven by sufficient representation learning, embedding a structural bottleneck into the generative architecture. Theoretically, we prove that the standard SDR condition is equivalent to a law-preserving generative factorization, which is achieved at the global optimum of the population Belted Engression objective. Furthermore, by uncovering a localized Bernstein-type control for the energy-score loss, we establish finite-sample convergence rates that are sharper than those of existing results. We also prove that this belted architecture is strictly smaller, operating with an asymptotically vanishing parameter count relative to the unstructured baseline. Extensive simulations and real-world applications demonstrate that Belted Engression achieves superior distributional prediction and SDR recovery with fewer trainable parameters.
Wenxi Tan, Bing Li, Lingzhou Xue
Department of Statistics, The Pennsylvania State University
Learning representations that capture both intrinsic data geometry and target-relevant structure remains a fundamental challenge, particularly in settings where data reduction must balance compression with predictive fidelity. While distributional reduction-encompassing joint clustering and dimensionality reduction-offers a principled way to summarize data, its supervised variants remain relatively under-explored, despite the importance of retaining task-relevant signal for downstream prediction and decision-making. We propose Supervised Distributional Reduction (SDR), an algorithm for learning target-aware representations by combining optimal transport with explicit dependence maximization. SDR builds on the Fused Gromov-Wasserstein (FGW) objective to align the relational structure of the input distribution with a set of representative points, while augmenting it with a direct dependence term that encourages the learned embeddings to capture predictive signal more explicitly. This results in compact representations that reflect both geometric structure and supervision. Beyond representation learning, SDR naturally induces a data-dependent, non-stationary geometry that can be leveraged for settings such as Gaussian Process (GP) modelling. By redefining distances through target-aware distributional alignment, SDR enables the construction of adaptive kernels that respond to local variations in both data geometry and supervision, offering an optimal transport-based perspective on non-stationary kernel design.
Sufficient dimension reduction (SDR) makes high-dimensional regression tractable by projecting the covariates onto a low-dimensional subspace that preserves the conditional mean of the response. Existing gradient-based estimators either operate in the ambient space and suffer from the curse of dimensionality, or localize in the reduced space at a per-outer-iteration cost at least quadratic in the sample size. We show that minimizers of the population Minimum Average Variance Estimation (MAVE) risk approximate the same Grassmannian target as the Outer Product of Gradients (OPG), and recast the empirical criterion as a smooth maximization on the Stiefel manifold with closed-form Riemannian gradient. The resulting algorithm, SMAVE, combines sparse projected-space nearest-neighbor localization with Riemannian stochastic gradient ascent. A simplified version comes with almost-sure convergence and a non-asymptotic rate matching the standard non-convex stochastic first-order scaling. Empirically, SMAVE matches or improves on RMAVE's synthetic subspace recovery at moderate-to-high ambient dimension, and on four real datasets it uniformly improves over OPG and is competitive with or outperforms RMAVE at orders of magnitude lower runtime.
Thibault Pautrel, François Portier
Laboratoire des Signaux et Systèmes (L2S), CentraleSupélec, Université Paris-Saclay, Gif-sur-Yvette, France · ENSAI, CREST, Bruz, France