Contextual Distributionally Robust Optimization with Causal and Continuous Structure
Authors: Fenglin Zhang, Jie Wang
Organizations: School of Artificial Intelligence The Chinese University of Hong Kong, Shenzhen Shenzhen, 518172, China · School of Artificial Intelligence, School of Data Science The Chinese University of Hong Kong, Shenzhen Shenzhen, 518172, China
We propose a framework for contextual distributionally robust optimization (DRO) that considers the causal and continuous structure of the underlying distribution, and we develop an interpretable and tractable decision rule. We first introduce the causal Sinkhorn discrepancy (CSD), an entropy-regularized causal Wasserstein distance that encourages continuous transport plans while preserving causal consistency. We then formulate a contextual DRO model with a CSD-based ambiguity set, termed Causal Sinkhorn DRO (Causal-SDRO), and derive its strong dual reformulation, where the worst-case distribution is characterized as a mixture of Gibbs distributions. To obtain an (infinite-dimensional) optimal policy, we propose a soft regression forest (SRF) decision rule: it preserves the interpretability of classical decision trees while being fully parametric, differentiable, and Lipschitz-smooth, enabling intrinsic interpretation from both global and local perspectives. To solve the Causal-SDRO with parametric decision rules, we develop an efficient stochastic compositional gradient algorithm that converges to an ε-stationary point at a rate of O(ε−4), matching that of standard stochastic gradient descent. Finally, we validate our method through numerical experiments on synthetic and real-world datasets, demonstrating its superior performance and interpretability.
Figures & tables
Figure 1 : Visualization for causal and non-causal Sinkhorn transport plans.
Figure 2 : Structure of distributions P (red points in 2a-2c), Pλ∗ , Pλ,SDRO∗ , Pλ,Causal-WDRO∗ , and Pλ,KL-DRO∗ . A covariate may correspond to multiple demand values. ( p=2 , λ=0.5 , ϵ=0.05 , and sample size N=100 )
Figure 3 : Structure of a soft regression tree t with depth D(t) .
Figure 4 : Out-of-sample performance of the decision rules on the newsvendor problem
Figure 5 : True distribution vs. Trained decision rules for 2-Causal-SDRO ( dx=10 )
Figure 6 : Out-of-sample performance of the newsvendor problem with different parameters ( N=400,dx=10 )
Figure 7 : Comparison of different DRO models on the newsvendor problem
Figure 8 : Out-of-sample performance of the decision rules on the inventory substitution problem
ω
Methods
Average Loss ↓
Sharpe ↑
Mean ↑
stdDev ↓
CVaR ↓
1
PT
0.468
4.125
0.176
0.784
1.971
EW
1.425
1.163
0.068
1.195
2.450
MV
1.009
1.578
0.072
1.008
1.971
CMV
0.984
1.156
0.061
1.003
2.108
1-Causal-SDRO
0.950
1.352
0.067
0.985
2.004
2-Causal-SDRO
0.925
1.559
0.081
0.978
1.965
Table 1 : Average performance for mean-variance methods
Figure 9 : Out-of-sample performance of the decision rules on the portfolio problem
Figure 10 : The structure of the trained SRT with three layers. Here, the solid black lines denote the model structure, while the red and blue dashed lines illustrate the decision routes and their corresponding selection probabilities for two input covariates x1 and x2 , respectively.
Figure 11 : Global feature importance comparison for the trained SRT
Figure 12 : Local feature attribution comparison for the trained SRT
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 13 : Out-of-sample Performance of the inventory substitution problem with different parameters ( N=200,dx=10 )
Figure 14 : Out-of-sample Performance of the portfolio problem with different parameters ( ω=5 )
Figure 15 : Comparison of different DRO models on the inventory substitute problem
Figure 16 : Comparison of different DRO models on the portfolio problem
We study distributionally robust optimization (DRO) for robust inference when the worst-case distribution is continuous, leading to significant computational challenges due to the infinite-dimensional nature of the optimization problem. Unlike traditional discrete DRO approaches, which often suffer from scalability issues, limited generalization, and costly worst-case inference, our framework exploits Brenier's theorem to characterize the least favorable distribution as the pushforward of a transport map from a continuous reference measure. This characterization motivates our study of the minimax problem in Wasserstein space. We propose an iterative algorithmic framework with multiple variants and establish global convergence guarantees under mild assumptions, deriving complexity bounds in terms of subgradient evaluations and inexact Jordan-Kinderlehrer-Otto updates. Numerical results with neural network-based transport maps demonstrate that the proposed method enables both stable training of robust classifiers and effective worst-case inference for classification tasks.
Linglingzhi Zhu, Yunqin Zhu, Yao Xie
1H. Milton Stewart School of Industrial and Systems Engineering Georgia Institute of Technology
Generative models are increasingly adopted in distributionally robust optimization (DRO), but existing approaches trade off model compatibility and adversarial structure: methods that accept arbitrary samplers do not restrict worst-case laws to a generator family, while generator-parameterized adversaries rely on model-specific access such as likelihoods, scores, or training data. We propose Generative Distributionally Robust Optimization (GDRO), a principled framework that accepts any sampleable conditional generator as the nominal model and restricts worst-case laws to a chosen conditional generator family. The key is the sampler-Sinkhorn pairing: samplers represent the conditional laws exactly, while Sinkhorn divergence compares their induced distributions without likelihood access and can be estimated from samples alone. The resulting population problem admits a direct finite-sample approximation and differentiable primal-dual implementation at the active decision context. For Lipschitz losses, the population Sinkhorn radius bounds downstream degradation. Across explicit and implicit generators, our method reduces rare-context inventory regret by 60% and SocialGAN navigation collisions by 50% relative to nominal decisions.
Ziwei Zhang, Jonathan Yu-Meng Li, Zhihao Jin
1Telfer School of Management, University of Ottawa · Department of Electrical and Computer Engineering, Western University
Distributionally robust optimization (DRO) is widely used for decision-making under uncertainty, but its adversarial focus on worst-case loss can lead to overly conservative policies. To mitigate this, we study ex-ante Distributionally Robust Regret Optimization (DRRO) with Wasserstein ambiguity sets, designed to balance robustness with upside potential. We develop a theory of Wasserstein DRRO (WDRRO) paralleling Wasserstein DRO. Under smoothness and regularity, WDRRO selects among ERM optima by a first-order gradient-discrepancy rule. If the ERM optimizer is unique, first-order sensitivity vanishes and a second-order expansion governs deviations. For convex quadratics ERM and DRRO coincide for any radius. We then study regimes where these assumptions fail: nondifferentiable max-affine losses, discrete references, and larger radii, where WDRRO can differ from ERM and WDRO. We show that computing WDRRO regret is NP-hard even without bilinear terms. Nevertheless, we develop exact algorithms, a tractable convex relaxation with guarantees, and experiments showing tightness and loss-dependent behavior.
Lukas-Benedikt Fiechtner, Jose Blanchet
Institute of Computational and Mathematical Engineering Stanford University · Institute of Computational and Mathematical Engineering Department of Management Science and Engineering Stanford University