Schrödinger bridge (SB) learns stochastic transport between prescribed initial and target distributions. When the initial distribution shifts at test time, the learned dynamics can fail to recover the target distribution. We introduce the Distributionally Robust Schrödinger Bridge (DRSB), which learns a single controller that accounts for uncertainty in the initial distribution. The DRSB objective consists of control energy and a KL penalty between the resulting terminal distribution and the target distribution. DRSB seeks a single controller that minimizes the worst-case value of this objective as the initial distribution varies within an ambiguity set around the nominal distribution. We derive an exact variational formulation of this objective and connect its fixed-terminal-cost subproblem to stochastic optimal control and distributionally robust optimization. This formulation motivates an alternating algorithm that updates the adversarial initial distribution, estimates the terminal log-density ratio, and trains the controller. We develop Wasserstein and Sinkhorn variants using stochastic control optimality conditions to approximate the gradients required for adversarial updates. Experiments on two-dimensional transport tasks and image-to-image translation show improved robustness to input perturbations relative to standard SB, with a tradeoff in nominal performance. On Gaussian mixture transport, Sinkhorn DRSB also achieves lower mean sliced Wasserstein distance than fixed-level noise augmentation at both tested unseen noise levels.
Figures & tables
Figure 1: Deployment under initial distribution uncertainty. Blue circles denote nominal samples from μ , and purple squares denote shifted test samples from μtest . The desired target ν is unchanged. (a) A standard SB controller matches ν for μ , but this matching need not persist under an initial distribution shift. (b) DRSB trains one controller to reduce the worst-case sum of control energy and terminal mismatch over an ambiguity set of initial distributions.
Gaussian-to-Gaussian
GMM-to-GMM
Input
DSBM
DRSB
DSBM
DRSB
Org
0.038 ± 0.007
0.123 ± 0.012
0.902 ± 0.242
1.165 ± 0.138
Noisy
0.684 ± 0.013
0.638 ± 0.016
2.407 ± 0.243
2.212 ± 0.413
Shifted
0.666 ± 0.011
0.501 ± 0.014
5.917 ± 0.127
5.840 ± 0.095
Worst
0.951 ± 0.013
0.710 ± 0.013
1.988 ± 0.100
1.478 ± 0.100
Table 1: SWD ( ↓ ) under input perturbations (mean ± std, 10 seeds). Org denotes the original input; the other conditions are defined in the text.
Figure 3
DSBM
DSBM + Aug
DRSB-S
δ
δtrain=1
δtrain=2
Org
0.902 ± 0.242
0.850 ± 0.178
0.907 ± 0.153
1.165 ± 0.138
1.0
1.287 ± 0.112
1.042 ± 0.246
1.119 ± 0.196
1.150 ± 0.206
2.0
1.842 ± 0.086
1.881 ± 0.145
1.800 ± 0.153
1.636 ± 0.157
3.0
2.175 ± 0.118
2.298 ± 0.146
2.220 ± 0.165
1.953 ± 0.095
4.0
2.407 ± 0.243
2.425 ± 0.129
2.308 ± 0.138
2.212 ± 0.413
Table 2: GMM-to-GMM SWD ( ↓ ) with noise augmentation (mean ± std, 10 seeds). Test levels δ=3,4 are unseen by both augmentation baselines.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Training dynamics. (a) Evolution of λ in DRSB-W. (b) ε -schedule for DRSB-S and corresponding λ evolution.
Model
Org
δ=0.25
0.50
1.00
DSBM
25.55 ± 0.50
26.59 ± 0.48
26.15 ± 0.49
26.81 ± 0.52
DRSB-S κ =0.005, ε =9
26.76 ± 0.49
26.28 ± 0.44
26.60 ± 0.52
27.15 ± 0.49
DRSB-S κ =0.005, ε =12
26.36 ± 0.48
26.48 ± 0.48
26.10 ± 0.50
26.66 ± 0.52
DRSB-S κ =0.005, ε =15
26.08 ± 0.47
26.57 ± 0.50
26.35 ± 0.53
26.40 ± 0.53
DRSB-S κ =0.32, ε =7
26.70 ± 0.48
25.93 ± 0.48
26.05 ± 0.47
26.05 ± 0.50
DRSB-S κ =0.32, ε =12
26.52 ± 0.50
25.82 ± 0.49
25.98 ± 0.49
25.94 ± 0.47
Appendix
Table 4: FID ( ↓ ) sensitivity to κ and ε on FFHQ Woman → Man under noise δ .
Model
δ =0.5
1.0
1.5
2.0
2.5
3.0
3.5
4.0
(a) Gaussian-to-Gaussian
DSBM
0.052 ± 0.006
0.072 ± 0.009
0.128 ± 0.012
0.214 ± 0.010
0.313 ± 0.015
0.431 ± 0.012
0.555 ± 0.013
0.684 ± 0.013
DRSB-W
0.080 ± 0.010
0.093 ± 0.012
0.138 ± 0.013
0.216 ± 0.013
0.315 ± 0.012
0.425 ± 0.013
0.548 ± 0.013
0.678 ± 0.013
DRSB-S
0.131 ± 0.011
0.134 ± 0.014
0.165 ± 0.015
0.227 ± 0.016
0.314 ± 0.015
0.414 ± 0.015
0.522 ± 0.015
0.638 ± 0.016
(b) GMM-to-GMM
DSBM
1.164 ± 0.137
1.287 ± 0.112
1.597 ± 0.077
1.842 ± 0.086
2.051 ± 0.116
2.175 ± 0.118
2.254 ± 0.108
2.407 ± 0.243
Appendix
Table 5: SWD ( ↓ ) on 2D transport tasks under varying L2 -norm noise levels δ (mean ± std over 10 seeds). Bold indicates best performance; gray indicates outperforming the baseline DSBM.
Model
Org
Up
Down
Left
Right
W-wst
S-wst
(a) Gaussian-to-Gaussian
DSBM
0.038 ± 0.007
0.665 ± 0.012
0.673 ± 0.020
0.666 ± 0.011
0.651 ± 0.009
1.081 ± 0.021
0.951 ± 0.013
DRSB-W
0.144 ± 0.007
0.691 ± 0.020
0.698 ± 0.020
0.517 ± 0.014
0.792 ± 0.014
0.972 ± 0.007
0.812 ± 0.012
DRSB-S
0.123 ± 0.012
0.647 ± 0.018
0.640 ± 0.014
0.501 ± 0.014
0.769 ± 0.016
0.881 ± 0.018
0.710 ± 0.013
(b) GMM-to-GMM
DSBM
0.902 ± 0.242
6.167 ± 0.233
5.798 ± 0.245
5.917 ± 0.127
5.917 ± 0.132
2.086 ± 0.166
1.988 ± 0.100
Appendix
Table 6: SWD ( ↓ ) on 2D transport tasks under distribution shifts (mean ± std over 10 seeds). Org: original distribution; Up/Down/Left/Right: directional shift; W-wst: DRSB-W worst case; S-wst: DRSB-S worst case. Bold indicates best performance; gray indicates outperforming the baseline DSBM.
Figure 6: Cost comparison within the Wasserstein ambiguity set. Left : Evaluated distributions with the ambiguity-set boundary (dashed). Right : Measured cost versus ρ .
Gaussian DSBM
Gaussian DRSB
FFHQ DSBM
FFHQ DRSB
Iterations
140
50
70
50
Gradient steps (total)
224k
101k
3.36M
320k
Trajectories
4,096
1,024
20,480
40,960
Wall-clock time
754 s
404 s
8.5 h
1.4 h
Stage time / DSBM time
1×
0.54×
1×
0.16×
Appendix
Table 7: Measured training costs. DRSB columns report the additional stage after DSBM pretraining. Trajectory counts follow the reported training configurations.
k
5
10
15
20
25
30
35
40
45
50
Rk
0.260
0.319
0.240
0.111
0.002
0.004
0.002
0.000
0.002
0.001
λk
6.601
5.696
4.535
3.627
3.427
3.457
3.477
3.488
3.494
3.483
Gk
0.752
1.232
1.735
2.142
2.068
1.757
1.672
1.631
1.524
1.550
Uk
22.94
22.92
23.19
23.74
24.10
24.22
24.29
24.36
24.32
24.30
Appendix
Table 8: Reported update diagnostics on the 2D Gaussian task.
k
5
10
20
30
40
50
L
270.99
276.16
304.99
311.99
311.44
311.37
Appendix
Table 9: Variational objective at the reported epochs.
k
10
20
30
40
50
cos(∇X0V,−uθ(0,X0))
0.9997
0.9996
0.9996
0.9997
0.9998
∥∇X0V∥/∥uθ(0,X0)∥
0.9310
1.0761
1.0249
1.0278
1.0089
Appendix
Table 10: SOC consistency diagnostics on the 2D Gaussian task.