Optimal Transport (OT) provides a principled framework for learning transformations between probability distributions from unpaired samples. In many applications, however, a single transformation must map several source distributions to a common target distribution. For example, image restoration might require handling different types of degradation without knowing the degradation of each input at inference time. Simple approaches of pooling the source distributions only encourage alignment with the target at the aggregate level and may leave individual sources misaligned. In our paper, we consider the simultaneous OT problem which formalizes the task of learning a shared transport map that minimizes the average transport cost while aligning each source distribution with a prescribed target. We propose a neural method for solving the simultaneous OT problem by learning a shared transport map that minimizes the average transport cost while aligning each source distribution with a prescribed target. We derive a max-min formulation for learning this map. We illustrate its application to image restoration, where a single model handles multiple degradation types using a common collection of clean target images.
Figures & tables
Figure 1: Restoration of selected CelebA 64×64 test images under bilinear downsampling with an unseen factor ( ×3 at test time versus ×4 during training). Each group shows the degraded input, conditional OT with classifier-based routing, SimNOT (ours), and the clean reference.
Figure 2: A schematic illustration of simultaneous OT.
Figure 3: Simultaneous transport from five Gaussian sources to a common Swiss roll target. Each panel shows the outputs of the same learned map for one source (colored points), overlaid with target samples (gray points). The visualization uses independent samples after 100K training iterations.
Method
Mean FID ↓
Max FID ↓
PSNR ↑
Degraded input
87.60
178.94
25.19
Pooled UOT
9.96
14.39
27.51
Conditional UOT + classifier
6.39
8.33
28.21
SimNOT (ours)
8.23
11.88
27.85
Table 1: CelebA restoration on held-out images with the training degradation parameters. Each of the five degradations is applied to the same full set of test images. Mean and Max FID aggregate the five source-specific FIDs. PSNR (dB) is averaged over images and degradation conditions. The best and second-best results among restoration methods are bold and underlined, respectively.
Bilinear ×3
Bilinear ×5
Method
FID ↓
LPIPS ↓
PSNR ↑
FID ↓
LPIPS ↓
PSNR ↑
Degraded input
92.82
0.1821
25.02
317.44
0.3717
20.93
Pooled UOT
27.71
0.0717
24.23
99.22
0.2342
20.51
Conditional UOT + classifier
35.16
0.1185
24.35
80.87
0.2670
20.80
SimNOT (ours)
18.44
0.0635
25.03
32.15
0.1134
20.93
Conditional UOT + known family
21.42
0.0637
23.99
77.23
0.1983
20.69
Table 2: Generalization to unseen bilinear downsampling factors. Models trained with factor ×4 are evaluated at factors ×3 and ×5 without retraining. Each setting uses all 20,259 test images. The last row supplies the bilinear family label directly to Conditional UOT. Bold indicates the best result among restoration methods.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Bicubic
Bilinear
JPEG
Blur
Noise
FID ↓
Degraded input
124.86
178.94
31.19
46.47
56.54
Pooled UOT
10.96
8.54
9.15
6.75
14.39
Conditional UOT + classifier
8.33
8.18
5.92
3.48
6.02
SimNOT (ours)
11.28
11.88
6.59
6.27
5.14
PSNR (dB) ↑
Appendix
Table 3: Source-specific results on the original degradation settings. The evaluation protocol is described in Appendix D.2 . The best and second-best results among restoration methods are bold and underlined, respectively. Rankings use unrounded values.
Method
FID ↓
LPIPS ↓
PSNR ↑
Bicubic ×3 (training: ×4 )
Degraded input
61.46
0.16835
25.24
Pooled UOT
21.06
0.07670
23.10
Conditional UOT + classifier
113.85
0.20161
20.25
SimNOT (ours)
18.72
0.06642
22.49
Conditional UOT + known family
9.27
0.04598
23.74
Appendix
Table 4: Generalization to milder degradations. All models are evaluated without retraining on 20,259 images per setting. The best and second-best results among blind restoration methods are bold and underlined, respectively. The known-family variant receives the true degradation family and is reported separately as a reference. Rankings use unrounded values.
Method
FID ↓
LPIPS ↓
PSNR ↑
Bicubic ×5 (training: ×4 )
Degraded input
238.78
0.41126
20.67
Pooled UOT
40.31
0.16003
20.09
Conditional UOT + classifier
177.46
0.27675
19.69
SimNOT (ours)
151.24
0.24771
19.49
Conditional UOT + known family
21.89
0.09326
20.37
Appendix
Table 5: Generalization to stronger degradations. The evaluation protocol follows Table 4 .
Degradation
Test parameter
Accuracy (%) ↑
Bicubic
×3
0.035
Bicubic
×5
0.000
Bilinear
×3
0.000
Bilinear
×5
0.000
JPEG
30
100.000
JPEG
20
100.000
Appendix
Table 6: Degradation classification accuracy under parameter shifts. Each setting contains 20,259 test images. The classifier is trained only on the fixed degradation parameters specified in Appendix D.2 .
Department of Mathematical Sciences, Seoul National University · IPAI (Interdisciplinary Program in Artificial Intelligence, Seoul National University) · Sungkyunkwan University