Optimal Transport (OT) provides a principled framework for learning transformations between probability distributions from unpaired samples. In many applications, however, a single transformation must map several source distributions to a common target distribution. For example, image restoration might require handling different types of degradation without knowing the degradation of each input at inference time. Simple approaches of pooling the source distributions only encourage alignment with the target at the aggregate level and may leave individual sources misaligned. In our paper, we consider the simultaneous OT problem which formalizes the task of learning a shared transport map that minimizes the average transport cost while aligning each source distribution with a prescribed target. We propose a neural method for solving the simultaneous OT problem by learning a shared transport map that minimizes the average transport cost while aligning each source distribution with a prescribed target. We derive a max-min formulation for learning this map. We illustrate its application to image restoration, where a single model handles multiple degradation types using a common collection of clean target images.
Figures & tables
Figure 1: Restoration of selected CelebA 64×64 test images under bilinear downsampling with an unseen factor ( ×3 at test time versus ×4 during training). Each group shows the degraded input, conditional OT with classifier-based routing, SimNOT (ours), and the clean reference.
Figure 2: A schematic illustration of simultaneous OT.
Figure 3: Simultaneous transport from five Gaussian sources to a common Swiss roll target. Each panel shows the outputs of the same learned map for one source (colored points), overlaid with target samples (gray points). The visualization uses independent samples after 100K training iterations.
Method
Mean FID ↓
Max FID ↓
PSNR ↑
Degraded input
87.60
178.94
25.19
Pooled UOT
9.96
14.39
27.51
Conditional UOT + classifier
6.39
8.33
28.21
SimNOT (ours)
8.23
11.88
27.85
Table 1: CelebA restoration on held-out images with the training degradation parameters. Each of the five degradations is applied to the same full set of test images. Mean and Max FID aggregate the five source-specific FIDs. PSNR (dB) is averaged over images and degradation conditions. The best and second-best results among restoration methods are bold and underlined, respectively.
Bilinear ×3
Bilinear ×5
Method
FID ↓
LPIPS ↓
PSNR ↑
FID ↓
LPIPS ↓
PSNR ↑
Degraded input
92.82
0.1821
25.02
317.44
0.3717
20.93
Pooled UOT
27.71
0.0717
24.23
99.22
0.2342
20.51
Conditional UOT + classifier
35.16
0.1185
24.35
80.87
0.2670
20.80
SimNOT (ours)
18.44
0.0635
25.03
32.15
0.1134
20.93
Conditional UOT + known family
21.42
0.0637
23.99
77.23
0.1983
20.69
Table 2: Generalization to unseen bilinear downsampling factors. Models trained with factor ×4 are evaluated at factors ×3 and ×5 without retraining. Each setting uses all 20,259 test images. The last row supplies the bilinear family label directly to Conditional UOT. Bold indicates the best result among restoration methods.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Bicubic
Bilinear
JPEG
Blur
Noise
FID ↓
Degraded input
124.86
178.94
31.19
46.47
56.54
Pooled UOT
10.96
8.54
9.15
6.75
14.39
Conditional UOT + classifier
8.33
8.18
5.92
3.48
6.02
SimNOT (ours)
11.28
11.88
6.59
6.27
5.14
PSNR (dB) ↑
Appendix
Table 3: Source-specific results on the original degradation settings. The evaluation protocol is described in Appendix D.2 . The best and second-best results among restoration methods are bold and underlined, respectively. Rankings use unrounded values.
Method
FID ↓
LPIPS ↓
PSNR ↑
Bicubic ×3 (training: ×4 )
Degraded input
61.46
0.16835
25.24
Pooled UOT
21.06
0.07670
23.10
Conditional UOT + classifier
113.85
0.20161
20.25
SimNOT (ours)
18.72
0.06642
22.49
Conditional UOT + known family
9.27
0.04598
23.74
Appendix
Table 4: Generalization to milder degradations. All models are evaluated without retraining on 20,259 images per setting. The best and second-best results among blind restoration methods are bold and underlined, respectively. The known-family variant receives the true degradation family and is reported separately as a reference. Rankings use unrounded values.
Method
FID ↓
LPIPS ↓
PSNR ↑
Bicubic ×5 (training: ×4 )
Degraded input
238.78
0.41126
20.67
Pooled UOT
40.31
0.16003
20.09
Conditional UOT + classifier
177.46
0.27675
19.69
SimNOT (ours)
151.24
0.24771
19.49
Conditional UOT + known family
21.89
0.09326
20.37
Appendix
Table 5: Generalization to stronger degradations. The evaluation protocol follows Table 4 .
Degradation
Test parameter
Accuracy (%) ↑
Bicubic
×3
0.035
Bicubic
×5
0.000
Bilinear
×3
0.000
Bilinear
×5
0.000
JPEG
30
100.000
JPEG
20
100.000
Appendix
Table 6: Degradation classification accuracy under parameter shifts. Each setting contains 20,259 test images. The classifier is trained only on the fixed degradation parameters specified in Appendix D.2 .
We propose an implicit neural formulation of optimal transport that eliminates adversarial min--max optimization and multi-network architectures commonly used in existing approaches. Our key idea is to parameterize a single potential in the Kantorovich dual and reformulate the associated c-transform as a proximal fixed-point problem. This yields a stable single-network framework in which dual feasibility is enforced exactly through proximal optimality conditions rather than adversarial training. Despite the inner fixed-point computation, gradients can be computed without differentiating through the fixed-point iterations, enabling efficient training without requiring implicit differentiation. We further establish convergence of stochastic gradient descent. The resulting framework is efficient, scalable, and broadly applicable: it simultaneously recovers forward and backward transport maps and naturally extends to class-conditional settings. Experiments on high-dimensional Gaussian benchmarks, physical datasets, and image translation tasks demonstrate strong transport accuracy together with improved training stability and favorable computational and memory efficiency.
Yesom Park, Eric Gelphman, Stanley Osher +1
Department of Mathematics, University of California, Los Angeles · Department of Applied Mathematics and Statistics, Colorado School of Mines
We investigate unpaired image inverse problems, a challenging setting where only independent, non-paired sets of noisy measurements and clean target signals are available for training. We propose a novel inverse problem solver based on Unbalanced Optimal Transport, called Unbalanced Optimal Transport Map for Inverse Problems (UOTIP). Our method formulates the reconstruction task, predicting clean target signals from noisy measurements, as learning a UOT Map from noisy measurement distribution to clean signal distribution by incorporating a likelihood-based cost function. By relaxing the exact marginal constraint, the UOT framework provides key advantages to our model: robustness to multi-level observation noise, adaptability to class imbalance between noisy and clean datasets, and generalizability to diverse noise-type scenarios. Furthermore, we theoretically demonstrate that incorporating a quadratic cost term ensures the existence and uniqueness of the transport map by satisfying the twist condition, even for ill-posed inverse problems. Our experiments demonstrate that UOTIP achieves state-of-the-art performance on unpaired image inverse problem benchmarks, across linear and nonlinear inverse problems.
Donggyu Lee, Taekyung Lee, Jaewoong Choi
Department of Mathematical Sciences, Seoul National University · IPAI (Interdisciplinary Program in Artificial Intelligence, Seoul National University) · Sungkyunkwan University
Alignment plays a fundamental role in many machine learning problems, such as multi-network analysis, multimodal learning, and point cloud registration. Recent works increasingly leverage optimal transport (OT) for distributional alignment, whose effectiveness largely depends on sparse supervision that is hard or costly to obtain in practice. Existing works, however, largely overlook how to actively acquire high-quality supervision to improve their alignment performance under OT frameworks. In this paper, we propose a principled active alignment framework for optimal transport alignment called AvAtar. We quantify the informativeness of a candidate by measuring its gradient-based impact on the global alignment result, computed as the gradient propagation from the global alignment result to all possible supervisions of the candidate through the entropy-regularized OT formulation. While differentiating through OT is challenging given its constrained nature, we leverage the adjoint-state method to reformulate the computation to a linear system solvable by the conjugate gradient method with linear complexity and guaranteed convergence. By encoding the global alignment result via effective utility functions, AvAtar is applicable to general alignment problems under the OT framework. Extensive experiments on three representative alignment tasks demonstrate the effectiveness, scalability, and generalizability of the proposed AvAtar.
Qi Yu, Ruizhong Qiu, Zhichen Zeng +3
University of Illinois Urbana-Champaign · University of Florida · 3Arizona State University