stat.MLSep 27, 2026

Multi-Marginal Inverse Optimal Transport for Contrastive Learning Via Explicit Anchor-Positive-Negative Coupling

Authors: Ngoc-Hai Nguyen, Thuan Nguyen, Prakash Ishwar, Shuchin Aeron

Organizations: Department of Electrical and Computer Engineering Tufts University · Department of Engineering, Engineering Technology East Tennessee State University · Department of Electrical and Computer Engineering Boston University

Abstract

Inverse Optimal Transport (OT) based methods for representation learning learn representations such that the global OT coupling between a pair of data marginals in the representation space, concentrates on the positive pairs. This is in contrast to previous methods that primarily focused on pairwise matching. However, these methods DO NOT\textit{DO NOT} utilize negative pairs and hence are not truly contrastive in their approach. We show that this leads to issues of dimensional collapse and hence degraded downstream performance. To alleviate this, we develop a novel multi-marginal (MM) inverse OT (IOT) contrastive learning (CL) approach called Neg-MMIOT-CL, which learns representations such that the global multi-marginal OT (MMOT) coupling between a triple of data marginals, with respect to a carefully designed ground-cost between triplets of data points in the representation space, concentrates on the anchor-positive-negative triplets\textit{triplets}. For a latent class model, we empirically show that Neg-MMIOT-CL alleviates dimensional collapse. Furthermore, for a specific choice of ground cost for all triplets in representation space, we prove that the optimal representation configuration for Neg-MMIOT-CL exhibits equiangular property for within-class and across-class representations, which translates to Neural-Collapse when the representation dimension is larger than the number of classes minus one -- a result that is previously established only\textit{previously established only} for pairwise contrastive learning methods. Finally, we propose Neg-IOT-CL-PushPull, that is a computationally efficient alternative to Neg-MMIOT-CL, alleviating the high cost of computing MMOT plans needed during implementation. We apply these methods on both synthetic and real-world datasets and show significant improvements over existing OT-based contrastive learning methods.

Figures & tables

Appendix figures & tables21 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

May 11, 2026cs.LG

Optimal Representations for Generalized Contrastive Learning with Imbalanced Datasets

In this paper, we provide a computable characterization of the geometry of optimal representations in Contrastive Learning (CL) when the classes are imbalanced. When classes are balanced and the representation dimension is greater than the number of classes, it is well-known that the optimal representations exhibit Neural Collapse (NC), i.e., representations from the same class collapse to their class means and the class means form an Equiangular Tight Frame (ETF). For imbalanced classes and a large, generalized family of CL losses, we prove that the optimal representations of all samples from the same class collapse to their class means and their geometry exhibits an angular symmetry structure that is determined by the relative class proportions. In general, we show that the geometry can be determined by solving a convex optimization problem. Exploiting this symmetry structure, we analytically investigate a special case where class imbalance is extreme and prove that CL exhibits a phenomenon called Minority Collapse (MC) where all samples from the minority classes (classes with small probabilities) collapse into a single vector, whenever the class imbalance exceeds a threshold, which in turn depends on the regularity properties of the CL loss used and on the number of negative samples. Numerical results are provided to illustrate these phenomena and corroborate the theoretical results. We conclude by identifying a number of open problems.
Sep 21, 2026cs.CV

Positive Pair Geometry Matters: Optimal Transport for Contrastive Learning of Visual Representations

Contrastive self-supervised learning has achieved strong performance by learning representations from multiple augmented views of the same image. However, most existing methods construct positive pairs using independently sampled stochastic augmentations, which may alter semantic content and ignore the intrinsic geometry of the data distribution. In this work, we propose OTCLR, an optimal transport-aware framework for contrastive learning representations that generates geometry-consistent positive samples. Instead of directly contrasting two randomly augmented views, we construct intermediate views between the original image and its augmented variants through entropic optimal-transport displacement interpolation. These transport-interpolated samples serve as positive views that better preserve image structure while explicitly modeling spatial distributional geometry. To further promote smooth representation learning, we evaluate auxiliary Sinkhorn regularization terms that encourage transport-interpolated views to remain consistent with their endpoint images. The proposed method can be incorporated into standard contrastive learning pipelines without modifying the encoder architecture. Experiments on multiple benchmark datasets show that our approach improves representation quality and transfer learning performance compared with conventional augmentation-based contrastive learning baselines.
Sep 30, 2026cs.DS

Optimal VC Dimension of Contrastive Learning with Margin

Contrastive learning is a successful paradigm for learning dd-dimensional geometric representations from a collection of anchor--positive--negative'' triplets $(i,j^{+},k^{-})$, indicating that item ii is closer to jj than to kk.'' Despite its success, understanding why contrastive learning leads to representations of high \textit{generalization} quality---beyond the often pessimistic predictions from PAC-learning---remains a central question. Recently, \citet*{alon2024optimal} proved that, for PAC-learning dd-dimensional Euclidean representations of nn-point datasets, Θ(min⁡(nd,n2))Θ(\min(nd, n^2)) triplets are necessary and sufficient, while they posed as an open question whether their VC dimension bounds for the more realistic setting of \textit{contrastive learning with a margin} can be improved. For a margin parameter α>0α>0, a triplet (i,j+,k−)α(i,j^{+},k^{-})_α is satisfied by the embedding φ:[n]→Rdφ:[n]\rightarrow \mathbb{R}^{d}, if ∥φ(i)−φ(k)∥2>(1+α)⋅∥φ(i)−φ(j)∥2\|φ(i)-φ(k)\|_2>(1+α)\cdot\|φ(i)-φ(j)\|_2. In this work, we resolve their question by proving that the VC dimension of contrastive learning under any margin α∈(0,1)α\in(0,1) is in fact O(n/α2)O(n/α^2), improving on the previous bound of O(nlog⁡(n)/α2)O(n\log(n)/α^2). We also establish that the bounds are optimal up to constant factors, by providing a matching lower bound of Ω(nα2)Ω(\frac{n}{α^2}) (the previously known lower bound was Ω(nα)Ω(\frac{n}α)), for α≥max⁡(n−1/2,d−1/2)α\geq \max(n^{-1/2},d^{-1/2}).