cs.LGSep 30, 2026

Does a Shared Temperature Imply a Shared Angular Scale in Probabilistic Contrastive Learning?

Authors: Ningkang Peng, Qianfeng Yu, Jingyang Mao, Xiaoqian Peng, Tingyu Lu, Peirong Ma, Yanhui Gu

Organizations: Nanjing Normal University · Nanjing University of Chinese Medicine · Tohoku University

Abstract

In probabilistic contrastive learning, a shared temperature is commonly interpreted as a shared similarity scale, but this interpretation does not hold for high-dimensional distributional class representations. We study the exact von Mises-Fisher (vMF) probabilistic score used by ProCo when representation dimension and class concentration grow jointly. We prove that the score retains a class-dependent leading angular gain gc=Ac/τg_c=A_c/τ, where AcA_c is the mean resultant length. This gain enters Softmax competition, pairwise decision boundaries, and feature gradients. On real CIFAR-LT, ImageNet-LT, and iNaturalist representations, the theory accurately predicts boundary movements and local gradient changes under the full vMF score. Classwise temperature adjustment also changes the cosine-zero intercept and finite-dimensional response. We construct intercept-preserving and Pure Angular controls to separate the leading gain from these accompanying changes. Complete gain equalization yields a shared-scale cosine prototype rule at leading order; a finite-dimensional margin condition guarantees agreement of the two classifiers. Across 16 frozen representation settings, prediction agreement is 98.43-99.99%, with disagreements concentrated at small cosine margins. In controlled contrastive-only training with the training-frequency prior, Pure Angular editing improves both learned representations at all tested CIFAR-10/100 imbalance factors and retains positive changes on ImageNet-LT. Thus vMF concentration not only describes class distributions, but also forms a decision and learning scale in high-dimensional probabilistic contrastive learning.

Figures & tables

Appendix figures & tables30 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

May 11, 2026cs.LG

Optimal Representations for Generalized Contrastive Learning with Imbalanced Datasets

In this paper, we provide a computable characterization of the geometry of optimal representations in Contrastive Learning (CL) when the classes are imbalanced. When classes are balanced and the representation dimension is greater than the number of classes, it is well-known that the optimal representations exhibit Neural Collapse (NC), i.e., representations from the same class collapse to their class means and the class means form an Equiangular Tight Frame (ETF). For imbalanced classes and a large, generalized family of CL losses, we prove that the optimal representations of all samples from the same class collapse to their class means and their geometry exhibits an angular symmetry structure that is determined by the relative class proportions. In general, we show that the geometry can be determined by solving a convex optimization problem. Exploiting this symmetry structure, we analytically investigate a special case where class imbalance is extreme and prove that CL exhibits a phenomenon called Minority Collapse (MC) where all samples from the minority classes (classes with small probabilities) collapse into a single vector, whenever the class imbalance exceeds a threshold, which in turn depends on the regularity properties of the CL loss used and on the number of negative samples. Numerical results are provided to illustrate these phenomena and corroborate the theoretical results. We conclude by identifying a number of open problems.
Sep 30, 2026cs.DS

Optimal VC Dimension of Contrastive Learning with Margin

Contrastive learning is a successful paradigm for learning dd-dimensional geometric representations from a collection of anchor--positive--negative'' triplets $(i,j^{+},k^{-})$, indicating that item ii is closer to jj than to kk.'' Despite its success, understanding why contrastive learning leads to representations of high \textit{generalization} quality---beyond the often pessimistic predictions from PAC-learning---remains a central question. Recently, \citet*{alon2024optimal} proved that, for PAC-learning dd-dimensional Euclidean representations of nn-point datasets, Θ(min⁡(nd,n2))Θ(\min(nd, n^2)) triplets are necessary and sufficient, while they posed as an open question whether their VC dimension bounds for the more realistic setting of \textit{contrastive learning with a margin} can be improved. For a margin parameter α>0α>0, a triplet (i,j+,k−)α(i,j^{+},k^{-})_α is satisfied by the embedding φ:[n]→Rdφ:[n]\rightarrow \mathbb{R}^{d}, if ∥φ(i)−φ(k)∥2>(1+α)⋅∥φ(i)−φ(j)∥2\|φ(i)-φ(k)\|_2>(1+α)\cdot\|φ(i)-φ(j)\|_2. In this work, we resolve their question by proving that the VC dimension of contrastive learning under any margin α∈(0,1)α\in(0,1) is in fact O(n/α2)O(n/α^2), improving on the previous bound of O(nlog⁡(n)/α2)O(n\log(n)/α^2). We also establish that the bounds are optimal up to constant factors, by providing a matching lower bound of Ω(nα2)Ω(\frac{n}{α^2}) (the previously known lower bound was Ω(nα)Ω(\frac{n}α)), for α≥max⁡(n−1/2,d−1/2)α\geq \max(n^{-1/2},d^{-1/2}).
Sep 7, 2026cs.CV

TeMo: Temperature Modulation for Multimodal Contrastive Learning

Contrastive learning approaches achieve strong performance by training models to bring similar samples closer while pushing dissimilar samples apart. A crucial component of contrastive learning is the temperature hyperparameter ττ, which controls the penalty strength applied to negative samples. However, most existing methods either fix this hyperparameter or learn a global value during training. In this paper, we introduce TeMo, Temperature Modulation framework, a similarity-based modulation approach that adaptively adjusts the temperature for each positive-negative pair according to their similarity, enabling more fine-grained multimodal contrastive learning. Our approach seamlessly integrates temperature-modulated multimodal and unimodal losses with the standard multimodal contrastive loss by gradually transitioning between them. This design allows the model to capture both coarse- and fine-grained semantics at different training stages. Extensive experiments demonstrate that each component of TeMo consistently enhances performance across diverse zero-shot retrieval and classification tasks, establishing new state-of-the-art results.