Does a Shared Temperature Imply a Shared Angular Scale in Probabilistic Contrastive Learning?
Organizations: Nanjing Normal University · Nanjing University of Chinese Medicine · Tohoku University
Abstract
In probabilistic contrastive learning, a shared temperature is commonly interpreted as a shared similarity scale, but this interpretation does not hold for high-dimensional distributional class representations. We study the exact von Mises-Fisher (vMF) probabilistic score used by ProCo when representation dimension and class concentration grow jointly. We prove that the score retains a class-dependent leading angular gain , where is the mean resultant length. This gain enters Softmax competition, pairwise decision boundaries, and feature gradients. On real CIFAR-LT, ImageNet-LT, and iNaturalist representations, the theory accurately predicts boundary movements and local gradient changes under the full vMF score. Classwise temperature adjustment also changes the cosine-zero intercept and finite-dimensional response. We construct intercept-preserving and Pure Angular controls to separate the leading gain from these accompanying changes. Complete gain equalization yields a shared-scale cosine prototype rule at leading order; a finite-dimensional margin condition guarantees agreement of the two classifiers. Across 16 frozen representation settings, prediction agreement is 98.43-99.99%, with disagreements concentrated at small cosine margins. In controlled contrastive-only training with the training-frequency prior, Pure Angular editing improves both learned representations at all tested CIFAR-10/100 imbalance factors and retains positive changes on ImageNet-LT. Thus vMF concentration not only describes class distributions, but also forms a decision and learning scale in high-dimensional probabilistic contrastive learning.
Figures & tables
| Dataset / IF | Feature | ProCo | Direct temperature | Intercept- preserving | Pure Angular | |
|---|---|---|---|---|---|---|
| CIFAR-10 / 10 | Encoder | |||||
| Projection | ||||||
| CIFAR-10 / 50 | Encoder | |||||
| Projection | ||||||
| CIFAR-10 / 100 | Encoder | |||||
| Projection |
| Dataset / IF | Feature | Original vMF | Cosine | Pure Angular | Agreement |
|---|---|---|---|---|---|
| CIFAR-10 / 10 | Encoder | 85.21 | 85.82 | 85.90 | 99.74 |
| Projection | 87.85 | 87.84 | 87.83 | 99.99 | |
| CIFAR-10 / 50 | Encoder | 81.47 | 80.76 | 80.72 | 99.46 |
| Projection | 80.72 | 80.58 | 80.58 | 99.95 | |
| CIFAR-10 / 100 | Encoder | 78.76 | 78.19 | 78.02 | 99.51 |
| Projection | 77.81 | 77.59 | 77.58 | 99.93 |
Appendix figures & tables30 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset / IF | Feature | Original | ||||
|---|---|---|---|---|---|---|
| CIFAR-10 / 10 | Encoder | 85.21 | 85.92 | 85.84 | 85.90 | |
| Projection | 87.85 | 87.83 | 87.83 | 87.83 | ||
| CIFAR-10 / 50 | Encoder | 81.47 | 80.71 | 80.72 | 80.72 | |
| Projection | 80.72 | 80.58 | 80.58 | 80.58 | ||
| CIFAR-10 / 100 | Encoder | 78.76 | 78.00 | 78.12 | 78.02 | |
| Projection | 77.81 | 77.58 | 77.59 | 77.58 |
| Dataset / IF | Feature | Disagree | Low-margin | Certified / Total |
|---|---|---|---|---|
| CIFAR-10 / 10 | Encoder | 26 | 26 | 9,763 / 10,000 |
| Projection | 1 | 1 | 9,982 / 10,000 | |
| CIFAR-10 / 50 | Encoder | 54 | 54 | 9,717 / 10,000 |
| Projection | 5 | 5 | 9,969 / 10,000 | |
| CIFAR-10 / 100 | Encoder | 49 | 49 | 9,655 / 10,000 |
| Projection | 7 | 7 | 9,960 / 10,000 |
| Dataset / IF | Feature | Mean | Grad. error (%) | Median grad. cosine |
|---|---|---|---|---|
| CIFAR-100 / 10 | Encoder | 0.01141 | 1.555 | 0.99999712 |
| Projection | 0.01751 | 1.624 | 0.99996033 | |
| CIFAR-100 / 50 | Encoder | 0.00917 | 1.524 | 0.99999679 |
| Projection | 0.02206 | 1.592 | 0.99996161 | |
| CIFAR-100 / 100 | Encoder | 0.00832 | 1.448 | 0.99999724 |
| Projection | 0.02634 | 1.595 | 0.99995594 |
| ResNet-50 | ResNeXt-50 | ||
| 0 | -1 | -1.665 | -1.465 |
| 0 | 0 | +0.195 | +0.315 |
| 0 | 1 | +0.175 | -0.175 |
| 0.5 | -1 | -0.635 | -0.610 |
| 0.5 | 0 | +0.260 | +0.240 |
| 0.5 | 1 | +0.055 | -0.040 |
| CIFAR-10-LT | CIFAR-100-LT | |||
|---|---|---|---|---|
| Policy | IF50 | IF100 | IF50 | IF100 |
| ProCo | 87.57 | 85.39 | 55.78 | 51.34 |
| Reduced gain disparity | 87.72 | 85.10 | 56.62 | 51.72 |
| Increased gain disparity | 87.56 | 85.14 | 55.99 | 52.05 |
| Contrastive-only | Joint | |||||
|---|---|---|---|---|---|---|
| Realization | E | P | E | P | FC | |
| ProCo | 1 | 92.15 | 92.91 | 92.95 | 93.04 | 92.81 |
| Temperature | 0 | 92.09 | 93.13 | 92.77 | 92.60 | 92.54 |
| Temperature | 0.5 | 92.32 | 92.95 | 92.40 | 92.32 | 92.38 |
| Intercept-preserving | 0 | 92.46 | 93.01 | 93.01 | 93.04 | 92.85 |
| Intercept-preserving | 0.5 | 92.59 | 93.19 | 92.84 | 92.89 | 92.55 |
| Contrastive-only | Joint | |||||
|---|---|---|---|---|---|---|
| Realization | E | P | E | P | FC | |
| ProCo | 1 | 88.97 | 88.63 | 88.57 | 88.28 | 86.32 |
| Temperature | 0 | 88.34 | 87.58 | 88.66 | 88.01 | 85.96 |
| Temperature | 0.5 | 88.29 | 88.21 | 88.20 | 87.84 | 85.96 |
| Intercept-preserving | 0 | 88.61 | 88.20 | 88.36 | 87.96 | 85.62 |
| Intercept-preserving | 0.5 | 88.51 | 88.24 | 88.34 | 87.96 | 85.83 |
| Contrastive-only | Joint | |||||
|---|---|---|---|---|---|---|
| Realization | E | P | E | P | FC | |
| ProCo | 1 | 88.60 | 88.41 | 88.44 | 88.19 | 87.46 |
| Temperature | 0 | 88.60 | 88.94 | 88.23 | 88.00 | 87.29 |
| Temperature | 0.5 | 88.55 | 88.34 | 88.31 | 88.10 | 87.34 |
| Intercept-preserving | 0 | 88.74 | 89.02 | 88.23 | 87.96 | 87.41 |
| Intercept-preserving | 0.5 | 88.35 | 88.53 | 88.44 | 87.76 | 87.27 |
| Contrastive-only | Joint | |||||
|---|---|---|---|---|---|---|
| Realization | E | P | E | P | FC | |
| ProCo | 1 | 82.70 | 81.52 | 83.70 | 82.29 | 76.14 |
| Temperature | 0 | 76.82 | 68.89 | 83.63 | 81.89 | 75.81 |
| Temperature | 0.5 | 80.98 | 78.34 | 83.97 | 82.07 | 76.04 |
| Intercept-preserving | 0 | 78.18 | 72.34 | 83.78 | 81.73 | 75.23 |
| Intercept-preserving | 0.5 | 81.54 | 79.21 | 83.36 | 81.88 | 76.28 |
| Contrastive-only | Joint | |||||
|---|---|---|---|---|---|---|
| Realization | E | P | E | P | FC | |
| ProCo | 1 | 84.86 | 83.91 | 83.42 | 82.04 | 80.52 |
| Temperature | 0 | 84.57 | 83.61 | 83.36 | 82.95 | 80.46 |
| Temperature | 0.5 | 85.16 | 84.31 | 84.25 | 83.37 | 81.24 |
| Intercept-preserving | 0 | 84.60 | 84.19 | 84.13 | 83.07 | 80.81 |
| Intercept-preserving | 0.5 | 84.38 | 84.04 | 84.60 | 82.93 | 80.79 |
| Contrastive-only | Joint | |||||
|---|---|---|---|---|---|---|
| Realization | E | P | E | P | FC | |
| ProCo | 1 | 78.28 | 74.65 | 81.64 | 79.55 | 71.25 |
| Temperature | 0 | 70.46 | 57.88 | 81.10 | 76.62 | 71.58 |
| Temperature | 0.5 | 74.56 | 67.21 | 81.93 | 79.16 | 71.28 |
| Intercept-preserving | 0 | 70.65 | 58.17 | 81.44 | 79.14 | 71.44 |
| Intercept-preserving | 0.5 | 74.82 | 67.63 | 81.42 | 78.43 | 71.31 |
| Contrastive-only | Joint | |||||
|---|---|---|---|---|---|---|
| Realization | E | P | E | P | FC | |
| ProCo | 1 | 82.49 | 81.32 | 81.50 | 80.13 | 77.79 |
| Temperature | 0 | 82.43 | 80.56 | 81.56 | 79.72 | 77.39 |
| Temperature | 0.5 | 82.15 | 80.41 | 81.29 | 79.52 | 77.11 |
| Intercept-preserving | 0 | 82.64 | 81.49 | 80.90 | 79.31 | 76.71 |
| Intercept-preserving | 0.5 | 82.71 | 81.53 | 81.12 | 78.91 | 76.36 |
| Dataset / IF | ProCo | Direct Temp. | Preserve | Pure Angular | |
|---|---|---|---|---|---|
| CIFAR-10 / 10 | 87.46 | 87.29 | 87.41 | 87.24 | -0.22 |
| CIFAR-10 / 50 | 80.52 | 80.46 | 80.81 | 80.79 | +0.27 |
| CIFAR-10 / 100 | 77.79 | 77.39 | 76.71 | 76.50 | -1.29 |
| CIFAR-100 / 10 | 58.26 | 58.24 | 57.55 | 58.75 | +0.49 |
| CIFAR-100 / 50 | 47.63 | 46.73 | 48.82 | 48.10 | +0.47 |
| CIFAR-100 / 100 | 44.67 | 43.15 | 43.54 | 43.88 | -0.79 |
| Setting | CIFAR-10/100-LT | ImageNet-LT |
|---|---|---|
| Encoder | ResNet-32 | ResNet-50 |
| Encoder dimension | 64 | 2048 |
| Projection dimension | 128 | 1024 |
| Training epochs | 200 | 90 |
| Global batch size | 256 | 256 |
| Optimizer | SGD | SGD |
| Setting | Value or protocol |
|---|---|
| Linear-classifier epochs | 100; fixed final-epoch representation |
| Linear-classifier optimizer | SGD; batch size 256; learning rate 0.1; momentum 0.9; zero weight decay; recorded cosine schedule |
| Linear-classifier objective | Ordinary cross-entropy; no extra prior offset; no vMF correction |
| Main linear fitting pool | Corresponding long-tail training subset; ImageNet uses 115,846 fitting images |
| Main Encoder preprocessing | CIFAR: raw Encoder features; ImageNet: unit-normalized Encoder features in the main comparison |
| Projection preprocessing | Original unit-normalized Projection features |
Explore similar work
Optimal VC Dimension of Contrastive Learning with Margin
anchor--positive--negative'' triplets $(i,j^{+},k^{-})$, indicating that item is closer to than to .'' Despite its success, understanding why contrastive learning leads to representations of high \textit{generalization} quality---beyond the often pessimistic predictions from PAC-learning---remains a central question. Recently, \citet*{alon2024optimal} proved that, for PAC-learning -dimensional Euclidean representations of -point datasets, triplets are necessary and sufficient, while they posed as an open question whether their VC dimension bounds for the more realistic setting of \textit{contrastive learning with a margin} can be improved. For a margin parameter , a triplet is satisfied by the embedding , if . In this work, we resolve their question by proving that the VC dimension of contrastive learning under any margin is in fact , improving on the previous bound of . We also establish that the bounds are optimal up to constant factors, by providing a matching lower bound of (the previously known lower bound was ), for .