In probabilistic contrastive learning, a shared temperature is commonly interpreted as a shared similarity scale, but this interpretation does not hold for high-dimensional distributional class representations. We study the exact von Mises-Fisher (vMF) probabilistic score used by ProCo when representation dimension and class concentration grow jointly. We prove that the score retains a class-dependent leading angular gain gc=Ac/τ, where Ac is the mean resultant length. This gain enters Softmax competition, pairwise decision boundaries, and feature gradients. On real CIFAR-LT, ImageNet-LT, and iNaturalist representations, the theory accurately predicts boundary movements and local gradient changes under the full vMF score. Classwise temperature adjustment also changes the cosine-zero intercept and finite-dimensional response. We construct intercept-preserving and Pure Angular controls to separate the leading gain from these accompanying changes. Complete gain equalization yields a shared-scale cosine prototype rule at leading order; a finite-dimensional margin condition guarantees agreement of the two classifiers. Across 16 frozen representation settings, prediction agreement is 98.43-99.99%, with disagreements concentrated at small cosine margins. In controlled contrastive-only training with the training-frequency prior, Pure Angular editing improves both learned representations at all tested CIFAR-10/100 imbalance factors and retains positive changes on ImageNet-LT. Thus vMF concentration not only describes class distributions, but also forms a decision and learning scale in high-dimensional probabilistic contrastive learning.
Figures & tables
Dataset / IF
Feature
ProCo
Direct temperature β=0
Intercept- preserving β=0
Pure Angular β=0
Δ
CIFAR-10 / 10
Encoder
86.28±0.09
86.52±0.15
86.54±0.11
86.93±0.10
+0.65
Projection
87.10±0.06
87.55±0.08
87.65±0.16
87.96±0.13
+0.86
CIFAR-10 / 50
Encoder
75.73±0.12
75.93±0.19
76.02±0.10
76.14±0.17
+0.41
Projection
77.66±0.07
77.64±0.11
78.46±0.14
78.83±0.09
+1.17
CIFAR-10 / 100
Encoder
68.24±0.16
67.89±0.13
68.13±0.13
68.43±0.11
+0.19
Projection
72.65±0.10
72.48±0.14
72.74±0.08
73.12±0.16
+0.47
Table 1: Representation accuracy after contrastive-only training. Long-tail linear-probe accuracy (%) after feature learning with the training-frequency prior. All interventions use β=0 . Results report mean ± sample standard deviation over three seeds. Δ is the difference between the displayed Pure Angular and ProCo means in percentage points; darker green marks a larger increase. Readout protocols are given in Section 3.4 and Appendices I.2 and N .
Figure 1: Class-dependent angular gain at a shared temperature. (a) A training-selected CIFAR-100-LT pair in a boundary-sign-preserving spherical projection. Solid and dashed curves show original and equalized linear-mean boundaries at zero prior; shading shows the original regions. (b) Exact centered scores and linear-mean approximations for the same pair. (c) All-class empirical distributions of Ac/τ0 in three projection spaces, with equal weight per class.
Figure 2: Accuracy and residual orders of the joint expansion. Left: median log error ratio of second-order joint and fixed-dimensional approximations over the prespecified temperature–cosine grid. Right: observed residual orders for successive approximations; diamonds mark theoretical orders, boxes show interquartile ranges, and whiskers use 1.5 interquartile ranges. Formulas, validity domains, and fitting details are in Theorem 4 and Appendix J .
Figure 3: Predicted and exact boundary displacement. Left: exact and joint-leading margins near the roots of a fixed CIFAR-100-LT pair. Right: theory versus exact displacement for all preselected pairs with a unique root in both conditions at zero prior. The legend gives comparable-pair coverage. The CIFAR-10 exhaustive extension has MAE 0.0116∘ (45/45 pairs); for CIFAR-100, ImageNet-LT, and iNaturalist, displacement MAEs are 0.119∘ , 0.0027∘ , and 0.0042∘ , respectively.
Figure 4: Intercept drift under matched gain interventions. (a) Maximum absolute classwise change in the cosine-zero score (linear scale below 0.01, logarithmic above). Open circles denote four separately fixed original ProCo states; stars denote interventions on the independently trained IF10 uniform-prior temperature-failure state. (b) Exact cross-entropy on the same 500 queries from the failure state. Within each state, queries, statistics, and prior are fixed; the three interventions share gc∗=gˉ .
Dataset / IF
Feature
Original vMF
Cosine
Pure Angular
Agreement
CIFAR-10 / 10
Encoder
85.21
85.82
85.90
99.74
Projection
87.85
87.84
87.83
99.99
CIFAR-10 / 50
Encoder
81.47
80.76
80.72
99.46
Projection
80.72
80.58
80.58
99.95
CIFAR-10 / 100
Encoder
78.76
78.19
78.02
99.51
Projection
77.81
77.59
77.58
99.93
Table 2: Frozen classification after complete gain equalization. Accuracy and cosine–Pure Angular prediction agreement (%) on identical features and class statistics, with zero evaluation offset and β=0 . CIFAR and ImageNet use original contrastive-only models; iNaturalist uses the official ProCo model. Full realization comparisons are in Appendix G ; evaluation protocols are given in Appendices I and J .
Appendix figures & tables30 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Geometric effects of score parameters. Temperature, bias, concentration, and prototype angle affect pairwise and one-versus-rest boundaries; class count affects only the latter. Joint B0/B2 are truncations through orders zero/two.
Figure 6: Finite-dimensional approach to the tail-error transition. (a,b) The transition narrows around Rt=0 . (c) Exact one-dimensional errors approach the large-deviation exponents. (d) Exact ProCo and the leading rule agree. Zero Monte Carlo counts are displayed at half a count and are not treated as zero probability.
Figure 7: Reversing the prior ordering can reverse the head–tail exponent order; the prior condition in the theorem is substantive.
Dataset / IF
Feature
Original
qtemp
qpreserve
qangular
Δ
CIFAR-10 / 10
Encoder
85.21
85.92
85.84
85.90
+0.69
Projection
87.85
87.83
87.83
87.83
−0.02
CIFAR-10 / 50
Encoder
81.47
80.71
80.72
80.72
−0.75
Projection
80.72
80.58
80.58
80.58
−0.14
CIFAR-10 / 100
Encoder
78.76
78.00
78.12
78.02
−0.74
Projection
77.81
77.58
77.59
77.58
−0.23
Appendix
Table 3: Frozen accuracy across score realizations. Accuracy (%) at β=0 and zero prior offset. Δ is Pure Angular minus original vMF in percentage points; darker green indicates a larger increase.
Dataset / IF
Feature
Disagree
Low-margin
Certified / Total
CIFAR-10 / 10
Encoder
26
26
9,763 / 10,000
Projection
1
1
9,982 / 10,000
CIFAR-10 / 50
Encoder
54
54
9,717 / 10,000
Projection
5
5
9,969 / 10,000
CIFAR-10 / 100
Encoder
49
49
9,655 / 10,000
Projection
7
7
9,960 / 10,000
Appendix
Table 4: Prediction disagreements and margin certificates. “Low-margin” reports the number of disagreements falling in the lowest cosine-gap decile. Certified queries satisfy gˉmcos>D ; all certified queries agree.
Dataset / IF
Feature
Mean ∣ΔL∣
Grad. error (%)
Median grad. cosine
CIFAR-100 / 10
Encoder
0.01141
1.555
0.99999712
Projection
0.01751
1.624
0.99996033
CIFAR-100 / 50
Encoder
0.00917
1.524
0.99999679
Projection
0.02206
1.592
0.99996161
CIFAR-100 / 100
Encoder
0.00832
1.448
0.99999724
Projection
0.02634
1.595
0.99995594
Appendix
Table 5: Loss and gradient differences between Pure Angular and cosine. Complete equalization with zero prior offset. Gradient error is 100∥Gangular−Gcos∥F/∥Gangular∥F , with query tangent gradients as rows of G . The last column reports their median cosine similarity.
Figure 8: Variance coupling in the multiclass expansion. Retaining the second-order variance term recovers the predicted third-order residual decay.
Figure 9: Support size and self-inclusion effects. (a) Resultant-length estimation error; bars show between-class SD. (b) Current-query inclusion effects; bars show average within-class query SD.
Figure 10: Class geometry and vMF goodness of fit. Left: class frequency and fitted concentration. Right: held-out projection goodness of fit.
Figure 11: Margins and fitted-score errors. Prototype-centre and random-typical margins versus reconstructed pairwise errors.
β
γ
ResNet-50
ResNeXt-50
0
-1
-1.665
-1.465
0
0
+0.195
+0.315
0
1
+0.175
-0.175
0.5
-1
-0.635
-0.610
0.5
0
+0.260
+0.240
0.5
1
+0.055
-0.040
Appendix
Table 6: Frozen gain interventions across backbones. ImageNet-LT projection-score accuracy changes in percentage points relative to ProCo ( β=1 ) at the same prior.
Figure 12: Prior-dependent effects of angular gain. Frozen vMF accuracy changes relative to ProCo at the same prior strength γ .
Figure 13: Angular gain and local gradients. (a) Gain dispersion during training. (b) Relative error of leading gradient increments. Colors identify interventions; line styles identify training states. (c) Exact (solid) and leading (dashed) increments projected onto a common plane.
CIFAR-10-LT
CIFAR-100-LT
Policy
IF50
IF100
IF50
IF100
ProCo
87.57
85.39
55.78
51.34
Reduced gain disparity
87.72
85.10
56.62
51.72
Increased gain disparity
87.56
85.14
55.99
52.05
Appendix
Table 7: Historical power-temperature intervention. Original classification-head test accuracy (%).
Figure 14: Changes in common linear readouts. Accuracy differences from ProCo at epoch 200 within each feature layer on CIFAR-100-LT IF100. Each classifier is fitted independently under the same fixed protocol. These probes are distinct from the original classification heads in Table 7 .
Figure 15: Contrastive-only training across sampling and prior settings. Balanced-support probe accuracy changes from ProCo in percentage points; colors are centered at zero.
Figure 16: Joint training with a uniform prior. Accuracy changes from ProCo in percentage points. Encoder/Projection use balanced-support probes; FC is the jointly trained head.
Figure 17: Joint training with the training-frequency prior. Accuracy changes from ProCo in percentage points. Encoder/Projection use balanced-support probes; FC is the jointly trained head.
Figure 18: Score and gradient trajectories during training. IF50 with a uniform prior: cosine-zero intercept drift and leading-gradient approximation error.
Dataset / IF
ProCo
Direct Temp.
Preserve
Pure Angular
Δ
CIFAR-10 / 10
87.46
87.29
87.41
87.24
-0.22
CIFAR-10 / 50
80.52
80.46
80.81
80.79
+0.27
CIFAR-10 / 100
77.79
77.39
76.71
76.50
-1.29
CIFAR-100 / 10
58.26
58.24
57.55
58.75
+0.49
CIFAR-100 / 50
47.63
46.73
48.82
48.10
+0.47
CIFAR-100 / 100
44.67
43.15
43.54
43.88
-0.79
Appendix
Table 15: Joint-training FC-head accuracy (%) with the training-frequency prior. All modified policies use β=0 . Δ is Pure Angular minus ProCo.
πc=∑jnjnc.
Appendix
Algorithm 1 Controlled score realization and representation learning
Setting
CIFAR-10/100-LT
ImageNet-LT
Encoder
ResNet-32
ResNet-50
Encoder dimension
64
2048
Projection dimension
128
1024
Training epochs
200
90
Global batch size
256
256
Optimizer
SGD
SGD
Appendix
Table 16: Main feature-learning hyperparameters.
Setting
Value or protocol
Linear-classifier epochs
100; fixed final-epoch representation
Linear-classifier optimizer
SGD; batch size 256; learning rate 0.1; momentum 0.9; zero weight decay; recorded cosine schedule
Linear-classifier objective
Ordinary cross-entropy; no extra prior offset; no vMF correction
Main linear fitting pool
Corresponding long-tail training subset; ImageNet uses 115,846 fitting images
Main Encoder preprocessing
CIFAR: raw Encoder features; ImageNet: unit-normalized Encoder features in the main comparison
Department of Engineering, Engineering Technology East Tennessee State University · Department of Electrical and Computer Engineering Tufts University · Department of Electrical and Computer Engineering Worcester Polytechnic Institute +1