Embedding models, often obtained via self-supervised learning, extract general-purpose representations from data. Quantifying the reliability of these representations is crucial, as many downstream models rely on them as input for their own tasks. To this end, we introduce a formal definition of representation reliability: the representation for a given test point is considered to be reliable if the downstream models built on top of that representation can, on average, consistently generate accurate predictions for that test point across various downstream tasks. However, accessing the downstream data to quantify the representation reliability is often limited or restricted for various reasons. We propose training-free methods for estimating the representation reliability without access to the downstream data. Our method is based on the concept of neighborhood consistency (NC) across distinct pre-trained representation spaces. The key insight is to find shared neighboring points as anchors to align these representation spaces before comparing them. We provide theoretical justifications for NC and develop two practical approaches: (1) directly computing NC when multiple pre-trained models are available, and (2) a perturbation-based NC (PNC), which creates synthetic ensembles from a single model through isotropic Gaussian noise, avoiding the computational cost of training deep ensembles. We further propose PNC-spread tuning, which systematically determines the perturbation magnitude by maximizing the spread of the PNC scores on a reference set. We demonstrate through comprehensive numerical experiments that our methods effectively capture the representation reliability with a high degree of correlation, achieving robust and favorable performance compared with baseline methods.
Figures & tables
Figure 1 : Illustration of representation reliability ( Reli ) and neighborhood consistency ( NC ). For a test point x∗ and a collection of pre-trained backbone models H={h1,⋯,hM} , the representation reliability is defined as the average performance of downstream models when using the representations of x∗ provided by the backbones in H . Our NC estimates Reli without requiring any prior knowledge of the downstream tasks. It operates by measuring the number of consistent neighbors of x∗ among reference points across different representation spaces.
Pretraining Algorithms
Method
ResNet-18
ResNet-50
CIFAR-10
CIFAR-100
CIFAR-10
CIFAR-100
Entropy
Brier
Entropy
Brier
Entropy
Brier
Entropy
Brier
SimCLR
NC100
0.3865
0.3262
0.3130
0.2590
0.3786
0.3206
0.3070
0.2516
Dist1
0.3021
0.2544
0.0757
0.0635
0.2750
0.2376
0.1100
0.0919
Norm
0.2907
0.2429
0.1198
0.0649
0.2789
0.2385
0.1311
0.0710
FV
-0.0476
-0.0458
-0.1891
-0.1710
-0.0402
-0.0393
-0.1984
-0.1870
Table 1 : Comparison of our neighborhood consistency ( NC100 ) with baseline methods in terms of their correlation with the representation reliability. We use Kendall’s τ coefficient to measure this correlation. Downstream tasks are in-distribution tasks where the model is pre-trained and fine-tuned on the same dataset. Performance on these tasks is evaluated using either negative predictive entropy or negative Brier score. The highest and second-highest scores are highlighted in bold and underlined, respectively. As shown, our method consistently receives a favorable score compared with baselines.
Table 2 : Comparison of our neighborhood consistency ( NC100 ) with baseline methods in terms of their correlation with the representation reliability. Here the downstream tasks are transfer learning tasks, where embedding functions are pre-trained on TinyImagenet and fine-tuned on CIFAR-10 , CIFAR-100 , and STL-10 .
Figure 2 : Ablation studies on the (a)–(b) ensemble size ( M ), (c)–(d) the number of neighbors ( k ) for NCk (ours) and baselines, and (e)–(f) the robustness to the choice of distance metric for NCk . Brier score is used for the downstream performance metric. The comprehensive results can be found in Appendix D .
Pretraining Algorithms
Method
ResNet-18
ResNet-50
CIFAR-10
CIFAR-100
CIFAR-10
CIFAR-100
Entropy
Brier
Entropy
Brier
Entropy
Brier
Entropy
Brier
SimCLR
PNC100s
0.3262
0.2779
0.2175
0.1724
0.2961
0.2525
0.2396
0.1954
Dist1
0.2383
0.2058
0.0580
0.0572
0.2351
0.2049
0.0973
0.0838
Norm
0.2668
0.2157
0.1242
0.0620
0.2629
0.2156
0.1370
0.0690
BYOL
PNC100s
0.2306
0.1361
0.2184
0.1364
0.3473
0.1764
0.3130
0.1231
Table 3 : Quantifying the reliability of individual embedding functions without retraining (single model only). Kendall’s τ coefficient correlation between each method and the downstream performance of individual models on in-distribution binary classification tasks. The baselines ( Dist1 , Norm ) are computed from the single model without ensemble averaging. The highest and second-highest scores are highlighted in bold and underlined, respectively.
Table 4 : Quantifying the reliability of individual embedding functions without retraining (single model only). Kendall’s τ coefficient correlation between each method and the downstream performance on transfer learning binary classification tasks. Embedding functions are pre-trained on TinyImagenet and fine-tuned on CIFAR-10 , CIFAR-100 , and STL-10 .
Figure 3 : Validation of PNC -spread tuning for SimCLR on in-distribution binary classification tasks. Blue: Kendall’s τ coefficient between PNC100 scores and downstream Brier scores as a function of σpert . Red: standard deviation of PNC100 scores on Xref . The peak correlation aligns with the peak spread, supporting the spread-based tuning criterion. Dashed gray line: largest considered σpert that did not cause a floating-point overflow.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4 : Graphical visualization of the sketch for the proof of Theorem 2 . Let Zi and Zj denote the representation spaces defined by the embedding functions hi and hj , respectively. Suppose that there is a reliable neighboring point xr that is located close to the test point x∗ in each representation space. For any downstream task t , a reliable neighboring point xr serves as an anchor for comparing different representations zi∗=hi(x∗) and zj∗=hj(x∗) of the test point x∗ . The key idea is that yi,t∗ and yj,t∗ — the downstream predictions on the test point using the two different embedding functions — should be similar because the predictions yi,t∗ and yi,tr as well as yj,tr and yj,t∗ are similar due to the Lipschitz continuity of the downstream predictors. Additionally, since xr is a reliable point, the predictions yi,tr and yj,tr are similar. Thus, it follows that yi,t∗ and yj,t∗ are similar as well.
Pretraining Algorithms
Method
ResNet-18
CIFAR-10
CIFAR-100
Entropy
Brier
Entropy
Brier
SimCLR
NC100
0.3865 ± 0.0024
0.3262 ± 0.0022
0.3130 ± 0.0031
0.2590 ± 0.0028
Dist1
0.3021 ± 0.0063
0.2544 ± 0.0051
0.0757 ± 0.0084
0.0635 ± 0.0075
Norm
0.2907
0.2429
0.1198
0.0649
FV
-0.0476
-0.0458
-0.1891
-0.1710
Appendix
Table 5 : Comparison on the correlation between our neighborhood consistency ( NC100 ) and baseline methods in relation to performance on in-distribution downstream tasks.
Pretraining Algorithms
Method
ResNet-18
TinyImagenet→
CIFAR-10
CIFAR-100
STL-10
Entropy
Brier
Entropy
Brier
Entropy
Brier
SimCLR
NC100
0.1729 ± 0.0090
0.1292 ± 0.0089
0.1680 ± 0.0032
0.1343 ± 0.0032
0.2176 ± 0.0112
0.1513 ± 0.0106
Dist1
0.1311 ± 0.0127
0.0933 ± 0.0115
0.0138 ± 0.0071
-0.0251 ± 0.0049
0.1098 ± 0.0151
0.0520 ± 0.0116
Norm
0.2124
0.1642
0.1220
0.0507
0.1641
0.0852
Appendix
Table 6 : Comparison on the correlation between our neighborhood consistency ( NC100 ) and baseline methods in relation to performance on transfer learning tasks from TinyImagenet to CIFAR-10 , CIFAR-100 , and STL-10 .
Pretraining Algorithms
Method
ResNet-18
CIFAR-10
CIFAR-100
Entropy
Brier
Entropy
Brier
SimCLR
NC100
0.3848 ± 0.0016
0.3202 ± 0.0019
0.2998 ± 0.0051
0.2166 ± 0.0030
Dist1
0.3025 ± 0.0079
0.2533 ± 0.0055
0.0641 ± 0.0070
0.0698 ± 0.0072
Norm
0.3098
0.2462
0.1725
0.0467
FV
-0.0485
-0.0357
-0.1755
-0.1358
Appendix
Table 7 : Comparison on the correlation between our neighborhood consistency ( NC100 ) and baseline methods in relation to performance on in-distribution downstream tasks, for the multi-class downstream classification tasks. Our approach consistently demonstrates strong performance when compared to baseline methods, similar to the results observed when using the OVO binary classification tasks, as shown in Table 5 .
Pretraining Algorithms
Method
ResNet-18
TinyImagenet→
CIFAR-10
CIFAR-100
STL-10
Entropy
Brier
Entropy
Brier
Entropy
Brier
SimCLR
NC100
0.1673 ± 0.0080
0.1230 ± 0.0073
0.1486 ± 0.0046
0.1031 ± 0.0018
0.2012 ± 0.0106
0.1271 ± 0.0100
Dist1
0.1331 ± 0.0122
0.0982 ± 0.0116
0.0082 ± 0.0064
-0.0162 ± 0.0051
0.0633 ± 0.0122
0.0382 ± 0.0107
Norm
0.2294
0.1654
0.1527
0.0348
0.1431
0.0624
Appendix
Table 8 : Comparison on the correlation between our neighborhood consistency ( NC100 ) and baseline methods in relation to performance on transfer learning tasks from TinyImagenet to CIFAR-10 , CIFAR-100 , and STL-10 , for the multi-class downstream classification tasks.
Pretraining Algorithms
Method
ResNet-18
CIFAR-10
CIFAR-100
Entropy
Brier
Entropy
Brier
SimCLR
NC100
0.3237 ± 0.0023
0.2727 ± 0.0021
0.2868 ± 0.0014
0.2408 ± 0.0021
Dist1
-0.0590 ± 0.0071
-0.0460 ± 0.0065
-0.0317 ± 0.0061
0.0027 ± 0.0064
Norm
0.2907
0.2429
0.1198
0.0649
FV
-0.3224
-0.2704
-0.2370
-0.1723
Appendix
Table 9 : Comparison on the correlation between our neighborhood consistency ( NC100 ) and baseline methods in relation to performance on in-distribution downstream tasks, for the unnormalized representation. The overall correlation is weaker compared to the one observed when using normalized representation as presented in Table 5 . Nevertheless, our approach consistently shows a robust performance compared to baseline methods.
Pretraining Algorithms
Method
ResNet-18
TinyImagenet→
CIFAR-10
CIFAR-100
STL-10
Entropy
Brier
Entropy
Brier
Entropy
Brier
SimCLR
NC100
0.1392 ± 0.0072
0.1025 ± 0.0067
0.1520 ± 0.0052
0.1218 ± 0.0057
0.1540 ± 0.0103
0.1054 ± 0.0080
Dist1
-0.1318 ± 0.0092
-0.1101 ± 0.0090
-0.1067 ± 0.0079
-0.0671 ± 0.0053
-0.1211 ± 0.0120
-0.0838 ± 0.0086
Norm
0.2124
0.1642
0.1220
0.0507
0.1641
0.0852
Appendix
Table 10 : Comparison on the correlation between our neighborhood consistency ( NC100 ) and baseline methods in relation to performance on transfer learning tasks from TinyImagenet to CIFAR-10 , CIFAR-100 , and STL-10 , for the unnormalized representation.
Figure 5 : Ablation over the ensemble size ( M ) for NCk (ours) and baselines. Brier score is used for the downstream performance metric.
Figure 6 : Ablation over the number of neighbors ( k ) for NCk (ours) and baselines. Brier score is used for the downstream performance metric.
Pretraining Algorithms
Method
ResNet-18
CIFAR-10
CIFAR-100
Entropy
Brier
Entropy
Brier
SimCLR
PNC100s
0.3300 ± 0.0325
0.2716 ± 0.0218
0.2183 ± 0.0278
0.1427 ± 0.0159
Dist1
0.2407 ± 0.0251
0.2053 ± 0.0176
0.0520 ± 0.0315
0.0640 ± 0.0134
Norm
0.2903 ± 0.0421
0.2146 ± 0.0319
0.1808 ± 0.0308
0.0363 ± 0.0260
BYOL
PNC100s
0.3198 ± 0.0740
0.0863 ± 0.0161
0.2793 ± 0.0884
0.0494 ± 0.0161
Appendix
Table 11 : Quantifying the reliability of individual embedding functions without retraining. Kendall’s τ coefficient correlation between each method and the downstream performance of individual models on in-distribution multi-class classification tasks. The baselines ( Dist1 and Norm ) are computed from the single model without ensemble averaging. The highest and second-highest scores are highlighted in bold and underlined, respectively.
Pretraining Algorithms
Method
ResNet-18
TinyImagenet→
CIFAR-10
CIFAR-100
STL-10
Entropy
Brier
Entropy
Brier
Entropy
Brier
SimCLR
PNC100s
0.1769 ± 0.0154
0.1301 ± 0.0117
0.1747 ± 0.0080
0.0947 ± 0.0048
0.2251 ± 0.0198
0.1479 ± 0.0149
Dist1
0.1210 ± 0.0159
0.0882 ± 0.0127
-0.0228 ± 0.0100
-0.0188 ± 0.0058
0.0589 ± 0.0191
0.0357 ± 0.0150
Norm
0.2214 ± 0.0121
0.1534 ± 0.0091
0.1363 ± 0.0101
0.0166 ± 0.0071
0.1435 ± 0.0214
0.0617 ± 0.0168
Appendix
Table 12 : Quantifying the reliability of individual embedding functions without retraining. Kendall’s τ coefficient correlation between each method and the downstream performance of individual models on transfer learning multi-class classification tasks. The baselines ( Dist1 and Norm ) are computed from the single model without ensemble averaging. Embedding functions are pre-trained on TinyImagenet and fine-tuned on CIFAR-10 , CIFAR-100 , and STL-10 . The highest and second-highest scores are highlighted in bold and underlined, respectively.
Pretraining Algorithms
Method
ResNet-18
ResNet-50
CIFAR-10
CIFAR-100
CIFAR-10
CIFAR-100
Entropy
Brier
Entropy
Brier
Entropy
Brier
Entropy
Brier
SimCLR
PNC100o
0.3361
0.2863
0.2405
0.1929
0.3116
0.2658
0.2475
0.2004
PNC100s
0.3262
0.2779
0.2175
0.1724
0.2961
0.2525
0.2396
0.1954
BYOL
PNC100o
0.2458
0.1509
0.2624
0.1656
0.3530
0.1911
0.3157
0.1405
PNC100s
0.2306
0.1361
0.2184
0.1364
0.3473
0.1764
0.3130
0.1231
Appendix
Table 13 : Oracle ( PNC100o ) vs. spread-tuned ( PNC100s ) Kendall’s τ coefficient correlation with downstream performance on in-distribution binary classification tasks.
Table 14 : Oracle ( PNC100o ) vs. spread-tuned ( PNC100s ) Kendall’s τ coefficient correlation with downstream performance on transfer learning binary classification tasks.
Pretraining Algorithms
Method
ResNet-18
ResNet-50
CIFAR-10
CIFAR-100
CIFAR-10
CIFAR-100
Entropy
Brier
Entropy
Brier
Entropy
Brier
Entropy
Brier
SimCLR
PNC100o
0.3397
0.2806
0.2367
0.1594
0.3107
0.2613
0.2412
0.1641
PNC100s
0.3300
0.2716
0.2183
0.1427
0.2946
0.2474
0.2318
0.1605
BYOL
PNC100o
0.3365
0.1013
0.3468
0.0736
0.4493
0.1082
0.5505
0.0516
PNC100s
0.3198
0.0863
0.2793
0.0494
0.4408
0.0932
0.5472
0.0172
Appendix
Table 15 : Oracle ( PNC100o ) vs. spread-tuned ( PNC100s ) Kendall’s τ coefficient correlation with downstream performance on in-distribution multi-class classification tasks.
Table 16 : Oracle ( PNC100o ) vs. spread-tuned ( PNC100s ) Kendall’s τ coefficient correlation with downstream performance on transfer learning multi-class classification tasks.
Figure 7 : Validation of PNC -spread tuning for BYOL and MoCo on in-distribution binary classification tasks. Blue: Kendall’s τ coefficient between PNC100 scores and downstream Brier scores as a function of σpert . Red: standard deviation of PNC100 scores on Xref . The peak correlation aligns with the peak spread, supporting the spread-based tuning criterion. Dashed gray line: largest considered σpert that did not cause a floating-point overflow.
Downstream Task Type
Method
ResNet-18
ResNet-50
CIFAR-10
CIFAR-100
CIFAR-10
CIFAR-100
Entropy
Brier
Entropy
Brier
Entropy
Brier
Entropy
Brier
Binary
PNC100∥,o
0.3622
0.3062
0.2490
0.1968
0.3636
0.3107
0.2546
0.2016
PNC100∥,s
0.3574
0.3025
0.2479
0.1961
0.3618
0.3087
0.2546
0.2016
PNC100⊥,o
0.3267
0.2787
0.2349
0.1901
0.2873
0.2446
0.2458
0.2018
PNC100⊥,s
0.3193
0.2719
0.2119
0.1677
0.2637
0.2246
0.2113
0.1759
Appendix
Table 17 : Quantifying the reliability of individual embedding functions without retraining for different choices of P . The following are Kendall’s τ coefficient correlation between each method and the downstream performance of individual models on in-distribution classification tasks.
Table 18 : Quantifying the reliability of individual embedding functions without retraining for different choices of P . The following are Kendall’s τ coefficient correlation between each method and the downstream performance of individual models on transfer learning classification tasks. Embedding functions are pre-trained on TinyImagenet and fine-tuned on CIFAR-10 , CIFAR-100 , and STL-10 .
School of Information and Communication Technology, Hanoi University of Science and Technology, 1 Dai Co Viet Road, Hai Ba Trung, Hanoi 100000, Vietnam
Centre for Vision Speech and Signal Processing, University of Surrey, United Kingdom · International Audio Laboratories Erlangen∗, Erlangen, Germany · Fraunhofer Institute for Integrated Circuits IIS, Erlangen, Germany