Large-scale Vision-Language Models (VLMs) exhibit strong zero-shot recognition, yet their real-world deployment is challenged by distribution shifts. While Test-Time Adaptation (TTA) can mitigate this, existing VLM-based TTA methods operate under a closed-set assumption, failing in open-set scenarios where test streams contain both covariate-shifted in-distribution (csID) and out-of-distribution (csOOD) data. This leads to a critical difficulty: the model must discriminate unknown csOOD samples to avoid interference while simultaneously adapting to known csID classes for accuracy. Current open-set TTA (OSTTA) methods rely on hard thresholds for separation and entropy minimization for adaptation. These strategies are brittle, often misclassifying ambiguous csOOD samples and inducing overconfident predictions, and their parameter-update mechanism is computationally prohibitive for VLMs. To address these limitations, we propose Prototype-based Double-Check Separation (ProtoDCS), a robust framework for OSTTA that effectively separates csID and csOOD samples, enabling safe and efficient adaptation of VLMs to csID data. Our main contributions are: (1) a novel double-check separation mechanism employing probabilistic Gaussian Mixture Model (GMM) verification to replace brittle thresholding; and (2) an evidence-driven adaptation strategy utilizing uncertainty-aware loss and efficient prototype-level updates, mitigating overconfidence and reducing computational overhead. Extensive experiments on CIFAR-10/100-C and Tiny-ImageNet-C demonstrate that ProtoDCS achieves state-of-the-art performance, significantly boosting both known-class accuracy and OOD detection metrics. Code will be available at https://github.com/O-YangF/ProtoDCS.
Figures & tables
Fig. 1: Performance comparison on CIFAR-10-C. Our ProtoDCS simultaneously achieves the highest known-class accuracy (ACC) and the best OOD detection performance (AUROC), significantly outperforming all baseline methods. The top-right region indicates better performance.
Fig. 2: Overview of the proposed ProtoDCS. For an incoming unlabeled sample x , ProtoDCS first computes its initial openness score Sopen(x) . The First-Check stage uses this score for an initial separation: confident samples ( Sopen(x)<Θa ) populate the diversity-aware visual cache to build visual prototypes Pv , while trustworthy samples ( Sopen(x)<Θb ) undergo evidence-driven optimization to generate temporary prototypes Pv′,Pt′ for a refined openness assessment. The Final-Verification stage then leverages a GMM on the (re-evaluated) openness scores for all samples to make a probabilistic csID/csOOD decision. Only the sample confirmed as csID through this double-check process contribute to the final prototype evolution, forming a robust closed-loop adaptation system free from csOOD corruption.
Fig. 3: Uncertainty Landscapes in Feature Space. We compare the uncertainty surface constructed by (a) Standard Entropy Minimization and (b) our Evidence-Driven Uncertainty-Aware (EDUA) loss. The background color intensity denotes the uncertainty level derived from the loss function, where Yellow/Light indicates lower uncertainty and Purple/Dark indicates higher uncertainty. Crucially, Entropy Minimization creates a deceptive landscape where csOOD samples fall into low-uncertainty regions. In contrast, ProtoDCS correctly maps the csOOD and boundary regions to high uncertainty (dark background), effectively filtering out risky samples during prototype evolution.
Method
CIFAR-10-C
CIFAR-100-C
Acc ↑
AUROC ↑
FPR@TPR95 ↓
OSCR ↑
Acc ↑
AUROC ↑
FPR@TPR95 ↓
OSCR ↑
CLIP (ViT)
66.39
81.63
63.82
13.89
37.14
55.42
91.40
18.87
TENT
61.96
80.66
67.07
53.70
33.03
67.45
86.26
25.08
SAR
58.97
76.59
73.91
49.81
25.15
54.69
89.12
14.11
EATA
59.76
79.92
67.76
52.41
32.16
66.55
86.04
24.21
TPT
64.50
79.06
72.55
55.62
34.39
42.18
87.90
23.89
TABLE I: Results of different methods on CIFAR benchmarks using the CLIP ViT-B/16 backbone. ↑ indicates that larger values are better, and vice versa. All values are percentages (%). Bold indicates the best.
Method
CIFAR-10-C
Acc ↑
AUROC ↑
FPR@TPR95 ↓
OSCR ↑
CLIP(RN)
41.17
58.91
89.44
18.34
TENT
21.40
61.94
88.44
14.44
SAR
19.46
67.54
82.71
14.73
EATA
18.64
57.60
91.17
12.11
TPT
19.80
44.23
96.38
12.13
TABLE II: Performance comparison using the CLIP ResNet-50 backbone, complementing the CLIP ViT-based results in Table I .
Method
Tiny-ImageNet-C
Acc ↑
AUROC ↑
FPR@TPR95 ↓
OSCR ↑
CLIP(ViT)
29.66
51.50
94.35
17.98
TENT
28.92
62.50
88.42
20.51
SAR
29.89
60.35
90.02
21.14
EATA
28.52
61.30
89.08
20.06
TPT
28.72
37.56
98.05
20.80
TABLE III: Results (%) of different methods on Tiny-ImageNet-C using the CLIP ViT-B/16 backbone. ↑ indicates that larger values are better, and vice versa. Bold indicates the best.
\Block 2-1Method
\Block 1-3 L uncertainty
\Block 1-4CIFAR-100-C
AU
EU
Lalign
Acc ↑
AUROC ↑
FPR ↓
OSCR ↑
\Block 5-1ProtoDCS
✓
36.42
75.31
75.21
32.47
✓
✓
37.23
75.73
75.38
31.50
✓
✓
36.42
76.29
75.21
32.49
✓
✓
37.09
76.29
76.11
32.94
✓
✓
✓
37.62
76.56
75.19
33.48
TABLE IV: Ablation of Evidence-driven Uncertainty-aware Loss Components. The baseline (Row 1) uses entropy minimization.
\Block 2-1Method
\Block 1-4 Core Components
\Block 1-4CIFAR-100-C
Check
Verify
Pv
Pt
Acc ↑
AUROC ↑
FPR ↓
OSCR ↑
\Block 5-1ProtoDCS
✓
✓
✓
36.59
71.88
85.53
32.85
✓
✓
✓
37.15
76.33
75.20
33.00
✓
✓
✓
37.15
76.29
76.11
32.96
✓
✓
✓
37.09
76.29
77.11
32.94
✓
✓
✓
✓
37.62
76.56
75.19
33.48
TABLE V: Ablation of Separation Mechanism and Prediction Strategy. We evaluate the impact of removing specific modules.
Fig. 4: t-SNE Visualization of the Visual Cache on CIFAR-100-C. This figure illustrates how the cached features evolve over time, comparing the state after 3,000 samples (left) with the state after 30,000 samples (right). As more data is processed, our diversity-aware update mechanism creates increasingly compact and representative class clusters.
Θa
Value
0.1
0.2
0.3
0.4
0.5
Acc (%)
37.58
37.61
37.62
37.60
37.57
Θb
Value
0.2
0.4
0.6
0.8
1.0
Acc (%)
37.61
37.62
37.62
37.62
37.60
τsim
Value
0.6
0.7
0.8
0.9
1.0
Acc (%)
37.58
37.58
37.60
37.62
37.60
TABLE VI: Sensitivity analysis of Gating and Cache hyperparameters ( Θa,Θb,τsim ) on CIFAR-100-C.
Θp
Value
0.3
0.4
0.5
0.6
0.7
0.9
Acc (%)
37.62
37.62
37.62
37.61
37.60
37.56
Wg
Value
10
30
50
100
250
500
Acc (%)
35.35
36.25
37.55
37.62
37.62
37.62
Wc
Value
10
30
50
100
250
500
Acc (%)
36.19
36.54
36.60
37.62
37.62
37.62
Θc
Value
0.4
0.5
0.6
0.7
0.8
0.9
TABLE VII: Sensitivity analysis of Verification ( Θp ), Window Sizes ( Wc,Wg ), and Loss Weights ( λalign,λAU ) on CIFAR-100-C.
Fig. 5: Calibration and Uncertainty Analysis on Tiny-ImageNet-C. This figure provides a tripartite analysis demonstrating the safe adaptation capability of ProtoDCS in open-set TTA. All models are implemented in the same open-set TTA setting. (a) Confidence Distribution: TPT exhibits a skewed, low-confidence distribution, indicating predictive confusion induced by Entropy Minimization (EM). ProtoDCS recovers a balanced, well-calibrated distribution through its evidence-driven uncertainty-aware (EDUA) loss. (b) Reliability Diagram: While C-TPT shows perfect calibration (diagonal alignment), ProtoDCS adopts a deliberately conservative strategy (curve above diagonal), systematically assigning underconfidence on uncertain samples to enhance safety against OOD data. (c) Accuracy-Uncertainty Curve: ProtoDCS achieves the highest AUC, proving its superior ability to rank samples by risk (high uncertainty for OOD/hard ID), which is crucial for reliable sample filtering during adaptation. Replacing the EDUA loss with standard entropy (ProtoDCS(entropy)) yields a lower AUC.
Fig. 6: Cold-Start Adaptation Dynamics on CIFAR-100-C. The red dashed line marks the GMM activation point (100th sample). The curve illustrates the transition from initial fluctuation to rapid performance climbing, and finally to a robust steady state, confirming the effectiveness of our progressive double-check mechanism.
Method
ImageNet
ImageNet-A
ImageNet-V2
ImageNet-R
ImageNet-S
Average
CLIP-ViT-B/16
66.73
47.87
60.86
73.98
46.09
59.11
TPT [ 4 ]
68.98
54.77
63.45
77.06
47.94
62.44
C-TPT [ 5 ]
69.30
52.90
63.40
78.00
48.50
62.42
DiffTPT [ 6 ]
70.30
55.68
65.10
75.00
46.80
62.28
TDA [ 9 ]
69.51
60.11
64.67
80.24
50.54
65.01
TPS [ 2 ]
70.19
60.08
64.73
80.27
49.95
65.04
TABLE VIII: Performance comparisons on ImageNet and its variants under a Closed-Set setting (Natural Distribution Shifts). bold and underline indicate the best and second-best results, respectively.
Method
Efficiency Metrics
Throughput ↑
Latency ↓
Memory ↓
CLIP (ViT)
172.48
5.80
356
TENT
45.66
21.90
12471
TPT
2.49
401.78
3795
DPE
80.93
12.40
356
ProtoDCS
55.86
17.90
372
TABLE IX: Efficiency evaluation on Tiny-ImageNet-C with ViT-B/16 backbone. We report Throughput (samples/sec ↑ ), Latency (ms/sample ↓ ), and Peak GPU Memory (MB ↓ ). bold and underlined indicate the best and second-best results among adaptation methods , respectively.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Fig. S1: Distribution of Final Openness Scores ( Sopen′ ) on CIFAR-10-C. This plot shows the score distribution for ID and OOD samples under the “brightness” corruption. The clear separation demonstrates the effectiveness of our temporary prototype refinement in enhancing the score’s discriminability.
ID:OOD
GMM Conv ↑
H-Dist ↑
Acc ↑
AUROC ↑
5:95
82.9
0.4151
59.40
88.22
10:90
84.8
0.5085
61.47
88.39
20:80
87.9
0.6490
67.52
88.57
30:70
92.7
0.6447
66.48
88.37
40:60
95.7
0.6748
67.55
88.34
50:50
95.3
0.6777
67.85
88.38
Appendix
TABLE S1: Evaluation of ProtoDCS across varying csID:csOOD skew ratios on CIFAR-10-C. ↑ indicates that larger values are better, and vice versa. All values are percentages (%). Bold indicates the best.
Fig. S2: Evaluation of ProtoDCS under varying csID:csOOD skew ratios on CIFAR-10-C. The background bar chart represents GMM convergence rate (left Y-axis), while the curves represent csID Acc and AUROC (right Y-axis).
Dataset
Sample Group
Avg Entropy
Entropy Offset
Avg EU
EU Offset
CIFAR-10-C
Correct csID
1.3698
-
0.6959
-
Misclassified csOOD
1.2193
-0.26 σ
0.7740
+0.96 σ (Intercepted)
CIFAR-100-C
Correct csID
3.2515
-
0.7351
-
Misclassified csOOD
2.9335
-0.32 σ
0.7563
+0.48 σ (Intercepted)
Tiny-ImageNet-C
Correct csID
3.7229
-
0.7040
-
Misclassified csOOD
3.0644
-0.57 σ
0.7054
+0.04 σ (Intercepted)
Appendix
TABLE S2: Quantitative evaluation of standard Softmax Entropy and Epistemic Uncertainty (EU) for correct in-distribution (csID) samples versus misclassified out-of-distribution (csOOD) samples. Standard deviation offsets ( σ -distance) relative to correct ID samples are reported. “Intercepted” indicates that these pseudo-high-confidence samples are successfully intercepted by our EDUA-guided separation mechanism.
Dataset
Total Samples
Rejected Samples
Rejection Rate (%)
ImageNet
50,000
0
0.00
ImageNet-A
7,500
0
0.00
ImageNet-R
30,000
0
0.00
ImageNet-S
50,000
0
0.00
ImageNet-V2
10,000
0
0.00
Appendix
TABLE S3: False OOD rejection-rate analysis of the ProtoDCS separation mechanism in pure closed-set scenarios.
Method
ImageNet
ImageNet-A
ImageNet-V2
ImageNet-R
ImageNet-S
Average
DPE
71.95
59.62
65.58
80.44
52.28
65.97
ProtoDCS (Original)
71.04
60.23
65.01
80.63
51.43
65.67
ProtoDCS (Modified)
71.93
59.61
65.60
80.43
52.25
65.96
Appendix
TABLE S4: Ablation study on the multimodal prediction fusion formulas in pure closed-set scenarios (Acc ↑ ).
Θa
Value
0.1
0.2
0.3
0.4
0.5
Acc (%)
30.05
31.25
31.21
30.02
27.85
Θb
Value
0.2
0.4
0.6
0.8
1.0
Acc (%)
29.08
31.18
31.21
31.02
30.75
τsim
Value
0.6
0.7
0.8
0.9
1.0
Acc (%)
31.02
31.15
31.21
31.21
31.05
Appendix
TABLE S5: Sensitivity analysis of Gating and Cache hyperparameters ( Θa,Θb,τsim ) on Tiny-ImageNet-C.
Fig. S3: Four representative incorrectly classified (pseudo-high-confidence) csOOD samples that exhibit pseudo-low values under standard Softmax entropy and their corresponding uncertainty indicators. Standard Softmax entropy and Epistemic Uncertainty (EU) are reported, along with their respective deviations relative to correct in-distribution samples ( σ -distance).
Θp
Value
0.3
0.4
0.5
0.6
0.7
0.9
Acc (%)
29.18
30.22
31.21
31.15
30.95
30.65
Wg
Value
10
30
50
100
250
500
Acc (%)
29.50
30.80
31.12
31.21
31.15
31.02
Wc
Value
10
30
50
100
250
500
Acc (%)
30.12
30.95
31.18
31.21
31.10
30.95
Θc
Value
0.3
0.4
0.5
0.6
0.7
0.8
Appendix
TABLE S6: Sensitivity analysis of Verification ( Θp ), Window Sizes ( Wc,Wg ), and Loss Weights ( λalign,λAU ) on Tiny-ImageNet-C.
Fig. S4: Dynamic evolution of the ID-Occupancy dynamics (cache quality) over 1500 test stream steps under different ID:OOD ratios (average of 15 corruptions on CIFAR-10-C). The vertical dotted line indicates the GMM activation point (step 100).
Test-time adaptation (TTA) has emerged as a promising paradigm for vision-language models (VLMs) to bridge the distribution gap between pre-training and test data. Recent works have focused on backpropagation-free TTA methods that rely on cache-based designs, but these introduce two key limitations. First, inference latency increases as the cache grows with the number of classes, leading to inefficiencies in large-scale settings. Second, suboptimal performance occurs when the cache contains insufficient or incorrect samples. In this paper, we present Prototype-Based Test-Time Adaptation (PTA), an efficient and effective TTA paradigm that uses a set of class-specific knowledge prototypes to accumulate knowledge from test samples. Particularly, knowledge prototypes are adaptively weighted based on the zero-shot class confidence of each test sample, incorporating the sample's visual features into the corresponding class-specific prototype. It is worth highlighting that the knowledge from past test samples is integrated and utilized solely in the prototypes, eliminating the overhead of cache population and retrieval that hinders the efficiency of existing TTA methods. This endows PTA with extremely high efficiency while achieving state-of-the-art performance on 15 image recognition benchmarks and 4 robust point cloud analysis benchmarks. For example, PTA improves CLIP's accuracy from 65.64% to 69.38% on 10 cross-domain benchmarks, while retaining 92% of CLIP's inference speed on large-scale ImageNet-1K. In contrast, the cache-based TDA achieves a lower accuracy of 67.97% and operates at only 50% of CLIP's inference speed.
Zhaohong Huang, Yuxin Zhang, Wenjing Liu +2
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University, 361005, P.R. China.
While test-time adaptation (TTA) empowers vision-language models to adapt without costly retraining, it remains highly vulnerable to out-of-distribution (OOD) outliers prevalent in real-world applications. This discrepancy motivates Noisy TTA (NTTA), an online task to filter noisy OOD samples on the fly while maximizing in-distribution (ID) classification accuracy. Existing zero-shot NTTA approaches typically rely on test-time discriminative training, leading to overconfident misclassifications and significantly degraded inference efficiency. To address these limitations, we propose a novel framework named Dual Distribution Estimation (DDE), shifting the zero-shot NTTA paradigm from instance-level learning to training-free Gaussian distribution modeling. DDE incorporates two novel modules: Positive Feature Distribution Estimation (PFDE) and Negative Label Distribution Estimation (NLDE). PFDE explicitly models class-wise inclusion and exclusion Gaussian distributions to formulate a calibrated contrastive score, robustly enhancing ID accuracy. In parallel, NLDE improves OOD identification by explicitly modeling the negative label distribution to mine highly discriminative labels, effectively mitigating spurious correlations. Extensive experiments show that on the large-scale ImageNet benchmark, DDE achieves an improvement of 3.70% in harmonic mean accuracy and reduces the FPR95 for OOD detection by 6.20%, while ensuring highly scalable and efficient online inference. Furthermore, DDE is zero-shot and training-free, demonstrating remarkable robustness in data-scarce scenarios. Codes are available at https://github.com/ZhuWenjie98/DDE.
Wenjie Zhu, Yabin Zhang, Liang Xu +3
The Hong Kong Polytechnic University · Harbin Institute of Technology (Shenzhen) · Shanghai Jiao Tong University +1
Vision-language models (VLMs) such as CLIP exhibit remarkable zero-shot capabilities, yet their performance frequently degrades sharply under unexpected test-time distribution shifts. While Test-Time Adaptation (TTA) offers a promising solution, continuously adapting VLMs over an unlabeled test stream presents fundamental challenges. Conventional top-1-centric updates often reinforce errors by corrupting the local semantic geometry among related classes, while iterative adaptation exacerbates progressive bias accumulation, ultimately driving the model toward mode collapse. To overcome these coupled vulnerabilities, we propose Local Margin Restoration (LMR), a lightweight, one-step TTA framework. At the sample level, our Protected Margin Restoration (PMR) objective recovers local semantic geometry by shielding plausible near-top candidates from external hard negatives. Concurrently, to combat stream-level degradation, we introduce a dual-stage stabilization mechanism, featuring an Adaptive Margin (AM) controller and Bias Correction (BC), to dynamically disrupt progressive bias accumulation and prevent mode collapse. Extensive experiments on CIFAR-C, ImageNet-C, and ImageNet variants demonstrate that LMR consistently outperforms state-of-the-art TTA baselines, proving exceptionally robust and efficient even in challenging low-batch test-time regimes. Our code is available at https://github.com/DennisHuangYan/LMR.
Yan Huang, Guowei Wang, Xu Wang +2
Guangzhou University Guangzhou, China · The Second Affiliated Hospital of Guangzhou University of Chinese Medicine Guangzhou, China · Jinan University Zhuhai, China +1