Test-Time Adaptation (TTA) and Generalized Category Discovery (GCD) are traditionally treated as disjoint problems: the former adapts models to domain shift assuming all test classes are known, while the latter discovers novel categories assuming labeled training data for known classes. However, real-world deployment rarely fits either setting. Motivated by this gap, we introduce Test-Time Generalized Category Discovery (TT-GCD), a unified and more realistic scenario where a vision-language model must adapt to distribution shifts, classify known categories using only textual supervision, and discover novel categories, all during test time and without access to labeled data. To address this challenging scenario, we propose PACT (Prototype Assignment for Category discovery at Test time), a fully unsupervised framework that casts known-class recognition and novel-class discovery via prototype assignment. PACT first re-aligns shifted visual features with the text-derived class representations of the VLM using confident zero-shot predictions. Known and novel categories are then both represented by prototypes in the visual embedding space, estimated from the unlabeled test stream, and each test image is assigned to the category whose prototype is most similar to its visual feature. Extensive experiments across corruption and domain-shift benchmarks demonstrate that PACT outperforms adapted state-of-the-art TTA and GCD methods, effectively bridging the gap between adaptation and discovery.
Figures & tables
Figure 1: The TT-GCD setting. A VLM trained on New York streets, with no labeled source data [ Left ], is deployed on an unlabeled Alberta stream [ Middle ] that combines domain shift (snow), known classes ( car ) and novel classes ( moose ). [ Right ] TTA misclassifies the novel class as a known class, Open-Set TTA discards it as an outlier, GCD needs labeled target-domain data, and TT-GCD adapts to the shift, classifies the known and discovers the novel.
Figure 2: Mean performance ( Known vs. novel ) over CIFAR-10-C, CIFAR-100-C and ImageNet-C. Dashed lines mark the best baseline on each axis.
Setting
Domain Shift
Novel Classes
Novel Discovery
No Labeled Train Data
TTA ( Wang et al., 2021 ; Shu et al., 2022 ; Maharana et al., 2025 ; Mishra et al., 2026 )
✓
✗
✗
✓
Open-Set TTA ( Lee et al., 2023 ; Gao et al., 2024 )
✓
✓
✗
✓
GCD ( Vaze et al., 2022 ; Chiaroni et al., 2023 ; Wang et al., 2024 )
✗
✓
✓
✗
GCD w/ Shift ( Wang et al., 2025 )
✓
✓
✓
✗
TT-GCD (Ours)
✓
✓
✓
✓
Table 1: Comparison of TT-GCD with related test-time and category-discovery settings. TT-GCD uniquely combines distribution shift, novel categories, novel-category discovery, and fully unlabeled test-time adaptation.
CIFAR-10-C
CIFAR-100-C
ImageNet-C
Method
All
Known
Novel
All
Known
Novel
All
Known
Novel
Test Time Adaptation
TENT++
33.41_{\text{\pm.10}}
60.06_{\text{\pm.12}}
6.75_{\text{\pm.25}}
15.78_{\text{\pm.01}}
17.16_{\text{\pm.02}}
14.41_{\text{\pm.04}}
13.14_{\text{\pm.07}}
19.58_{\text{\pm.12}}
6.69_{\text{\pm.03}}
BATCLIP++
29.24_{\text{\pm.08}}
46.88_{\text{\pm.01}}
11.59_{\text{\pm.16}}
15.57_{\text{\pm.01}}
15.31_{\text{\pm.07}}
15.83_{\text{\pm.09}}
15.74_{\text{\pm.06}}
18.83_{\text{\pm.05}}
12.64_{\text{\pm.07}}
SAT++
23.43_{\text{\pm.07}}
34.48_{\text{\pm.31}}
12.39_{\text{\pm.29}}
13.25_{\text{\pm.07}}
13.69_{\text{\pm.18}}
12.81_{\text{\pm.12}}
10.55_{\text{\pm.03}}
12.31_{\text{\pm.04}}
8.80_{\text{\pm.03}}
Open-Set Test Time Adaptation
Table 2: TT-GCD on corruption benchmarks (severity 5, mean over 15 corruption types). Clustering accuracy (%), mean over 3 seeds with standard deviation. Best in bold , second best underlined .
Method
Clipart
Infograph
Painting
Quickdraw
Real
Sketch
Average
TENT++
28.33_{\text{\pm.19}}
17.42_{\text{\pm.09}}
26.95_{\text{\pm.04}}
0.52_{\text{\pm.00}}
\underline{31.69}_{\text{\pm.03}}
23.58_{\text{\pm.14}}
21.41
BATCLIP++
\underline{29.09}_{\text{\pm.15}}
16.64_{\text{\pm.14}}
\underline{28.10}_{\text{\pm.31}}
\underline{6.04}_{\text{\pm.09}}
30.63_{\text{\pm.39}}
\underline{24.85}_{\text{\pm.40}}
22.56
SAT++
14.17_{\text{\pm.28}}
14.28_{\text{\pm.41}}
9.81_{\text{\pm.48}}
2.20_{\text{\pm.01}}
7.93_{\text{\pm.14}}
10.90_{\text{\pm.09}}
9.88
OSTTA++
26.82_{\text{\pm.34}}
19.19_{\text{\pm.18}}
25.69_{\text{\pm.40}}
5.71_{\text{\pm.21}}
27.21_{\text{\pm.67}}
23.92_{\text{\pm.12}}
21.42
UniEnt++
28.01_{\text{\pm.09}}
\underline{19.70}_{\text{\pm.14}}
26.81_{\text{\pm.17}}
5.50_{\text{\pm.25}}
30.24_{\text{\pm.34}}
24.80_{\text{\pm.37}}
22.51
SimGCD++
17.10_{\text{\pm 15.65}}
10.53_{\text{\pm 4.56}}
15.35_{\text{\pm 16.43}}
5.23_{\text{\pm.02}}
17.97_{\text{\pm 14.91}}
8.43_{\text{\pm 2.19}}
12.44
Table 3: TT-GCD under natural shift. DomainNet, per domain, K=172 . Clustering accuracy ( All , %), mean over 3 seeds with the standard deviation. Best in bold , second best underlined . Known and Novel per domain are reported in App. D.1 .
Method
B/16
L/14
BioCLIP-2
TENT++
9.54
15.01
31.97
BATCLIP++
11.57
15.52
28.63
SAT++
9.11
10.42
26.22
OSTTA++
8.86
14.00
28.82
UniEnt++
7.57
13.96
20.58
SimGCD++
13.74
17.85
24.76
Table 4: Fine-grained shift on CUB-C. All (%), severity 5.
Figure 3: (a) Time per batch of 128 (log scale) against mean All on DomainNet. (b) All (%) as a function of the buffer size W .
Figure 4: TT-GCD under estimated L . All accuracy (%) when the novel budget is estimated once per stream and shared by every method (App. D.3 ).
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Adapted parameters
Novel
Test-time objective
Test Time Adaptation
TENT++ [ Wang et al., 2021 ]
visual LN
✗
entropy minimization on known samples
BATCLIP++ [ Maharana et al., 2025 ]
visual + text LN
✗
entropy + image-to-text alignment + diversity
SAT++ [ Mishra et al., 2026 ]
visual LN
✗
Sinkhorn optimal-transport self-labeling
Open-Set Test Time Adaptation
OSTTA++ [ Lee et al., 2023 ]
visual LN
↑
known entropy min. − novel entropy + diversity
Appendix
Table 6: TT-GCD adaptation of each baseline, and PACT. LN denotes LayerNorm affine parameters. The Novel column indicates how the test-time objective treats novel samples: ✗ not at all, ↑ entropy maximization only, ✓ modeled as categories with their own parameters.
Table 7: Baseline hyperparameter sweeps on the four held-out CIFAR-100-C corruptions. baseline denotes the value of the original implementation. Best All (%) is on the held-out corruptions and is not comparable to Table 2 , which reports the 15 evaluation corruptions.
Dataset
Streams
∣C∣
K
L
Coarse-grained corruptions
CIFAR-10-C
15
10
5
5
CIFAR-100-C
15
100
50
50
ImageNet-C
15
1000
500
500
Fine-grained corruptions
CUB-C
7
200
100
100
Appendix
Table 8: Benchmarks. Each corruption type forms one test stream.
Parameter
Meaning
Value
Re-alignment
W
buffer size
1024
E
epochs
30
LR
learning rate
10−2
Streaming
emin
known activation threshold
1.5
Appendix
Table 9: PACT configuration, same for every benchmark.
CIFAR-100-C
DomainNet
Method
ms/batch
peak GB
ms/batch
peak GB
TENT++
1106
8.7
1050
10.3
BATCLIP++
1066
8.4
1181
10.5
SAT++
1014
5.6
1085
10.7
OSTTA++
1771
18.0
1673
18.1
UniEnt++
1758
18.0
1675
18.1
Appendix
Table 10: Cost per batch. Steady-state ms per batch of 128 and peak GPU memory (GB) on one H100 3g.40gb slice.
Method
Clipart
Infograph
Painting
Quickdraw
Real
Sketch
TENT++
\underline{49.29}_{\text{\pm.43}}
29.16_{\text{\pm.21}}
\underline{46.94}_{\text{\pm.06}}
0.82_{\text{\pm.30}}
\underline{57.02}_{\text{\pm.14}}
41.77_{\text{\pm.09}}
BATCLIP++
37.94_{\text{\pm.21}}
21.84_{\text{\pm.07}}
36.56_{\text{\pm.22}}
7.29_{\text{\pm.11}}
37.75_{\text{\pm 1.06}}
33.21_{\text{\pm.28}}
SAT++
16.79_{\text{\pm.14}}
13.51_{\text{\pm.52}}
11.53_{\text{\pm 1.37}}
2.57_{\text{\pm.28}}
9.32_{\text{\pm.40}}
13.59_{\text{\pm.89}}
OSTTA++
46.31_{\text{\pm.54}}
\underline{31.50}_{\text{\pm.74}}
43.63_{\text{\pm.93}}
\underline{8.68}_{\text{\pm.35}}
45.09_{\text{\pm.86}}
40.93_{\text{\pm.15}}
UniEnt++
48.79_{\text{\pm.19}}
31.41_{\text{\pm.54}}
45.98_{\text{\pm.36}}
7.98_{\text{\pm.40}}
51.87_{\text{\pm.91}}
\underline{42.35}_{\text{\pm.86}}
SimGCD++
23.70_{\text{\pm 23.75}}
10.65_{\text{\pm 4.88}}
22.43_{\text{\pm 27.54}}
7.30_{\text{\pm.54}}
25.42_{\text{\pm 23.25}}
10.95_{\text{\pm 3.69}}
Appendix
Table 11: DomainNet, Known accuracy (%) per domain, mean over 3 seeds with the standard deviation. Best in bold , second best underlined .
Method
Clipart
Infograph
Painting
Quickdraw
Real
Sketch
TENT++
6.65_{\text{\pm.07}}
5.93_{\text{\pm.13}}
7.18_{\text{\pm.11}}
0.23_{\text{\pm.30}}
6.22_{\text{\pm.08}}
7.22_{\text{\pm.29}}
BATCLIP++
\underline{19.95}_{\text{\pm.12}}
11.55_{\text{\pm.22}}
\underline{19.74}_{\text{\pm.41}}
\underline{4.80}_{\text{\pm.13}}
\underline{23.48}_{\text{\pm.93}}
\underline{17.34}_{\text{\pm.53}}
SAT++
11.46_{\text{\pm.57}}
\underline{15.04}_{\text{\pm.36}}
8.11_{\text{\pm.39}}
1.84_{\text{\pm.29}}
6.53_{\text{\pm.50}}
8.49_{\text{\pm.63}}
OSTTA++
6.68_{\text{\pm.18}}
7.13_{\text{\pm.61}}
7.95_{\text{\pm.17}}
2.77_{\text{\pm.19}}
9.25_{\text{\pm.49}}
8.62_{\text{\pm.11}}
UniEnt++
6.53_{\text{\pm.06}}
8.24_{\text{\pm.76}}
7.86_{\text{\pm.68}}
3.05_{\text{\pm.43}}
8.50_{\text{\pm.47}}
9.01_{\text{\pm 1.13}}
SimGCD++
10.28_{\text{\pm 7.31}}
10.42_{\text{\pm 4.26}}
8.35_{\text{\pm 5.43}}
3.17_{\text{\pm.51}}
10.49_{\text{\pm 6.55}}
6.15_{\text{\pm 1.02}}
Appendix
Table 12: DomainNet, Novel accuracy (%) per domain, mean over 3 seeds with the standard deviation. Best in bold , second best underlined .
CLIP B/16
CLIP L/14
BioCLIP-2 L/14
Method
All
Known
Novel
All
Known
Novel
All
Known
Novel
TENT++
9.54
12.03
7.04
15.01
21.81
8.21
31.97
59.26
4.73
BATCLIP++
11.57
11.37
11.77
15.52
17.03
14.01
28.63
45.70
11.60
SAT++
9.11
9.05
9.18
10.42
11.37
9.48
26.22
40.08
12.40
OSTTA++
8.86
11.58
6.15
14.00
20.92
7.11
28.82
53.15
4.53
UniEnt++
7.57
9.72
5.42
13.96
18.60
9.33
20.58
34.33
6.85
Appendix
Table 13: Backbone study on CUB-C (7 corruption types). Clustering accuracy (%). All is in Table 4 .
CIFAR-10-C
CIFAR-100-C
DomainNet
ImageNet-C
L^=5 , L=5
L^=27 , L=50
L^=135 , L=173
L^=113 , L=500
Method
All
Known
Novel
All
Known
Novel
All
Known
Novel
All
Known
Novel
Test Time Adaptation
TENT++
33.20
59.89
6.50
15.47
17.02
13.92
21.50
37.48
5.77
14.25
20.29
8.21
BATCLIP++
28.90
46.73
11.08
15.54
15.34
15.74
22.53
28.98
16.19
16.07
18.83
13.30
SAT++
23.27
34.20
12.33
13.19
13.71
12.67
9.85
11.15
8.57
11.48
13.17
9.80
Appendix
Table 14: TT-GCD under estimated L . Clustering accuracy (%). Best in bold , second best underlined .
CIFAR-100-C
DomainNet
Rate
All
Known
Novel
All
Known
Novel
Known rate η
0.05
28.99
33.74
24.25
46.64
53.70
39.68
0.10
29.07
34.62
23.51
46.39
54.87
38.01
0.30
28.23
35.47
20.99
43.65
56.74
30.70
0.60
27.22
36.09
18.35
40.26
57.27
23.40
Appendix
Table 15: Prototype update rates, varied one at a time around the fixed value. All, Known and Novel (%), mean over CIFAR-100-C corruptions and DomainNet domains.
Figure 5: Activation threshold emin on CIFAR-100-C and DomainNet.
Generalized Category Discovery (GCD) aims to categorize unlabelled instances from both known and unknown classes by transferring knowledge from labelled data of known classes. Existing methods assume all data comes from a single domain, yet real-world unlabelled data often exhibits domain shifts alongside semantic shifts. We study GCD under domain shifts and propose three frameworks that adapt foundation models, ranging from self-supervised vision models to vision-language models. (i) HiLo disentangles domain and semantic features through multi-level feature extraction and mutual information minimization, combined with PatchMix augmentation and curriculum sampling. (ii) HLPrompt extends HiLo with semantic-aware spatial prompt tuning to suppress background and domain noise. (iii) VLPrompt leverages vision-language models via factorized textual prompts and cross-modal consistency regularization. The three methods share core design principles while operating on different foundation backbones, making them suitable for different deployment scenarios. Extensive experiments on synthetic corruptions and real-world multi-domain shifts demonstrate consistent improvements over strong baselines. Project page: https://visual-ai.github.io/hilo/
Hongjun Wang, Po Hu, Kai Han
School of Computing and Data Science, The University of Hong Kong
Generalized Category Discovery (GCD) seeks to identify novel categories from unlabeled data while retaining the classification ability of seen categories. Prior GCD methods commonly leverage transferable representations from pre-trained models, adapting to downstream datasets via partial fine-tuning (updating only the final ViT block) and visual prompt tuning (appending learnable vectors to inputs). However, conventional partial fine-tuning offers limited flexibility, as it fails to adapt the entire model; meanwhile, visual prompt tuning is prone to overfitting, due to its sensitivity to initialization and inherently constrained capacity. To address these limitations, we propose LAGCD, a simple yet effective GCD approach that embeds a residual linear adapter into each ViT block. From the perspective of feature sparsity, we systematically show that non-linearity in conventional adapters impairs performance, whereas our linear adapter enhances it by enabling more flexible model capacity. We further introduce an auxiliary distribution alignment loss to mitigate the negative impact of biased predictions between seen and novel categories. Extensive experiments on both generic and fine-grained datasets confirm that LAGCD consistently improves performance over many sophisticated baselines. The source code is available at https://github.com/yebo0216best/LAGCD
Bo Ye, Kai Gan, Tong Wei +1
School of Computer Science and Engineering, Southeast University, Nanjing 210096, China, and the Key Laboratory of Computer Network and Information Integration (Southeast University), Ministry of Education, China.
Test-time adaptation (TTA) has emerged as a promising paradigm for vision-language models (VLMs) to bridge the distribution gap between pre-training and test data. Recent works have focused on backpropagation-free TTA methods that rely on cache-based designs, but these introduce two key limitations. First, inference latency increases as the cache grows with the number of classes, leading to inefficiencies in large-scale settings. Second, suboptimal performance occurs when the cache contains insufficient or incorrect samples. In this paper, we present Prototype-Based Test-Time Adaptation (PTA), an efficient and effective TTA paradigm that uses a set of class-specific knowledge prototypes to accumulate knowledge from test samples. Particularly, knowledge prototypes are adaptively weighted based on the zero-shot class confidence of each test sample, incorporating the sample's visual features into the corresponding class-specific prototype. It is worth highlighting that the knowledge from past test samples is integrated and utilized solely in the prototypes, eliminating the overhead of cache population and retrieval that hinders the efficiency of existing TTA methods. This endows PTA with extremely high efficiency while achieving state-of-the-art performance on 15 image recognition benchmarks and 4 robust point cloud analysis benchmarks. For example, PTA improves CLIP's accuracy from 65.64% to 69.38% on 10 cross-domain benchmarks, while retaining 92% of CLIP's inference speed on large-scale ImageNet-1K. In contrast, the cache-based TDA achieves a lower accuracy of 67.97% and operates at only 50% of CLIP's inference speed.
Zhaohong Huang, Yuxin Zhang, Wenjing Liu +2
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University, 361005, P.R. China.