Test-Time Training (TTT) adapts models to incoming test samples (e.g. out-of-distribution, (OOD)) when conventional fine-tuning is infeasible. Existing TTT methods for Vision-Language Models (VLMs) create supervision from several augmented views, each requiring forward (and often backward) passes through the entire VLM, incurring substantial computational cost. We observe that one forward pass with all the intermediate layer outputs already yields far more signal than the final embedding from all augmentations. We introduce Layer Query Network (LQN), a lightweight approach that can adapt a frozen VLM (teacher) in a single forward pass of the VLM via a small model (student). LQN uses Position-Aware Distillation (PAD) to mimic the teacher VLM's intermediate-layer spatial tokens by querying spatial coordinates of intermediate tokens. LQN additionally relies on Location Consistency Regularization (LCR), a self-supervision technique, replacing expensive O(H x W) image augmentation with O(1) coordinate sampling. Integrating these, LQN i) adapts and improves zero-shot CLIP ViT-B/16 by 9.8% Top-1 on OOD ImageNet, ii) outperforms the previous best GS-Bias on fine-grained classification by 3.9% Top-1, iii) achieves faster convergence than TPS for CLIP ResNet-50 (47 mins vs 55 mins), iv) generalizes adaptation to VLMs like SigLIP, EVA-CLIP, and CoCa, and lightweight students like MLP, ResNet, VGG, and v) extends to panoptic, instance, and semantic segmentation.
Figures & tables
Figure 1: Comparison with existing works: (a) Left: TPT trains prompts via multiple augmentations, requiring multiple forward/backward passes of VLM (b) Middle: LQN trains small student in single forward pass from VLM (w/o any augmentation), by iteratively distilling intermediate features of image encoder via positional encoding ( i,j,l coordinates for i,j spatial token at l layer) (c) Right: Self-supervision via varying src coordinates saves compute on expensive image augmentation and avoids over-reliance on VLMs (teacher) embedding on OOD samples.
Figure 2: Layer Query Network (LQN) 1) For a single test image, frozen teacher (VLM) extracts intermediate features across all layers. During test-time training , LQN runs for N iterations by randomly sampling different destination (dest) positions (i,j,l) , where (i,j) is the spatial location and l the layer depth. Position-Aware Distillation ( PAD ) takes the image, (0,0,0) as the src position, and trains student network to predict the teacher’s feature at dest . Location Consistency Regularization ( LCR ) self-supervises LQN to produce consistent features across random src (s). 2) After the final iteration (test-time inference), the student is queried at all spatial locations of the last layer L . Resultant features are processed by task-specific modules (eg, zero-shot classification/ segmentation).
Method
Augment.
ImageNet
ImageNet-A
ImageNet-V2
ImageNet-R
ImageNet-Sketch
Avg
OOD Avg
CLIP-B/16
CLIP-ViT-B/16
✗
66.7
47.8
60.8
73.9
46.0
59.0
57.1
Ensemble
✗
68.3
49.8
61.8
77.6
48.2
61.1
59.4
TPT
✓
68.9
54.7
63.4
77.0
47.9
62.4
60.8
Diff-TPT
✓
70.3
55.6
65.1
75.0
46.8
62.6
60.6
MTA + TPT
✓
70.0
58.0
64.2
78.3
49.6
64.0
62.5
APM
✗
68.1
52.1
67.2
76.5
49.3
62.6
61.3
Table 1: Natural distribution shifts: Top-1 Results for ImageNet and its distribution-shifted variants (ImageNet-A/-V2/-R/-Sketch). Requirements column specifies resources needed during test-time; Augment. : multiple augmented views of each test sample; LQN doesn’t need any augmented views for test samples. Use of the CLIP variant as a teacher is shown on the left. Avg: across all cols, OOD Avg: all cols except ImageNet.
Method
Augment.
Flower102
DTD
Pets
UCF101
Caltech101
Food101
SUN397
Aircraft
EuroSAT
Avg
CLIP-ViT-B/16
✗
67.4
44.3
88.3
65.1
93.4
83.7
62.6
23.7
42.0
63.4
Ensemble
✗
67.0
45.0
86.9
65.2
93.6
82.9
65.6
23.2
50.4
64.4
TPT
✓
69.0
47.8
87.8
68.0
94.2
84.7
65.5
24.8
42.4
64.9
DiffTPT
✓
70.1
47.0
88.2
62.6
92.4
87.2
65.7
25.6
43.1
64.7
MTA
✓
68.0
45.9
88.2
68.6
94.2
85.0
66.6
25.2
45.3
65.2
APM
✗
62.0
48.9
81.6
72.6
89.6
84.2
65.7
29.7
55.7
65.5
Table 2: Cross-dataset generalization from ImageNet to fine-grained classification tasks. Top-1 accuracy reported. CLIP-B/16 as teacher. Avg. averages all columns.
Panoptic
Instance
Semantic
COCO
ADE20K
COCO
Cityscapes
ADE20K
Method
PQ ↑
PQ ↑
AP ↑
mIoU ↑
mIoU ↑
EoMT
58.3
51.7
48.8
84.2
58.4
LQN (Our)
64.5
57.1
55.2
87.5
64.3
LQN (Pretrained, Our )
64.9
58.6
56.7
88.2
65.8
Table 3: Dense-task Segmentation Input resolution of 1280×1280 ; semantic segmentation has 1024×1024 . PQ: Panoptic Quality, AP: Average Precision. LQN performs TTT on EoMT (teacher).
ImageNet and OOD
Fine-grained classification
Method
ImageNet
IN-A
IN-V2
IN-R
IN-Sketch
Avg
OOD Avg
Flowers102
DTD
Pets
UCF101
Caltech101
Food101
SUN397
Aircraft
EuroSAT
Avg
SigLIP
76.0
45.3
68.9
90.3
67.9
69.7
68.1
85.8
64.7
94.1
72.5
90.5
89.8
69.8
43.8
43.8
72.8
+ LQN
79.2
48.7
72.4
93.4
71.6
73.1
71.5
89.2
68.5
97.4
76.2
94.1
93.1
73.5
47.3
47.6
76.3
EVA-CLIP
76.1
64.6
68.9
89.1
63.3
72.4
71.5
72.0
59.3
93.7
74.1
90.4
89.7
71.9
28.5
69.9
72.2
+ LQN
79.8
68.9
73.1
92.5
67.4
76.3
75.5
76.4
63.2
97.2
78.6
94.3
93.7
75.9
32.7
74.1
76.2
CoCa
63.6
21.5
55.7
73.2
51.3
53.1
50.4
64.7
53.3
89.1
61.4
89.1
77.3
66.1
18.8
45.3
62.8
Table 4: Generalization across various VLMs teacher LQN+{SigLIP, EVA-CLIP for various classification benchmarks. Similar averaging as table 1 and table 2 .
Figure 3: Wall-clock time CLIP ResNet-50 on ImageNet.
Figure 4: Left: Segmentation Mask Quality: EoMT vs LQN (Above) LQN can even segment persons occluded behind the window of the bus. (Bottom) LQN semantically groups visual elements of the scene ( e.g. walls), whereas EoMT falls short. Middle: Cost vs. Performance. Left) TPT’s 63 augmentations result in high overhead. Right: Generalization: Performance improves even at unseen final layers (COCO PQ), showing LQN’s ability to generalize beyond TTT teacher layers.
Figure 5: Left: Token visualizations of the teacher model EoMT (top), LQN (middle), and their difference ( Error , bottom) showing LQN’s predictions matching EoMT. Right (i) Augmenting test image during TTT harms LQN (red), but improves CLIP ViT-B/16 (blue). (ii) Distilling more teacher locations improves performance, verifying benefits of intermediate layer distillation. naug & nlocation represent the # of augmentations and the # of coordinates to distill in PAD.
Parts
Food-101
COCO
PAD (- src )
71.8
54.9
PAD + △
76.0
58.1
PAD + LCR
89.2
64.5
Table 5: LCR variations W/o src , only PAD can be applied. △ : LCR w/ Gaussian noise instead of src .
Parts
Food-101
COCO
PAD (- src )
71.8
54.9
PAD + △
76.0
58.1
PAD + LCR
89.2
64.5
Table 5: LCR variations W/o src , only PAD can be applied. △ : LCR w/ Gaussian noise instead of src .
Student
Food-101
VGG
84.9
ResNet 18
85.4
ResNet 34
87.2
MLP
90.8
Table 6: Student Architectures .
Loss
Food-101
COCO
L1
75.8
52.0
MSE
85.7
61.1
Cosine
89.2
64.5
Table 7: Type of loss in PAD & LCR
Pos.
Food-101
COCO
Coord.
33.1
27.9
Sin-Cos
88.6
63.8
RoPE
89.2
64.5
Table 8: Positional Encodings . Coord. feeds int positions.
Sampling Strategy
Food-101
COCO
Sampling more from shallow layers
81.9
43.7
Sampling more from deeper layers
87.8
52.0
Uniform Sampling
89.2
64.5
Table 9: Sampling Strategies for VLM Layers.
Sampling Strategy
Food-101
COCO
Sampling more from shallow layers
81.9
43.7
Sampling more from deeper layers
87.8
52.0
Uniform Sampling
89.2
64.5
Table 9: Sampling Strategies for VLM Layers.
Configuration
Food-101
COCO
PAD + LCR (w/o weight Reg.)
87.4
62.6
PAD + LCR (w/ weight Reg.)
89.2
64.5
Table 10: Effect of Weight Regularization (Reg.) on LCR ( Equation 4 )
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Requirements
ImageNet ↑
TDA [CVPR’24]
History
69.5
DMN-ZS [CVPR’24]
History
72.2
DPE [NeurIPS’24]
History
71.9
DynaPrompt [ICLR’25]
History
72.9
CoOp [IJCV’22]
Labeled Data
71.5
CoCoOp [CVPR’22]
Labeled Data
71.0
Appendix
Table 11: Test Time Augmentation . History means cumulative of test samples. CLIP-B/16 (teacher).
Layer
Feature Dimension
nkernels
Stride
Padding
(H × W × C)
Input / Output
Input
h×w×c
Encoder
Conv
h/s×w/s×d
1
s
0 / 0
Decoder
Linear
(h/s∗w/s+dp+dp†) * 4096
-
-
-
Linear
4096 * 4096
-
-
-
Linear
4096 * 2048
-
-
-
Appendix
Table 12: LQN architecture for TTT : with input dimensions h,w,c and feature dimension dp : dimensionality of positional encoding. s : stride of convolutional filter in encoder, dc : dimension of the intermediate token of teacher on which LQN learns. LQN contains two locations, (src,dest) , so additional dp term is added to the input. †: LQN-[src] contains only a single location, so this term is not added.
Number of Test samples
50000 (Imagenet Splits), variable for other datasets.
Testing iterations
15
Batch Size
1
Learning Rate
1e-4
Optimizer
Adam
Feature Output size d
768/1024
Positional Encoding size
768/1024
Appendix
Table 13: LQN hyperparameters during test-time-training.
Dataset
ImageNet
Flowers102
DTD
Caltech101
Aircraft
IN-A
IN-v2
Pets
Sun397
Eurosat
IN-R
UCF101
Food101
α
0.7
0.5
0.7
0.5
0.7
0.6
0.6
0.6
0.6
0.6
0.8
0.8
0.8
ϵ (in LCR)
1e-7
1e-7
1e-7
1e-7
1e-7
1e-7
1e-7
1e-7
1e-7
1e-7
1e-7
1e-7
1e-7
Appendix
Table 14: Values of α , i.e. the weight of LCR loss for different datasets used in the experiments. α=0.7 , for the remaining datasets not shown here, for eg, COCO, ADE20k. Optimal values were found by grid search and choosing the parameters that gave lowest values of the LCR loss. Note that this procedure does not use any labels at test-time.
Figure 6: Fixed src performs best.
N
5
8
12
15
20
25
Accuracy
80.5
83.9
88.6
89.2
87.8
81.0
Appendix
Table 15: Model performance (Accuracy) across different iteration counts N .
Figure 7: Qualitative results on LQN. LQN reconstructs RGB images with lowest pixelwise loss (numbers marked in yellow) as compared to MAE/APM. This illustrates the extreme scenario when input images contain repeating textures and tests whether the model encodes semantics consistently. Note that these results are not cherry-picked.
Memory
ORIN
AGX Thor
Ampere
GFlops
EoMT (teacher)
2.1 GB
0.5 sec
0.1sec
0.08sec
4146
LQN (student) / iter
600 MB
0.68sec
0.37sec
0.29sec
24.5
LQN total (15 iters)
2.7 GB
8.9sec
4.3sec
2.88sec
4384
Appendix
Table 16: Memory and GFlop Analysis.
Model
CLIP ViT-L
2 Layers
4 Layers
8 Layers
10 Layers
Accuracy
74.9
75.8
76.5
77.3
77.0
Appendix
Table 17: Accuracy comparison across different layer configurations.
Model
Base
+ LQN
Vision-Mamba-T
78.3
80.9
Vision-Mamba-S
81.4
83.2
Appendix
Table 18: Performance comparison of Vision-Mamba models with LQN.
P
brigh
cont
defoc
elast
fog
frost
gauss
glass
impul
jpeg
motn
pixel
shot
snow
zoom
Average
Joint Train
✓
62.3
4.5
26.7
39.9
25.7
30.0
5.8
16.3
5.8
45.3
30.9
45.9
7.1
25.1
31.8
24.8
Fine-Tune
✓
67.5
7.8
33.9
32.4
36.4
38.2
22.0
15.7
23.9
51.2
37.4
51.9
23.7
37.6
37.1
33.7
ViT Probe
✓
68.3
6.4
24.2
31.6
38.6
38.4
17.4
18.4
18.2
51.2
32.2
49.7
18.2
35.9
32.2
29.2
TTT-MAE
✓
69.1
9.8
34.4
50.7
44.7
50.7
30.5
36.9
32.4
63.0
41.9
63.0
33.0
42.8
45.9
44.4
OpenCLIP VIT-L/14(t)
✗
71.9
47.0
50.3
32.7
58.3
46.9
26.0
26.5
28.1
62.7
37.7
58.3
28.2
50.4
37.9
42.1
APM
✗
77.4
51.9
56.6
37.9
64.8
53.2
28.7
31.4
33.0
68.4
44.1
64.5
33.1
56.9
43.9
50.3
Appendix
Table 19: LQN’s performance on ImageNet-C, level 5 . The first three rows are fixed models without test-time training. The third row, ViT probing, is the baseline used in ( Gandelsman et al., 2022 ) . A ✓ in P means that method leveraged pre-trained weights on clean variant of train set aka, Image-net and downstream-ttt on corrupted version. OpenCLIP VIT-L/14 is generally more robust. LQN does better than APM.
P
brigh
cont
defoc
elast
fog
frost
gauss
glass
impul
jpeg
motn
pixel
shot
snow
zoom
Average
Baseline
✓
73.1
33.1
35.8
56.9
54.2
45.2
39.6
26.0
38.2
62.0
43.2
60.3
32.2
44.2
40.7
47.4
TTT-MAE
✓
72.7
39.6
45.7
64.9
58.3
52.6
48.5
42.8
47.6
67.0
50.5
66.6
42.4
45.7
51.5
53.2
OpenCLIP VIT-L/14
✗
74.2
64.2
58.7
57.8
66.3
52.8
45.3
34.6
45.2
68.9
46.6
63.9
41.1
56.2
45.6
54.8
APM
✗
79.2
70.4
64.9
63.7
72.3
58.6
51.2
40.4
51.3
74.1
53.0
70.0
46.7
62.5
51.8
59.6
LQN (Ours)
✗
80.9
72.1
66.5
65.2
73.9
59.9
52.9
41.9
52.8
75.3
54.2
71.4
47.9
63.9
53.3
61.2
Appendix
Table 20: Performance on ImageNet-C, level 4 . The first two rows are from the supplementary materials of ( Gandelsman et al., 2022 ) . A ✓ in column P indicates use of pre-trained weights on clean ImageNet followed by TTT on corrupted inputs. OpenCLIP ViT-L/14 shows stronger robustness than earlier models. LQN does better than APM.
P
brigh
cont
defoc
elast
fog
frost
gauss
glass
impul
jpeg
motn
pixel
shot
snow
zoom
Average
Baseline
✓
75.8
62.7
49.5
67.1
59.8
47.6
57.1
35.0
57.4
68.6
60.2
70.1
54.3
54.7
48.0
57.6
TTT-MAE
✓
75.8
64.4
59.4
71.2
64.0
54.0
63.6
50.7
64.2
71.3
64.2
73.1
61.8
58.0
57.4
64.4
OpenCLIP VIT-L/14
✗
75.8
71.8
65.5
67.7
69.0
54.7
58.9
42.4
59.5
72.8
59.9
69.7
58.2
63.5
51.8
62.5
APM
✗
80.5
77.2
71.3
73.3
74.8
60.6
64.7
48.5
65.4
77.8
61.6
75.2
64.1
69.3
58.0
68.5
LQN (Ours)
✗
81.9
78.8
72.8
74.9
76.2
61.9
66.3
49.9
66.9
79.1
62.8
76.7
65.5
70.5
59.5
69.9
Appendix
Table 21: Performance on ImageNet-C, level 3 . The first two rows are from the supplementary materials of ( Gandelsman et al., 2022 ) . A ✓ in column P indicates that the method used pre-trained weights on clean ImageNet and applied TTT on the corrupted set. OpenCLIP ViT-L/14 is more robust than earlier models. LQN does better than APM.
P
brigh
cont
defoc
elast
fog
frost
gauss
glass
impul
jpeg
motn
pixel
shot
snow
zoom
Average
Baseline
✓
77.4
71.2
62.3
51.0
66.3
58.4
68.6
59.2
64.9
70.4
70.6
74.7
66.2
54.2
55.2
64.1
TTT-MAE
✓
77.8
71.5
69.4
49.7
69.8
62.7
72.5
66.4
70.0
72.7
72.3
76.2
70.6
58.7
63.6
68.3
OpenCLIP VIT-L/14
✗
76.6
74.4
71.4
53.8
72.0
62.6
67.6
64.0
64.6
73.8
69.0
72.8
66.4
61.8
58.3
66.1
APM
✗
81.1
79.4
76.6
59.4
77.3
68.2
73.1
70.0
70.3
78.6
74.5
77.8
72.0
67.8
64.3
72.4
LQN (Ours)
✗
82.4
81.0
78.4
60.9
78.9
69.8
74.7
71.3
71.7
80.1
75.8
79.2
73.4
69.1
65.8
74.1
Appendix
Table 22: Performance on ImageNet-C, level 2 . The first two rows are from the supplementary materials of ( Gandelsman et al., 2022 ) . A ✓ in column P indicates that the method used pre-trained weights on clean ImageNet and performed TTT on the corrupted version. OpenCLIP ViT-L/14 is generally more robust than earlier models. LQN does better than APM.
P
brigh
cont
defoc
elast
fog
frost
gauss
glass
impul
jpeg
motn
pixel
shot
snow
zoom
Average
Baseline
✓
78.5
74.5
68.1
73.9
70.5
70.6
74.8
68.6
72.3
73.0
75.2
75.9
73.6
69.3
63.7
71.4
TTT-MAE
✓
78.9
74.7
72.5
74.7
72.9
72.2
76.8
72.2
75.5
74.5
75.8
77.0
75.9
71.9
69.3
73.1
OpenCLIP VIT-L/14
✗
77.3
75.4
73.5
73.1
73.5
71.4
71.9
70.2
69.9
75.1
73.7
74.2
71.9
71.2
65.2
71.1
APM
✗
81.6
80.3
78.6
78.0
78.6
76.6
77.2
75.7
75.1
79.6
78.7
79.1
76.9
76.4
70.7
76.0
LQN (Ours)
✗
83.2
82.0
80.3
79.6
80.1
78.0
78.6
77.0
76.5
80.9
80.2
80.5
78.4
77.9
72.3
77.6
Appendix
Table 23: APM and LQN performance on ImageNet-C, level 1 . The first two rows are reproduced from the supplementary materials of ( Gandelsman et al., 2022 ) . A ✓ in column P indicates that the method used pre-trained weights on clean ImageNet and applied TTT on the corrupted set. OpenCLIP VIT-L/14 is generally more robust than earlier models. LQN does better than APM.
Method
orig
gauss
shot
impul
defoc
glass
motn
zoom
snow
frost
fog
brit
contr
elas
pixel
jpeg
Avg
TTT-Online
8.2
25.8
22.6
30.6
14.6
34.4
18.3
17.1
20.0
18.0
16.9
11.2
15.6
21.6
18.1
21.2
19.1
UDA-SS
9.0
28.2
26.5
20.8
15.6
43.7
24.5
23.8
25.0
24.9
17.2
12.7
11.6
22.1
20.3
22.6
21.4
Zeroshot
CLIP ViT-L/14
4.63
35.4
32.3
21.9
19.3
49.7
19.3
17.3
17.0
15.1
21.6
8.4
15.9
34.6
25.0
27.4
24.5
CLIP ViT-L/14 (t)
APM
3.5
21.9
30.1
13.7
15.2
34.1
11.9
11.1
15.0
9.0
13.5
5.8
9.5
23.0
15.8
17.0
14.8
Appendix
Table 24: CIFAR-10-C results at highest severity level of 5. We report Error Rate (%, lower is better). The t- model acts as the teacher for APM and LQN variants. TTT was performed on the test set with randomly initialized weights. APM and LQN weights were reinitialized after each TTT iteration to prevent information leakage.
Figure 8: Qualitative results on LQN. Panoptic segmentation on COCO-Val set. LQN obtains more semantically-detailed masks than the EoMT/APM baselines via test-time-training. Masks visualized after 15 iterations of TTT on both APM/LQN. EoMT is a fully-supervised, fixed baseline. TTT on APM/LQN is performed with EoMT as the teacher. The red regions are the regions where a kind reader can focus on to see the comparison in the prediction quality. Starting from above, LQN easily segments the books in the bookshelf as white region, spoon as a white outline on the plate, distinguishes between the people in the background, and easily partitions chairs into two distinct parts, whereas other EoMT and APM group them together.
Figure 9: Qualitative results on LQN. Panoptic segmentation on COCO-Val set. LQN obtains more semantically detailed masks than the EoMT/APM baselines via test-time training. Masks visualized after 15 iterations of TTT on both APM/LQN. EoMT is a fully-supervised, fixed baseline. TTT on APM/LQN is performed with EoMT as the teacher. The red regions are the regions where a kind reader can focus on to see the comparison in the prediction quality. Starting from above, (first row) LQN gets the passengers in the bus in the background, even though they are partially occluded; (second row) easily segments the building into the blue regions; (third row) partitions buildings into a green colored region, whereas the other two methods don’t detect it as all/term it as background; (fourth row) gets the window in the proper shape.
Figure 10: Qualitative results on LQN. Panoptic segmentation on COCO-Val set. LQN obtains more semantically-detailed masks than the EoMT/APM baselines via test-time-training. Masks visualized after 15 iterations of TTT on both APM/LQN. EoMT is a fully-supervised, fixed baseline. TTT on APM/LQN is performed with EoMT as the teacher. The red regions are the regions where a kind reader can focus on to see the comparison in the prediction quality. Starting from above, (first row) LQN properly segments the suitcases, (second row) Precise boundaries of the cellphone in the person’s hand, even though part of the cellphone is gripped very tightly by the person’s hand. (third row) segments the base of the tree, whereas the other two methods don’t detect it at all (fourth row) is able to understand the fine regions which correspond to the bag/floor. In contrast, EoMT/APM confuse that some part of the purse is also the background.
Vision-language models (VLMs) such as CLIP exhibit remarkable zero-shot capabilities, yet their performance frequently degrades sharply under unexpected test-time distribution shifts. While Test-Time Adaptation (TTA) offers a promising solution, continuously adapting VLMs over an unlabeled test stream presents fundamental challenges. Conventional top-1-centric updates often reinforce errors by corrupting the local semantic geometry among related classes, while iterative adaptation exacerbates progressive bias accumulation, ultimately driving the model toward mode collapse. To overcome these coupled vulnerabilities, we propose Local Margin Restoration (LMR), a lightweight, one-step TTA framework. At the sample level, our Protected Margin Restoration (PMR) objective recovers local semantic geometry by shielding plausible near-top candidates from external hard negatives. Concurrently, to combat stream-level degradation, we introduce a dual-stage stabilization mechanism, featuring an Adaptive Margin (AM) controller and Bias Correction (BC), to dynamically disrupt progressive bias accumulation and prevent mode collapse. Extensive experiments on CIFAR-C, ImageNet-C, and ImageNet variants demonstrate that LMR consistently outperforms state-of-the-art TTA baselines, proving exceptionally robust and efficient even in challenging low-batch test-time regimes. Our code is available at https://github.com/DennisHuangYan/LMR.
Yan Huang, Guowei Wang, Xu Wang +2
Guangzhou University Guangzhou, China · The Second Affiliated Hospital of Guangzhou University of Chinese Medicine Guangzhou, China · Jinan University Zhuhai, China +1
Test-time adaptation (TTA) has emerged as a promising paradigm for vision-language models (VLMs) to bridge the distribution gap between pre-training and test data. Recent works have focused on backpropagation-free TTA methods that rely on cache-based designs, but these introduce two key limitations. First, inference latency increases as the cache grows with the number of classes, leading to inefficiencies in large-scale settings. Second, suboptimal performance occurs when the cache contains insufficient or incorrect samples. In this paper, we present Prototype-Based Test-Time Adaptation (PTA), an efficient and effective TTA paradigm that uses a set of class-specific knowledge prototypes to accumulate knowledge from test samples. Particularly, knowledge prototypes are adaptively weighted based on the zero-shot class confidence of each test sample, incorporating the sample's visual features into the corresponding class-specific prototype. It is worth highlighting that the knowledge from past test samples is integrated and utilized solely in the prototypes, eliminating the overhead of cache population and retrieval that hinders the efficiency of existing TTA methods. This endows PTA with extremely high efficiency while achieving state-of-the-art performance on 15 image recognition benchmarks and 4 robust point cloud analysis benchmarks. For example, PTA improves CLIP's accuracy from 65.64% to 69.38% on 10 cross-domain benchmarks, while retaining 92% of CLIP's inference speed on large-scale ImageNet-1K. In contrast, the cache-based TDA achieves a lower accuracy of 67.97% and operates at only 50% of CLIP's inference speed.
Zhaohong Huang, Yuxin Zhang, Wenjing Liu +2
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University, 361005, P.R. China.
Vision-Language Models (VLMs), such as CLIP, have achieved significant zero-shot performance on downstream tasks with various fine-tuning adaptation methods. However, recent studies have proven that adversarial attacks can significantly degrade the inference ability of VLMs, posing substantial risks to their practical applications. Prevalent test-time adaptation methods typically rely on multi-view augmentation to implement various fine-tuning strategies, which struggle to identify semantic information and are prone to destroying discriminative regions in fine-grained scenarios. To address these limitations, we propose Attention-Guided Test-Time Prompt Tuning (A-TPT), a semantics-preserving method designed for test-time adaptation. We first refine the gradient attention rollout mechanism to identify semantically meaningful regions surviving under adversarial attacks. Furthermore, we leverage them to guide the spatially varying augmentation intensities and multi-view ensemble for prompt tuning and inference. Extensive experiments demonstrate that A-TPT outperforms existing test-time adaptation methods on both adversarial and clean data. Codes are available at https://github.com/SEU-VIPGroup/A-TPT .
Jia-Wei Hai, Yijun Wang, Xiu-Shen Wei
School of Computer Science and Engineering, and Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications, Southeast University, China. · School of Intelligence Science and Engineering, Southeast University, China.