Active test-time adaptation (ATTA) improves robustness under distribution shift by updating a deployed model during inference while selectively querying supervision. However, most existing ATTA methods implicitly assume that supervision can be requested for every incoming test batch, which can incur substantial annotation cost over long test streams. In this work, we introduce \emph{budgeted ATTA} in which labels are available for only a fraction of test batches. This formulation shifts the central challenge from deciding \emph{what} to label within a batch to deciding \emph{when} supervision should be applied over time. To address this challenge, we propose a budget-aware approach \emph{WISE-ATTA} that allocates supervision over the test stream based on lightweight signals computed online, prioritizing periods where supervision is likely to be most useful. When a batch is selected for supervision, we further employ a drift-based sample selection criterion that targets samples exhibiting ongoing, unconverged adaptation dynamics, enabling effective updates from a single labeled example. We evaluate this approach on synthetic corruptions (ImageNet-C) and natural distribution shifts (ImageNet-R/K/A). Across settings, WISE-ATTA achieves competitive or improved performance compared to recent ATTA methods while requiring substantially fewer labels. Overall, we find that the timing of supervision is a key, yet underexplored, aspect of active test-time adaptation. Code: https://github.com/Muhammad-Huzaifaa/WISE-ATTA
Figures & tables
Method
Replay Buffer
Labels / Batch
Batch Sel.
Avg. Err. ↓
CEMA [ 3 ]
✓
teacher
✗
59.2
SimATTA [ 9 ]
✓
3
✗
58.2
HILTTA [ 22 ]
✗
3
✗
58.1
EATTA [ 43 ]
✗
1
✗
58.0
WISE-ATTA (Ours)
✗
≤0.5
✓
57.2
Table 1 : Comparison of ATTA methods on ImageNet-K. WISE-ATTA achieves the lowest average error with fewer labels.
# Labels
Method
ImageNet-R
ImageNet-K
ImageNet-A
Avg. Error
RN50-BN
ViT-B-16
RN50-BN
ViT-B-16
RN50-BN
ViT-B-16
RN50-BN
ViT-B-16
Non-Active
TENT [ 41 ]
57.8
53.4
69.5
65.6
99.9
77.4
75.7
65.5
CoTTA [ 45 ]
57.3
55.4
69.9
98.2
99.8
79.3
75.7
77.6
SAR [ 29 ]
57.2
48.8
68.5
70.4
99.9
74.9
75.2
64.7
ETA [ 28 ]
54.0
48.8
64.3
59.4
99.8
75.9
72.7
61.4
-
CEMA † [ 3 ]
51.4
44.6
65.6
60.0
97.7
72.9
71.6
59.2
Table 3 : Generalization under natural distribution shifts: FTTA error (%) on ImageNet-R/K/A with RN50-BN and ViT-B-16.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Utility Function
ImageNet-R
ImageNet-K
RN50-BN
ViT-B-16
RN50-BN
ViT-B-16
Mean Entropy
52.82
44.81
65.54
58.31
Top- k Entropy
52.77
44.62
65.62
58.85
Low- k Entropy
52.88
44.26
65.46
58.69
Mean Drift
52.85
44.32
65.23
58.78
Top- k Drift
52.25
43.94
65.09
58.44
Appendix
Table 4 : Effect of different batch utility functions (error %, ↓ ).
Setting
Value
ImageNet-R
ImageNet-K
ImageNet-A
Avg.
RN50-BN
ViT-B-16
RN50-BN
ViT-B-16
RN50-BN
ViT-B-16
Slack
w/o slack
52.15
43.13
64.37
58.04
98.56
71.24
64.58
w/ slack
52.03
43.38
64.29
58.11
98.41
71.13
64.56
History window W
1
52.59
44.40
65.01
58.50
97.89
71.84
65.04
50
52.06
44.07
65.16
58.10
98.32
70.81
64.75
100
52.42
43.70
64.91
58.35
98.39
70.57
64.72
Appendix
Table 5 : Sensitivity to batch-selection hyperparameters on ImageNet-R/K/A with RN50-BN and ViT-B-16 (FTTA error %, ↓ ). The default value used throughout the paper is highlighted in gray. WISE-ATTA is robust across all three hyperparameters — the spread in average error is at most 0.5 points across each sweep.
Sample Selection
Batch Selection
ImageNet-R
ImageNet-K
ImageNet-A
Avg. Error
RN50-BN
ViT-B-16
RN50-BN
ViT-B-16
RN50-BN
ViT-B-16
RN50-BN
ViT-B-16
EATTA [ 43 ]
Uniform
53.1
44.7
65.8
59.0
98.1
72.5
72.3
58.7
Random
53.1
44.8
65.6
58.8
98.6
72.4
72.4
58.7
WISE-ATTA
Uniform
52.8
44.9
65.6
58.6
98.2
71.9
72.2
58.5
Random
52.6
44.6
65.4
58.3
98.4
71.8
72.2
58.2
Budget-paced
52.2
44.2
65.3
58.0
98.2
70.6
71.9
57.6
Appendix
Table 7 : Batch selection strategies under a fixed annotation budget. FTTA error (%) on ImageNet-R/K/A with RN50-BN and ViT-B-16.
Setting
Method
ImageNet-R
ImageNet-K
Avg.
Δ Avg.
RN50-BN
ViT-B-16
RN50-BN
ViT-B-16
With update
Uniform
55.0
48.1
68.3
60.8
58.0
–
Random
54.5
48.7
68.2
60.2
57.9
–
WISE-ATTA
53.4
44.7
66.2
58.7
55.7
–
Without update
Uniform
55.0
48.2
68.6
61.0
58.2
+0.2
Random
55.9
48.1
68.4
60.8
58.3
+0.4
Appendix
Table 8 : Effect of skipping unsupervised updates on non-selected batches at label ratio r=0.2 (FTTA error %, ↓ ). With update applies the unsupervised loss on non-selected batches; Without update skips the backward pass entirely. WISE-ATTA degrades the least when unsupervised updates are removed ( +0.1 vs. +0.2 to +0.4 for the baselines), indicating that its gains do not rely on the unsupervised signal from skipped batches.
Batch
Method
ImageNet-R
ImageNet-K
Avg. Error
RN50-BN
ViT-B-16
RN50-BN
ViT-B-16
RN50-BN
ViT-B-16
16
Uniform
56.5
41.8
56.5
41.8
56.5
41.8
Random
57.5
42.1
57.5
42.1
57.5
42.1
WISE-ATTA
55.1
41.5
55.1
41.5
55.1
41.5
32
Uniform
53.0
43.2
65.1
57.5
59.0
50.4
Random
52.7
42.9
65.2
57.6
59.0
50.3
Appendix
Table 9 : Batch-size ablation on ImageNet-R/K with RN50-BN and ViT-B-16 (FTTA error %, ↓ ). WISE-ATTA consistently outperforms uniform and random batch selection across all batch sizes.
Delay
ImageNet-R
ImageNet-K
RN50-BN
ViT-B-16
RN50-BN
ViT-B-16
EATTA
WISE+Unif.
WISE-ATTA
EATTA
WISE+Unif.
WISE-ATTA
EATTA
WISE+Unif.
WISE-ATTA
EATTA
WISE+Unif.
WISE-ATTA
0
54.1
53.1
52.3
47.0
44.5
43.4
65.6
65.6
64.6
60.3
59.0
58.0
50
54.2
54.1
52.9
46.5
46.3
47.2
67.0
66.1
65.3
81.3
68.4
67.9
100
54.2
54.5
54.0
54.6
56.2
57.8
66.7
66.7
65.2
85.8
74.1
71.9
150
54.7
54.8
54.6
70.1
59.8
62.9
66.8
66.7
65.8
88.3
89.3
92.2
Appendix
Table 10 : Effect of annotation delay at a fixed labeling rate of r=0.5 . FTTA error (%, ↓ ) on ImageNet-R and ImageNet-K with RN50-BN and ViT-B-16, as a function of the delay (in batches) between batch selection and label availability. WISE-ATTA is highlighted in gray.
Computer Science and Engineering, University of California San Diego · New Laboratory of Pattern Recognition, Institute of Automation, Chinese Academy of Sciences · College of Computer Science and Electronic Engineering, Hunan University
The University of Texas at Dallas, USA · LIVIA ETS Montreal, ILLS International Laboratory on Learning Systems (ILLS), McGill - ÉTS - MILA - CNRS - Université Paris-Saclay - CentraleSupélec · Tulane University, USA +2