TC-ADA: One-Shot Active Domain Adaptation for Semantic Segmentation
Authors: Weihao Yan, Yeqiang Qian, Yueyuan Li, Tao Li, Chunxiang Wang, Ming Yang
Organizations: School of Automation and Intelligent Sensing, Shanghai Jiao Tong University, Key Laboratory of System Control and Information Processing, Ministry of Education of China, Shanghai, 200240, China
Manual dense annotation remains a major obstacle to deploying semantic segmentation models in new driving environments. Active domain adaptation (ADA) seeks label-efficient transfer by annotating only a selected portion of the target domain. Existing ADA methods commonly implement this process through multiple rounds of acquisition, annotation, and retraining. We study a practical one-shot image-level setting that selects and densely annotates a fixed target subset in a single round, followed by uninterrupted adaptation. Within this setting, we develop Target-Calibrated Active Domain Adaptation (TC-ADA) as a joint design of complete-image acquisition and target-calibrated adaptation. Stage1 uses visual representations from a vision foundation model (VFM) together with semantic predictions from a fixed unsupervised domain adaptation model to select representative and informative target images without target annotations. Stage2 jointly uses labeled source data, labeled target data, and the remaining unlabeled target data, while calibrating source and target supervision under limited target labels. Extensive experiments across five synthetic-to-real and real-to-real driving transfers show consistent improvements over representative ADA baselines. With only 23 to 46 labeled target images on four transfers and 140 on Mapillary, TC-ADA stays within 1.9 mean intersection over union (mIoU) points of target-only full supervision. Code will be available at https://github.com/ywher/TC-ADA.
Figures & tables
Fig. 1: Motivation and protocol overview of TC-ADA. It replaces iterative acquisition–annotation–retraining cycles with a single complete-image annotation round followed by uninterrupted adaptation. TC-ADA reaches 77.66% mean intersection over union (mIoU) using only 25 of 1,600 ACDC target images, close to the 77.72% full-supervision reference.
Fig. 2: Overview of TC-ADA. Stage 1 ranks complete target images using perspective-aware VFM features, density–diversity coverage, and reliable rare-class evidence. Stage 2 combines source and target supervision through two mixed branches, with budget-aware target-supervision attenuation, target-dominant mixing, and target-aware initialization. Distillation and progressive acquisition are optional extensions.
Shift
Transfer
Cls.
#Source
#Target
#Val
S2R
GTA → Cityscapes
19
24,966
2,975
500
SYNTHIA → Cityscapes
16
9,400
2,975
500
R2R
Cityscapes → ACDC
19
2,975
1,600
406
Cityscapes → MUSES
19
2,975
1,500
250
Cityscapes → Mapillary
19
2,975
18,000
2,000
TABLE I: Driving adaptation datasets and split sizes.
Method
Type
Segmentor
Params (M)
Train. (M)
Src. data
Tgt. data
Annotation unit
Annot. rounds
G2C
S2C
C2A
C2Mu
C2Map
Avg.
1/64 (46)
1/64 (46)
1/64 (25)
1/64 (23)
1/128 (140)
Full-data references (HRDA/DINOv3-B)
Target full
Ref.
HRDA/DINOv3-B
92.31
6.64
–
✓
Image
–
82.49
82.81
77.72
79.25
80.04
80.46
Source+target full
Ref.
HRDA/DINOv3-B
92.31
6.64
✓
✓
Image
–
81.77
82.50
78.26
80.72
80.44
80.74
Native-protocol reproductions
AllSpark [ 31 ]
Semi
SegFormer/MiT-B5
84.61
84.61
–
✓
Image
1
62.57
68.11
44.67
31.51
66.76
54.72
TABLE II: Protocol-aware comparison at the lowest target-label budgets (mIoU, %). Parentheses give labeled-image counts for image-level protocols.
(a) GTA → Cityscapes-19 Synthetic-to-real; eval: val
Ratio (images)
1/64 (46)
1/32 (93)
1/16 (186)
1/8 (372)
UniMatchV2 [ 27 ]
72.65
76.95
80.17
80.95
UniMatchV2+Source [ 27 ]
75.86
78.41
79.17
80.10
HALO-Image [ 8 ]
73.00
74.27
76.39
77.47
MADAv2-HRDA [ 9 ]
76.27
77.66
78.82
79.60
TABLE IV: Image-level comparison across target-label budgets.
Layout
CN
1/64
1/16
Avg.
Storage ↓ (MiB)
Global
–
76.95
77.49
77.22
19
+ G2x2
yes
77.19
78.04
77.62
94
+ G4x4
yes
77.37
77.79
77.58
319
+ G4x8
yes
76.93
78.02
77.47
619
+ G8x8
yes
77.58
77.83
77.70
1219
+ G2x4
yes
77.66
78.08
77.87
169
TABLE V: Embedding-layout ablation on C2A with sampling fixed (mIoU, %).
Sampler
C2A
C2Map
1/64
1/16
Avg.
1/128
Random
76.58
77.56
77.07
78.75
Uniform
77.14
77.83
77.48
79.01
Entropy
76.65
77.56
77.11
78.54
K-center [ 36 ]
77.27
77.57
77.42
78.09
Density-k-center [ 33 ]
77.17
77.69
77.43
78.90
TABLE VI: Comparison of Stage 1 sampling strategies (mIoU, %).
Fig. 3: Qualitative comparison at the lowest budgets. From top to bottom: G2C/C2A/C2Mu at 1/64 and C2Map at 1/128. From left to right: input image, ground truth, HALO [ 8 ] , MADAv2 [ 9 ] , UniMatchV2 [ 27 ] , and TC-ADA.
Selection
Class entropy ↑
Rare enrich. ↑
Diversity ↑
1/64
1/16
1/64
1/16
1/64
1/16
Random
1.8351
1.8569
1.43
1.10
0.1352
0.1369
Uniform
1.8932
1.9046
1.93
1.44
0.1220
0.1432
Entropy
1.6971
1.7169
0.69
0.88
0.1309
0.1257
K-center [ 36 ]
1.9846
1.9957
3.11
2.42
0.1715
0.1546
Density-k-center [ 33 ]
1.8816
1.9695
1.67
2.07
0.1731
0.1528
TABLE VII: Semantic coverage and feature diversity of selected ACDC subsets.
Evidence
C2A
G2C
C2Map
C2Mu
Δ DKC
1/64
1/16
1/64
1/16
1/128
1/16
None (DKC)
77.17
77.69
80.21
81.20
78.90
78.28
0.00
Uncertainty
77.13
78.46
79.76
81.25
78.62
77.90
-0.13
Entropy
77.62
77.19
79.99
81.66
78.84
78.29
+0.01
Confidence
76.98
77.13
80.17
81.33
79.29
78.49
+0.07
Rare coverage (R)
77.66
78.08
80.64
81.24
79.20
79.40
+0.52
TABLE VIII: Individual semantic cues added separately to DKC with fixed G2x4 embeddings (mIoU, %).
Configuration
Area
Conf.
Rarity
1/64
1/16
Avg.
DKC baseline
–
–
–
77.17
77.69
77.43
w/o area scaling
–
✓
✓
77.46
77.67
77.57
w/o class confidence
✓
–
✓
76.78
77.93
77.36
w/o rarity weighting
✓
✓
–
77.73
77.82
77.78
Full R
✓
✓
✓
77.66
78.08
77.87
TABLE IX: Internal components of reliable rare-class coverage on C2A (mIoU, %).
Method
Mode
Rounds
G2C 1/64
C2A 1/64
C2Map 1/128
Avg.
mIoU
min
mIoU
min
mIoU
min
mIoU
min
RIPU-Image [ 6 ]
Online
5
74.57
93.7
66.00
53.5
75.43
738.0
72.00
295.1
D2ADA-Image [ 7 ]
Online
5
68.79
1095.5
65.82
232.4
70.85
381.9
68.49
569.9
HALO-Image [ 8 ]
Online
5
73.00
70.2
70.30
36.3
75.90
411.0
73.07
172.5
MADAv2-HRDA [ 9 ]
Offline
1
76.27
165.4
70.42
26.2
76.59
138.9
74.43
110.2
TC-ADA (R-DKC)
Offline
1
80.64
17.7
77.66
8.9
79.20
107.0
79.17
44.5
TABLE X: ADA accuracy (mIoU, %) and acquisition cost under matched image-level budgets.
Setting
Stage 2 design
Init.
C2A
G2C
S→Tu
S→Tl
T-Sup
T-Mix
Atten.
TDM
1/64
1/64
Target full supervision
–
–
–
–
–
–
–
77.72
82.49
Source only
–
–
–
–
–
–
–
70.32
69.79
DACS ( S→Tu )
✓
–
–
–
–
–
scratch
72.92
77.13
+ labeled-target context
✓
✓
–
–
–
–
scratch
76.56
79.37
+ direct T-Sup
✓
✓
✓
–
–
–
scratch
75.87
78.60
TABLE XI: Ablation of target-calibrated adaptation on C2A and G2C (mIoU, %).
Transfer
Budget
VFM encode
UDA stats
R-DKC ranking
Total
G2C
1/64
2.93±0.07
14.57±0.02
0.23±0.00
17.72±0.09
C2A
1/64
1.21±0.16
7.53±0.04
0.13±0.01
8.87±0.19
C2Map
1/128
20.70±0.32
83.14±0.40
3.19±0.11
107.02±0.74
TABLE XII: Breakdown of our one-shot acquisition time (minutes).
Method
Varied factor
G2C 1/64
C2A 1/64
Random
Selection seed
80.07±0.36
76.86±0.33
TC-ADA
Training seed
80.55±0.11
77.67±0.09
TABLE XIII: Three-run reliability at the lowest budget (mIoU, %; mean ± sample standard deviation).
G2C
S2C
Setting
1/64
1/32
1/16
1/64
1/32
1/16
One-shot
80.64
80.70
81.24
81.31
81.75
82.15
Progressive
80.49
81.25
81.67
81.38
81.91
82.36
Δ
-0.15
+0.55
+0.43
+0.07
+0.16
+0.21
Avg. extra time (min)
23.2
27.7
TABLE XIV: Progressive acquisition on synthetic-to-real transfers (mIoU, %).
Model
Full sup
Sup@r
Semi@r
TC-ADA@r
+ KD
GTA → Cityscapes; r=1/8 ; target eval: val
DINOv3-L
83.88
81.88
83.39
83.55
–
DINOv2-B
82.47
79.88
81.54
82.15
–
PIDNet-S
75.59
67.74
71.96
72.80
74.45
Cityscapes → ACDC; r=1/16 ; target eval: val
DINOv3-L
80.27
76.34
78.73
80.44
–
TABLE XV: Generalization across segmentation models (mIoU, %).
Test-Time Domain Adaptation (TTDA) aims to adapt Deep Neural Networks to distribution shifts using only streaming, unlabeled test data in real time. Current methods for semantic segmentation tasks suffer from critical limitations. Entropy minimization techniques require costly backpropagation, risking catastrophic forgetting and producing noisy segmentation boundaries. Memory-bank methods, while backpropagation-free, exhibit slow adaptation, requiring numerous samples to converge and struggle to handle continuous domain shifts. We introduce TestMate, a novel, real-time, and backpropagation-free TTDA framework that overcomes these issues. TestMate leverages generalization capability of a lightweight Visual Foundation Model to guide the adaptation. We use a zero-shot instance segmentation YOLOv8-seg based model to generate unlabeled mask proposals for objects and their parts at multiple scales in real time. These proposals are fused with the primary model via a heuristic, size-ordered competitive scheme, where small, high-confidence regions dominate and refine predictions in surrounding larger, less certain areas. This paremeter-free mechanism enables immediate adaptation from the first frame, inherently avoids catastrophic forgetting and effectively preserves fine object details and boundaries, even for small objects. TestMate can be used as a standalone, efficient refinement module or seamlessly integrated into existing TTDA methods to significantly boost their performance. We demonstrate state-of-the-art results across two benchmark datasets, proving TestMate's effectiveness in three distinct adaptation tasks: TTDA, Source-Free Domain Adaptation (SFDA), and online-TTDA. Code is available.
Dimitrios Fotiou, Vasileios Mygdalis, Ioannis Pitas
Semantic segmentation provides pixel-level scene understanding essential for autonomous driving and fine-grained perception tasks. However, training segmentation models requires costly, labor-intensive annotations on real-world datasets. Unsupervised Domain Adaptation (UDA) addresses this by training models on labeled synthetic data and adapting them to unlabeled real images. While conceptually simple, adaptation is challenging due to the domain gap, i.e., differences in visual appearance and scene structure between synthetic and real data. Prior approaches bridge this gap through pixel-level mixing or feature-level contrastive learning. Yet, these techniques suffer from two major limitations: (1) reliance on high-confidence pseudo-labels restricts learning to a subset of the target domain, and (2) prototype-based contrastive methods initialize class prototypes from source-trained models, yielding biased and unstable anchors during adaptation. To address these issues, we propose a dual-foundation UDA framework that leverages two complementary foundation models. First, we employ the Segment Anything Model (SAM) with superpixel-guided prompting to enable learning from a broader range of target pixels beyond high-confidence predictions. Second, we incorporate DINOv3 to construct stable, domain-invariant class prototypes through its robust representation learning. Our method achieves consistent improvements of +1.3% and +1.4% mIoU over strong UDA baselines on GTA-to-Cityscapes and SYNTHIA-to-Cityscapes, respectively.
Yerin Cheon, Aruna Balasubramanian, Francois Rameau
Stony Brook University, Stony Brook, NY, USA · SUNY Korea, The State University of New York, Korea
Test-time adaptation (TTA) aims to align a model to shifting test domains using only unlabeled streaming data. Most existing methods implicitly infer a single global domain distribution, ignoring the multidimensional and sample-specific nature of real-world domain shifts, leading to fragile adaptation. We propose DOME, an effective domain encoder that explicitly models each sample's domain in a zero-shot manner. DOME leverages vision-language pretraining to extract dense, continuous representations, parameterizes domains as distributional variables, and introduces a momentum-updated sparse domain bank for disentangled supervision. By injecting these explicit domain cues into downstream models, even a basic entropy-minimization TTA strategy achieves state-of-the-art performance across ImageNet-C, ImageNet-R, and ImageNet-Sketch, outperforming complex TTA approaches. Our results demonstrate that robust adaptation stems not from intricate adaptation algorithms, but from explicit, structured domain representation.