Robust and reliable perception is essential for autonomous robots operating in real-world environments, particularly in long-term missions where environmental conditions may change significantly over time. Although recent advances in Visual Foundation Models (VFMs) have improved open-vocabulary semantic segmentation, these models can still suffer from domain shift, which can significantly degrade performance if they are not adapted to the current environment. Training-free domain adaptation is a relevant paradigm for adaptation, consisting of adjusting the model online using lightweight adapters. Recent approaches apply this on a per-frame basis, which is impractical for deployments on resource-constrained robotic hardware. To tackle this, we propose a multi-signal domain shift detection method for training-free continual test-time adaptation (TF-CTTA) in open-vocabulary segmentation. Our method leverages temporal coherence across consecutive frames by monitoring and combining complementary aspects of domain shift (visual change, adapter mismatch, and semantic drift) to trigger adaptation only when needed. We validate our approach on a benchmark including indoor and outdoor environments and using real robotic data. We demonstrate that our approach maintains segmentation accuracy while substantially reducing adaptations, making training-free adaptation practical and feasible for long-term, real-world robotic deployments.
Fig. 2: General overview of the framework for training-free continual test-time adaptation, incorporating SemLA [ 7 ] and our domain shift detector.
t
Test
Scannet → ACDC
Scannet++ → ACDC
Scannet → MUSES
Scannet++ → MUSES
All
Condition
E1I
E2I
E3I
E4O
E5I
E6I
E7I
E8O
E9I
E10I
E11I
E12O
E13I
E14I
E15I
E16O
Mean
Zero-shot
mIoU
39.0
39.3
39.9
43.1
37.9
36.8
36.3
43.8
38.5
39.0
38.7
35.6
37.5
37.7
36.4
37.6
38.6
NA
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
Entropy [ 8 ]
mIoU
44.6
43.4
41.5
42.0
36.4
35.0
35.0
42.3
38.6
38.3
39.9
33.4
36.7
35.7
36.3
35.4
38.4
NA
2
1
1
0
0
0
0
0
0
0
0
0
0
0
0
0
0.3
TABLE I: Performance comparison in training-free continual test-time adaptation settings
t
Test
Scannet++ → ACDC
All
Condition
E1I
E2I
E3I
E4O
Mean
Zero-shot
mIoU
33.6
36.0
34.0
43.0
36.6
NA
0
0
0
0
0
Multi-signal (ours)
mIoU
36.8
38.9
36.2
51.9
41.0
NA
5
3
3
2
3.2
TABLE II: Performance in out-of-distribution settings
Fig. 3: Our multi-signal domain shift detector on a real-world sequence from a robot sensor (Odin) on our university campus, covering indoor and outdoor environments (2735 frames, 12 confirmed adaptations out of 42 S1 activations). Each image pair covers consecutive adaptation events: the first frame is where adaptation is triggered, the second is from the stable period before the next trigger, showing that the domain remains consistent in between.
Fig. 4: Trade-off between segmentation accuracy and number of adaptations, averaged over 50 sequences, counting 8 domains each. For each point, we report the signals involved and the values of mIoU and NA .
Power mode
Models dim.
Domain nav.
Seg. model
Adapt. cost
FPS (per-frame)
FPS (N A=4 )
High
Large
150ms
230ms
550ms
1.1
2.6
Small
90ms
140ms
550ms
1.3
4.2
Low
Large
350ms
840ms
700ms
0.5
0.8
Small
110ms
490ms
700ms
0.7
1.7
Performance evaluated on a Jetson AGX Orin using two power modalities: 60W (High) and 30W (Low).
TABLE III: Latency of the per-frame vs. ours adaptations (N A=4 )
Test-Time Domain Adaptation (TTDA) aims to adapt Deep Neural Networks to distribution shifts using only streaming, unlabeled test data in real time. Current methods for semantic segmentation tasks suffer from critical limitations. Entropy minimization techniques require costly backpropagation, risking catastrophic forgetting and producing noisy segmentation boundaries. Memory-bank methods, while backpropagation-free, exhibit slow adaptation, requiring numerous samples to converge and struggle to handle continuous domain shifts. We introduce TestMate, a novel, real-time, and backpropagation-free TTDA framework that overcomes these issues. TestMate leverages generalization capability of a lightweight Visual Foundation Model to guide the adaptation. We use a zero-shot instance segmentation YOLOv8-seg based model to generate unlabeled mask proposals for objects and their parts at multiple scales in real time. These proposals are fused with the primary model via a heuristic, size-ordered competitive scheme, where small, high-confidence regions dominate and refine predictions in surrounding larger, less certain areas. This paremeter-free mechanism enables immediate adaptation from the first frame, inherently avoids catastrophic forgetting and effectively preserves fine object details and boundaries, even for small objects. TestMate can be used as a standalone, efficient refinement module or seamlessly integrated into existing TTDA methods to significantly boost their performance. We demonstrate state-of-the-art results across two benchmark datasets, proving TestMate's effectiveness in three distinct adaptation tasks: TTDA, Source-Free Domain Adaptation (SFDA), and online-TTDA. Code is available.
Dimitrios Fotiou, Vasileios Mygdalis, Ioannis Pitas
Test-Time Adaptation (TTA) aims to mitigate distributional shifts between training and test domains during inference time. However, existing TTA methods fall short in the realistic scenario where models face both continually changing domains and the simultaneous emergence of unknown semantic classes, a challenging setting we term Open-set Continual Test-Time Adaptation (OCTTA). The coupling of domain and semantic shifts often collapses the feature space, severely degrading both classification and out-of-distribution detection. To tackle this, we propose DOmain COmpensation (DOCO), a lightweight and effective framework that robustly performs domain adaptation and OOD detection in a synergistic, closed loop. DOCO first performs dynamic, adaptation-conditioned sample splitting to separate likely ID from OOD samples. Then, using only the ID samples, it learns a domain compensation prompt by aligning feature statistics with the source domain, guided by a structural preservation regularizer that prevents semantic distortion. This learned prompt is then propagated to the OOD samples within the same batch, effectively isolating their semantic novelty for more reliable detection. Extensive experiments on multiple challenging benchmarks demonstrate that DOCO outperforms prior CTTA and OSTTA methods, establishing a new state-of-the-art for the demanding OCTTA setting.
Yingkai Yang, Chaoqi Chen, Hui Huang
College of Computer Science and Software Engineering, Shenzhen University
Open-vocabulary semantic segmentation (OVSS) relies on vision-language alignment to recognize arbitrary text-defined categories, yet this alignment is fragile under continual test-time distribution shift. Our diagnostic analysis reveals that entropy minimization drives patch-level class collapse, continual updates erode vision-language alignment, and redundant gradients from low-shift samples waste computation. We propose Diversify, Anchor, and Filter (DAF), a stabilization framework that augments entropy-based adaptation with a marginal diversity loss that resists collapse, a cross-modal anchor consistency loss that constrains feature drift relative to a frozen source model, and feature salience filtering that skips low-value backward passes to offset part of the source-anchor overhead. We evaluate on five datasets spanning natural scenes, autonomous driving, underwater imagery, and remote sensing with their corrupted variants. Across the evaluated continual shifts, DAF remains stable where entropy minimization collapses, improving mIoU by over 8 points on Pascal VOC20-C, over 9 points on LoveDA, and over 3 points on Foggy Cityscapes compared to the source model, and is robust to aggressive adaptation and learning rate choices.
Chandler Timm C. Doloriel, Yunbei Zhang, Sarthak Kumar Maharana +5
Faculty of Science and Technology (REALTEK), Norwegian University of Life Sciences (NMBU) · Tulane University · University of Texas at Dallas