Source-Free Universal Domain Adaptation (SF-UniDA) extends Universal Domain Adaptation by removing access to source data at adaptation time while still handling label-set mismatches between domains. Despite growing interest in this setting for image data, no benchmark exists for time series, which are more challenging. We present the first SF-UniDA benchmark on time series. In addition, we provide the first study of pretrained foundation models as feature extractors for time series domain adaptation. In this context, we identify a critical and previously underexplored limitation of all existing SF-UniDA methods: the inference threshold for unknown-sample rejection is highly sensitive. We address this by proposing a plug-in auto-thresholding module that can be integrated into any SF-UniDA method. Experiments on three well-known time series datasets confirm the suitability of this module. They also highlight that foundation models do not systematically outperform classical backbones and that SF-UniDA tailored for time series is yet to be developed.
Figures & tables
Fig. 1 : Overview of the SF-UniDA benchmark for time series. Task-specific backbone architectures and pretrained foundation models are evaluated as feature extractors under the SF-UniDA methods. Both fixed thresholding and auto-thresholding are applied to each SF-UniDA method. The performance is measured with H-score over 3 time series datasets.
Model
Type
Architecture
Param.
Input
Pre-training objective
Task
MOMENT [ 5 ]
Encoder
T5
40M
Patches
Reconstruction
General-purpose
Mantis [ 3 ]
Encoder
ViT
8M
Tokens
Contrastive learning
Classification
Chronos [ 1 ]
Seq2Seq
T5 and LLM
8M
Scalar quantized
Autoregression
Forecasting
Table 1 : Comparison of MOMENT, Mantis, and Chronos time series foundation models. General-purpose tasks include classification, forecasting, anomaly detection, and imputation.
Method
Key idea
Unknown detection
Source-free
Threshold-free (Inference)
UMAD [ 9 ]
Dual-head classifier consistency
Consistency + MixUP
Weak
✓
GLC [ 14 ]
Global One Vs All + local k-NN
One Vs All clustering
Full
✗
GLC++ [ 15 ]
GLC + contrastive loss
One Vs All clustering
Full
✗
LEAD [ 13 ]
Orthogonal feature decomposition
Gaussian Mixture Model
Full
✗
Table 2 : Comparison of SF-UniDA methods. Threshold-free refers to inference only. “Weak” source-free indicates UMAD requires a dual-head architecture and auxiliary orthogonal loss during source pre-training.
Fig. 2 : Source model confidence distributions on target data (no adaptation) for known and unknown samples. Left: CNN / EDF (Time Series). Right: ResNet50 / Office-31 (Images).
Fig. 3 : Threshold sensitivity. Left: Mean H-score as a function of τ for GLC on time series (HAR) and image (Office) data with a CNN backbone and two foundation models, showing that the optimal threshold range is far narrower for time series. Right: Per-scenario H-score as a function of τ for GLC with a CNN backbone on HAR, illustrating that the optimal threshold also varies across adaptation scenarios.
Datasets
Methods
Backbones
Foundation Models
CNN
TFE
S3
TSLANet
Mantis
Moment
Chronos
HAR
UniJDOT
61.0 ∗ (–)
64.6 ∗ (–)
54.1 ∗ (–)
59.7 ∗ (–)
82.5 (–)
30.1 (–)
72.2 (–)
UMAD
52.7 ( ↓ 2.5)
64.6 ( ↑ 29.7)
36.1 ( ↓ 7.5)
43.1 ( ↑ 11.5)
51.2 ( ↑ 16.2)
24.2 ( ↑ 11.4)
46.0 ( ↑ 13.3)
GLC
46.2 ( ↑ 4.2)
42.2 ( ↑ 29.0)
52.3 ( ↑ 11.2)
54.0 ( ↑ 5.8)
53.4 ( ↑ 15.8)
16.7 ( ↑ 16.6)
50.6 ( ↑ 28.9)
GLC++
44.7 ( ↑ 5.4)
36.4 ( ↑ 20.8)
50.7 ( ↑ 5.0)
54.1 ( ↑ 17.8)
59.6 ( ↑ 51.1)
13.6 ( ↑ 10.8)
50.9 ( ↑ 33.8)
LEAD
43.8 ( ↑ 7.2)
42.3 ( ↑ 28.1)
50.4 ( ↑ 4.4)
53.1 ( ↑ 15.5)
57.8 ( ↑ 13.3)
13.7 ( ↑ 11.5)
50.6 ( ↑ 30.7)
Table 3 : H-score (%) with Auto-Thresholding (Yen’s criterion) over 10 seeds. Parentheses show gain vs. fixed thresholding: ( ↑ ) gain, ( ↓ ) loss, ( = ) no change. Bold and italic indicate the best and second-best architectures for each method, respectively. UniJDOT in gray is a UniDA method for which backbone results marked with ∗ are reported directly from [ 11 ] .
Threshold
Backbones
Foundation M.
CNN
TFE
S3Layer
Mantis
Chronos
yen
46.2
42.2
52.3
53.4
50.6
otsu
44.5
42.9
49.8
45.7
42.9
li
42.0
41.3
55.3
48.0
40.3
Table 4 : H-score (%) over 10 seeds for GLC on HAR across backbones and thresholding methods
The goal of source-free domain adaptation (SFDA) for time-series data is to transfer knowledge from a pre-trained source model to an unlabeled target domain without requiring access to source data, while addressing feature shift and temporal drift inherent in the signals. Although existing approaches have explored temporal dynamics in unsupervised source-free adaptation, they largely overlook spectral shifts in time-series data. Towards this end, we propose a novel approach termed temporal-Spectral Alignment with Frequency Adaptation (SAFA) for source-free time-series domain adaptation. Specifically, we first model the source domain at multiple scales by jointly capturing temporal dependencies and spectral characteristics. To adapt time-series data in the target domain, we introduce a trainable frequency adaptation module that modulates the phase and amplitude of target signals in the frequency domain to align them with the source distribution. Extensive experiments on multiple benchmark datasets demonstrate the efficacy and robustness of SAFA.
Source-free universal domain adaptation (SF-UniDA) adapts a pre-trained source model to an unlabeled target domain under both covariate and label shifts, without access to source data. However, existing SF-UniDA methods rely on inefficient techniques such as threshold tuning and clustering. Foundation models (FMs), known for their generalization and zero-shot capabilities, remain underexplored in SF-UniDA. In this paper, we propose a framework that leverages foundation models (LFM) for SF-UniDA. We use a vision-language model (VLM) to compute similarities between target samples and text labels, including those for unknown classes generated by prompting a large language model. The label shift type is determined by analyzing the coefficient of variation of a similarity-based sample-level score. Unknown samples are identified using a binary Gaussian mixture model fitted to another similarity-based metric. Under a consensus strategy, the pseudo-labels generated by the VLM are refined by the target model initialized with the pre-trained source model, integrating knowledge from both the source domain and foundation models. Finally, these refined pseudo-labels are used to train the target model. Extensive experiments across all possible label shifts and multiple benchmarks demonstrate the effectiveness and superiority of our proposed LFM framework. Our code is available at https://github.com/iamjingli/LFM.
Jing Li, Pan Liu, Meng Zhao +7
School of Computer Science and Engineering, Tianjin University of Technology, Tianjin 300384, China · Engineering Research Center of Learning-Based Intelligent System, Ministry of Education of the People’s Republic of China, Tianjin University of Technology, Tianjin 300384, China · School of Artificial Intelligence, Tianjin University, Tianjin 300350, China +1
Unsupervised time-series domain adaptation (DA) addresses the challenge of transferring a classifier from a labeled source domain to an unlabeled target domain under distribution shifts induced by different users, sensors, devices, acquisition conditions, or temporal dynamics. Existing methods typically mitigate this shift by aligning marginal feature distributions through adversarial training, optimal transport, or moment-based discrepancies. In this paper, we propose Class-Conditional Path Distribution Alignment (CPDA), a non-adversarial discrepancy-based framework that aligns source and target class-conditional latent path distributions rather than only global feature marginals. CPDA introduces a composite signature-spectral kernel that jointly captures pooled semantic features, temporal path structure, frequency-domain information, and low-rank path-signature dynamics, while using source labels and target soft pseudo-labels to perform class-preserving alignment. We further provide a theoretical analysis showing that CPDA defines a valid kernel discrepancy, admits existing moment-matching methods as restricted cases, and yields a class-conditional target-risk bound. Extensive experiments with CNN, ResNet18, and TCN backbones on 13 different time-series DA benchmarks demonstrate the effectiveness of CPDA against 30 discrepancy, adversarial, and pseudo-labeling baselines.
Felix Ott, Christopher Mutschler
Fraunhofer Institute for Integrated Circuits IIS, 90411 Nürnberg, Germany · Machine Learning and Positioning Systems Department, University of Technology Nürnberg (UTN), 90461 Nürnberg