Self-supervised pre-training of time series models is currently dominated by next-token prediction and reconstruction objectives. In continuous-valued domains, these paradigms often waste model capacity on high-frequency, point-wise noise at the expense of learning invariant structure. While invariance-based self-distillation has proven highly effective in computer vision, its application to temporal data remains largely underexplored. Effectively adapting such methods to time series requires carefully designed augmentations: spatial operations like cropping can shift the timing of repeating cycles or distort the signal, while basic jittering may provide limited variation. We introduce Wavelet-based self-distillation for time series (WinoTS), an invariance-based pre-training paradigm designed specifically for temporal signals. At its core, WinoTS leverages time-frequency augmentations to construct multi-scale structural views without distorting underlying signal dynamics. Across extensive evaluations, WinoTS outperforms state-of-the-art baselines in long-term forecasting, cross-domain zero-shot transfer, and unsupervised anomaly detection. Notably, linear probing on frozen WinoTS representations frequently surpasses fully supervised models trained from scratch. Systematic ablations demonstrate that WinoTS is a flexible, architecture-agnostic framework yielding gains across time series backbones, and establish that time-frequency transformations provide a principled alternative to vision-style spatial augmentations.
Figures & tables
Figure 1 : Architectural pipeline of WinoTS . The wavelet generator produces Q easy and V hard views. The teacher processes only easy views, while the student processes both. Each teacher easy view is matched to all student views except its identical easy counterpart; every hard view is matched to all teacher views.
Ours
Self-Supervised
Supervised
Dataset
WinoTS
WinoTS -LP
TimeSiam
TS2Vec
TimeMixer
TimeBase
SparseTSF
PatchTST
DLinear
iTransformer
FEDformer
TimesNet
Autoformer
(ours)
(ours)
Dong et al. [ 2024 ]
Yue et al. [ 2022 ]
Wang et al. [ 2024 ]
Huang et al. [ 2025 ]
Lin et al. [ 2024 ]
Nie et al. [ 2023 ]
Zeng et al. [ 2023 ]
Liu et al. [ 2023 ]
Zhou et al. [ 2022 ]
Wu et al. [ 2023 ]
Wu et al. [ 2021 ]
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
ETTh1
0.411
0.423
0.416
0.423
0.426
0.443
0.817
0.669
0.436
0.444
0.409
0.415
0.423
0.430
0.437
0.438
0.465
0.469
0.448
0.446
0.484
0.490
0.471
0.465
0.568
0.518
ETTh2
0.347
0.386
0.347
0.385
0.362
0.402
1.957
1.108
0.349
0.391
0.358
0.399
0.372
0.406
0.380
0.407
0.458
0.462
0.379
0.397
0.419
0.460
0.408
0.405
0.526
0.505
ETTm1
0.346
0.379
0.362
0.386
0.348
0.384
0.670
0.583
0.362
0.388
0.367
0.386
0.368
0.394
0.393
0.401
0.377
0.396
0.408
0.410
0.441
0.459
0.403
0.412
0.665
0.545
Table 1 : Average in-domain forecasting over horizons {96,192,336,720} with context length 336 (lower is better). Red and blue indicate best and second-best results. WinoTS uses full fine-tuning; WinoTS -LP freezes the backbone and trains only the forecasting head. TimeSiam and TS2Vec are self-supervised; the remaining baselines are supervised.
Backbone Family
Encoder Backbone
Avg. MSE / MAE Reduction (%)
MLP
TimeMixer
+3.2% / +2.7%
Transformer
PatchTST
+3.5% / +2.4%
iTransformer
+2.6% / +2.9%
TCN
TS2Vec
+58.0% / +44.0%
Table 2 : Relative MSE/MAE reduction of WinoTS fine-tuning compared to training the same backbone architectures from scratch under their optimal supervised configurations. Results are averaged across five datasets (ETTh1, ETTh2, ETTm1, ETTm2, Weather) and four forecasting horizons ( H∈{96,192,336,720} ). Positive values indicate performance gains.
Transfer
WinoTS
WinoTS -LP
TimeMixer
TimeBase
SparseTSF
PatchTST
Autoformer
FEDformer
iTransformer
TimesNet
(Ours)
(Ours)
Wang et al. [ 2024 ]
Huang et al. [ 2025 ]
Lin et al. [ 2024 ]
Nie et al. [ 2023 ]
Wu et al. [ 2021 ]
Zhou et al. [ 2022 ]
Liu et al. [ 2023 ]
Wu et al. [ 2023 ]
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
ETTh1 → ETTh2
0.3496
0.3888
0.3520
0.3902
0.3778
0.4000
0.3560
0.3991
0.3729
0.4075
0.3790
0.4020
0.4721
0.4810
0.4573
0.4696
0.3752
0.3991
0.4229
0.4307
ETTh1 → ETTm1
0.7027
0.5461
0.7032
0.5457
0.7866
0.5738
0.7420
0.5647
0.8111
0.5644
0.8000
0.5890
0.7730
0.5870
0.7627
0.5802
0.8307
0.5851
0.9413
0.6222
ETTh1 → ETTm2
0.2951
0.3481
0.2955
0.3482
0.3143
0.3559
0.3060
0.3611
0.3096
0.3616
0.3140
0.3570
0.3650
0.4070
0.3532
0.3902
0.3216
0.3632
0.3524
0.3837
ETTh2 → ETTh1
0.4473
0.4512
0.4513
0.4536
0.6379
0.5457
0.4150
0.4156
0.5039
0.4739
0.6410
0.5490
0.7140
0.5820
0.6856
0.5760
0.6673
0.5676
0.8393
0.6436
Table 3 : Cross-domain zero-shot forecasting transfer on ETT datasets (lower is better). Best and second-best results are highlighted in red and blue , respectively. Each source → target row denotes pre-training on the source dataset followed by direct evaluation on the target dataset, without target-domain parameter updates.
Dataset
WinoTS
iTransformer
DLinear
Autoformer
TimesNet
FEDformer
Crossformer
Reformer
(Ours)
Liu et al. [ 2023 ]
Zeng et al. [ 2023 ]
Wu et al. [ 2021 ]
Wu et al. [ 2023 ]
Zhou et al. [ 2022 ]
Zhang and Yan [ 2023 ]
Kitaev et al. (2020)
P
R
F1
P
R
F1
P
R
F1
P
R
F1
P
R
F1
P
R
F1
P
R
F1
P
R
F1
SMD
84.04
77.34
80.55
68.22
63.79
65.93
70.13
69.34
69.73
67.91
41.63
51.62
79.28
54.20
64.39
60.86
52.23
56.22
62.99
62.32
62.65
64.03
61.56
62.77
MSL
88.47
69.39
77.78
53.62
13.79
21.94
69.45
25.38
37.18
82.14
41.25
54.92
62.38
19.03
29.17
81.72
39.34
53.11
77.20
27.22
40.25
78.22
35.51
48.85
SMAP
92.36
64.56
76.00
57.84
8.45
14.75
69.09
14.14
23.47
73.57
20.13
31.61
68.28
13.39
22.38
70.91
16.94
27.35
70.42
16.24
26.39
77.41
22.56
34.94
SWaT
34.72
8.59
13.78
6.67
1.34
2.23
5.96
1.19
1.98
9.83
1.95
3.26
4.17
0.84
1.39
10.07
2.00
3.34
28.92
7.37
11.75
9.76
1.93
3.23
Table 4 : Anomaly detection results. We report point-adjusted precision (P), recall (R), and F1-score (%). Best and second-best results are shown in red and blue , respectively.
wavelets
jitter
jitter+crop
gaussian+crop
Dataset
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
ETTh1
0.416
0.423
0.417
0.426
0.422
0.426
0.424
0.428
ETTh2
0.347
0.385
0.349
0.387
0.349
0.388
0.360
0.392
ETTm1
0.364
0.386
0.397
0.414
0.438
0.443
0.398
0.418
ETTm2
0.252
0.308
0.247
0.308
0.253
0.312
0.274
0.328
Weather
0.238
0.273
0.242
0.278
0.303
0.320
0.263
0.294
Table 5 : Wavelet versus vision-like augmentations. We compare wavelet-based augmentations to jitter and cropping. Best results in red , second-best in blue .
Dataset
Invariance- based (Ours)
Invariance- based+MAE
MAE
NTP
JEPA
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
ETTh1
0.416
0.423
0.418
0.424
0.465
0.465
0.428
0.435
0.442
0.447
ETTh2
0.347
0.385
0.355
0.389
0.449
0.457
0.407
0.427
0.390
0.426
ETTm1
0.364
0.386
0.370
0.388
0.349
0.387
0.346
0.381
0.375
0.393
ETTm2
0.252
0.308
0.258
0.313
0.294
0.344
0.263
0.320
0.274
0.332
Weather
0.238
0.273
0.241
0.274
0.266
0.294
0.226
0.265
0.229
0.266
Table 6 : Ablation of WinoTS ’s objective (Invariance-based self distillation). In-domain forecasting MSE/MAE averaged over horizons {96,192,336,720} . Lower is better; best in red , second-best in blue .
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A1 : Qualitative comparison of wavelet-domain and vision-style view generation on a standardized periodic input window. The left panel shows the original input signal. In the middle panel, the WINO-TS easy teacher view, produced through wavelet-detail soft thresholding, and the hard student view, produced through wavelet-detail perturbation, retain the dominant period, temporal ordering, and full sequence support. Their differences are concentrated primarily in localized amplitude and high-frequency fluctuations. In the right panel, a representative vision-style augmentation pipeline based on temporal cropping, contrast modification, jitter, and additive noise alters the support and point-wise alignment of the two views, making their periodic correspondence less direct. The example illustrates the intended inductive bias of WINO-TS: preserving coarse temporal organization and global context while varying localized fine-scale content. It is provided as a qualitative sanity check rather than as general evidence of semantic preservation.
Wavelet
Family
VM
L
Primary property
sym4
Symlet
4
8
Reduced phase asymmetry
sym6
Symlet
6
12
Reduced phase asymmetry
sym8
Symlet
8
16
Reduced phase asymmetry
db4
Daubechies
4
8
Minimum-phase design
db6
Daubechies
6
12
Minimum-phase design
coif2
Coiflet
4
12
Moment constraints on ϕ and ψ
Appendix
Table A1 : The WINO-TS wavelet pool P . VM denotes the number of vanishing moments of the wavelet function ψ , and L denotes filter length in samples.
Hyperparameter
Value
Optimizer
AdamW
Base learning rate
5×10−4 (linear batch scaling, /256 )
LR schedule
3 -epoch warmup → cosine to 1×10−6
Weight decay
cosine 0.04→0.1
Gradient clipping
max-norm 3.0
Epochs
80
Appendix
Table A2 : WINO-TS pre-training hyperparameters.
Ours (synth, FT)
Ours (synth, LP)
Survey (synth, LP)
Dataset
MSE
MAE
MSE
MAE
MSE
MAE
ETTh1
0.417
0.425
0.521
0.489
0.438
0.446
ETTh2
0.365
0.402
0.410
0.428
0.362
0.401
ETTm1
0.351
0.378
0.361
0.389
0.357
0.384
ETTm2
0.250
0.310
0.259
0.319
0.253
0.311
Weather
0.227
0.261
0.242
0.276
0.235
0.272
Appendix
Table A3 : Forecasting (MSE/MAE, avg. over horizons {96,192,336,720} ) with synthetic pre-training. Ours (FT) = best fine-tuned DINO config per dataset; Ours (LP) and survey Major and others [2026] are linear-probe. Each MSE/MAE pair is from a single run. Lower is better.
K=512
K=1024 (Ours)
K=2048
K=8192
Dataset
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
ETTh1
0.430
0.431
0.416
0.423
0.431
0.433
0.435
0.434
ETTh2
0.369
0.395
0.347
0.385
0.377
0.399
0.378
0.400
ETTm1
0.364
0.387
0.364
0.386
0.359
0.383
0.357
0.384
ETTm2
0.248
0.311
0.252
0.308
0.248
0.310
0.249
0.309
Weather
0.238
0.273
0.238
0.273
0.269
0.298
0.241
0.275
Appendix
Table A4 : Effect of the DINO head output dimension K (out_dim) on in-domain forecasting (linear probe, context 336, MSE/MAE averaged over horizons {96,192,336,720} ). Ours uses K=1024 ; best per row in bold .
Ours
Self-Supervised
Supervised
Dataset
H
WINO-TS
WINO- TS-LP
TimeSiam
TS2Vec
TimeMixer
TimeBase
SparseTSF
PatchTST
DLinear
iTransformer
FEDformer
TimesNet
Autoformer
(Ours)
(Ours)
Dong et al. [ 2024 ]
Yue et al. [ 2022 ]
Wang et al. [ 2024 ]
Huang et al. [ 2025 ]
Lin et al. [ 2024 ]
Nie et al. [ 2023 ]
Zeng et al. [ 2023 ]
Liu et al. [ 2023 ]
Zhou et al. [ 2022 ]
Wu et al. [ 2023 ]
Wu et al. [ 2021 ]
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
ETTh1
96
0.369
0.394
0.374
0.396
0.376
0.405
0.647
0.578
0.379
0.402
0.370
0.385
0.374
0.394
0.385
0.401
0.380
0.402
0.385
0.403
0.388
0.430
0.402
0.422
0.501
0.475
192
0.405
0.413
0.410
0.417
0.414
0.429
0.739
0.628
0.426
0.433
0.401
0.406
0.419
0.422
0.427
0.427
0.412
0.422
0.440
0.436
0.463
0.475
0.472
0.467
0.526
0.495
336
0.425
0.430
0.429
0.429
0.439
0.450
0.852
0.688
0.431
0.439
0.418
0.417
0.432
0.431
0.465
0.450
0.496
0.490
0.479
0.459
0.492
0.488
0.513
0.484
0.595
0.543
Appendix
Table A5 : In-domain forecasting performance per prediction length {96,192,336,720} (lower is better). All methods use the same input context length of 336. Best and second-best per row are shown in red and blue . Ours : WINO-TS (full fine-tune) and WINO-TS-LP (frozen backbone, trained head). Self-supervised : TimeSiam and TS2Vec. Supervised : end-to-end baselines.
In-domain WINO-TS-FT
Synthetic WINO-TS-FT
Synthetic DINO+MAE
Dataset
MSE
MAE
MSE
MAE
MSE
MAE
ETTh1
0.411
0.423
0.417
0.425
0.418
0.423
ETTh2
0.347
0.386
0.365
0.402
0.359
0.404
ETTm1
0.346
0.379
0.350
0.378
0.345
0.377
ETTm2
0.250
0.308
0.250
0.309
0.249
0.308
Weather
0.224
0.262
0.226
0.261
0.227
0.261
Appendix
Table A6 : Effect of the pre-training data source and objective. Each entry reports MSE/MAE; lower is better. Best values for each dataset and metric are shown in bold .
Dataset
Reported samples
EthanolConcentration
261
SpokenArabicDigits
6599
FaceDetection
5890
JapaneseVowels
270
SelfRegulationSCP1
268
SelfRegulationSCP2
200
Appendix
Table A7 : Classification datasets and reported sample counts.
Dataset
WINO-TS
iTransformer
(Ours)
Liu et al. [2023]
EthanolConcentration
0.2970
0.2810
SpokenArabicDigits
0.9900
0.9827
FaceDetection
0.6700
0.6592
JapaneseVowels
0.9570
0.9811
SelfRegulationSCP1
0.8770
0.9113
Appendix
Table A8 : Classification accuracy after full end-to-end fine-tuning initialized from a WINO-TS-pretrained backbone. Both the backbone and classification head are updated. Higher is better; best results are in bold .
Dataset
TimeMixer(MLP)
PatchTST
iTransformer
TS2Vec(TCN)
Wang et al. [ 2024 ]
Nie et al. [ 2023 ]
Liu et al. [ 2023 ]
Yue et al. [ 2022 ]
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
ETTh1
0.416
0.423
0.422
0.432
0.656
0.558
0.541
0.505
ETTh2
0.347
0.385
0.357
0.392
0.427
0.447
0.388
0.423
ETTm1
0.364
0.386
0.370
0.388
0.427
0.423
0.431
0.432
ETTm2
0.252
0.308
0.252
0.309
0.290
0.344
0.286
0.341
Appendix
Table A9 : Backbone ablation for WINO-TS (DINO pre-training, full augmentation family). In-domain forecasting using linear probe, with MSE/MAE averaged over horizons {96,192,336,720} . The DINO objective and augmentation are held fixed — all backbones use the full wavelet-augmentation family (pool { sym4, sym6, sym8, db4, db6, coif2 } ) — while only the encoder backbone is swapped. Lower is better; best in red , second-best in blue .
Daub.
Zero-out
Full pool
Dataset
MSE
MAE
MSE
MAE
MSE
MAE
ETTh1
0.416
0.423
0.415
0.422
0.416
0.423
ETTh2
0.347
0.385
0.348
0.387
0.347
0.385
ETTm1
0.363
0.386
0.359
0.385
0.364
0.386
ETTm2
0.253
0.308
0.257
0.312
0.252
0.308
Weather
0.237
0.272
0.238
0.273
0.238
0.273
Appendix
Table A10 : Augmentation-family ablation for WINO-TS with linear probe on the DINO objective. Each cell reports MSE/MAE averaged over horizons {96,192,336,720} . Lower is better; best results are in bold .
Dataset
WINO-TS
WINO- TS+MAE
MAE
NTP
JEPA
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
ETTh1
0.416
0.423
0.418
0.424
0.465
0.465
0.428
0.435
0.442
0.447
ETTh2
0.347
0.385
0.355
0.389
0.449
0.457
0.407
0.427
0.390
0.426
ETTm1
0.364
0.386
0.370
0.388
0.349
0.387
0.346
0.381
0.375
0.393
ETTm2
0.252
0.308
0.258
0.313
0.294
0.344
0.263
0.320
0.274
0.332
Weather
0.238
0.273
0.241
0.274
0.266
0.294
0.226
0.265
0.229
0.266
Appendix
Table A11 : Pre-training objective ablation (WINO-TS, TimeMixer backbone, full augmentation pool, Linear probe). In-domain forecasting MSE/MAE averaged over horizons {96,192,336,720} . DINO : pure self-distillation — the student matches the teacher’s centered/sharpened prototype distribution across augmented views, with no reconstruction ( ϕ=1 ). DINO+MAE : adds a masked-autoencoding term in which the student reconstructs masked patches against the raw signal, blended as ϕDINO+(1−ϕ)MAE ( ϕ=0.6 ). MAE , NTP and JEPA are non-distillation baselines: pure masked autoencoding (reconstruct masked patches, no teacher) , next-token prediction (autoregressive forecasting-style pre-training) and joint-embedding predictive architecture respectively. Lower is better; best in red , second-best in blue .
Ours
Wavelet pool
Transform
DWT/Mixed
db
sym
SWT
MODWT
Dataset
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
ETTh1
0.416
0.423
0.416
0.423
0.416
0.423
0.415
0.422
0.417
0.423
ETTh2
0.347
0.385
0.347
0.386
0.347
0.386
0.347
0.386
0.346
0.385
ETTm1
0.364
0.386
0.364
0.386
0.365
0.387
0.360
0.385
0.364
0.386
ETTm2
0.252
0.308
0.253
0.309
0.258
0.312
0.257
0.312
0.254
0.309
Appendix
Table A12 : Wavelet augmentation ablation for WINO-TS (DINO pre-training, TimeMixer backbone, Linear probe). In-domain forecasting MSE/MAE averaged over horizons {96,192,336,720} . The first column is the default WINO-TS (DWT with the Mixed pool { sym4,sym6,sym8,db4,db6,coif2 } ) and is the shared reference. Wavelet pool fixes the DWT and sweeps the pool (db ={ db4,db6,db8 } ; sym ={ sym4,sym6,sym8 } ). Transform fixes the Mixed pool and swaps the transform (SWT / MODWT). Lower is better; best in red , second-best in blue .
Dataset
ρ=0.2
ρ=0.4
ρ=0.6
ρ=0.8
ρ=1.0
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
ETTh1
0.428
0.430
0.464
0.458
0.416
0.423
0.419
0.426
0.465
0.458
ETTh2
0.348
0.387
0.370
0.396
0.347
0.385
0.370
0.396
0.370
0.396
ETTm1
0.371
0.391
0.383
0.403
0.364
0.386
0.382
0.404
0.384
0.405
ETTm2
0.250
0.310
0.249
0.310
0.252
0.308
0.250
0.310
0.250
0.310
Weather
0.243
0.273
0.244
0.273
0.238
0.273
0.243
0.278
0.243
0.278
Appendix
Table A13 : Shrinkage-strength ( ρ ) ablation for WINO-TS (DINO pre-training, TimeMixer backbone, DWT with the Mixed pool, Linear probe). We vary the soft-threshold shrinkage ratio ρ that sets the per-level denoising threshold of the teacher’s easy view: detail coefficients below ρ⋅max(∣detail∣) are shrunk toward zero, so small ρ preserves high-frequency detail (weak invariance) while large ρ enforces aggressive low-pass smoothing (strong invariance); ρ=0.6 is the WINO-TS default. All other factors—transform, wavelet pool, objective, and backbone—are held fixed. Each cell reports MSE/MAE averaged over horizons {96,192,336,720} ; lower is better, best per row in bold .
J=2
J=3
J=4
Dataset
MSE
MAE
MSE
MAE
MSE
MAE
ETTh1
0.426
0.429
0.416
0.423
0.438
0.440
ETTh2
0.370
0.395
0.347
0.385
0.361
0.391
ETTm1
0.361
0.386
0.364
0.386
0.362
0.385
ETTm2
0.255
0.312
0.252
0.308
0.248
0.308
Weather
0.238
0.272
0.238
0.273
0.239
0.273
Appendix
Table A14 : Wavelet-depth ablation using linear probe. Each cell reports MSE/MAE averaged over horizons {96,192,336,720} . Lower is better; best results are in bold .
Fixed basis
Shared sampled
Independent sampled
Dataset
MSE
MAE
MSE
MAE
MSE
MAE
ETTh1
0.429
0.431
0.420
0.429
0.416
0.423
ETTh2
0.369
0.394
0.363
0.391
0.347
0.385
ETTm1
0.362
0.387
0.363
0.387
0.364
0.386
ETTm2
0.247
0.308
0.275
0.328
0.252
0.308
Weather
0.243
0.279
0.243
0.279
0.238
0.273
Appendix
Table A15 : Wavelet-pool and basis-sampling ablation (Linear probe). Fixed uses a single wavelet basis for both views. Shared sampled draws one wavelet from P and uses it for both easy and hard views. Independent sampled is the WINO-TS default, drawing weasy and whard independently. Each cell reports MSE/MAE averaged over horizons {96,192,336,720} . Lower is better; best results are in bold .
Default
Same-view
Symmetric
Dataset
MSE
MAE
MSE
MAE
MSE
MAE
ETTh1
0.416
0.423
0.435
0.434
0.420
0.430
ETTh2
0.347
0.385
0.363
0.391
0.364
0.392
ETTm1
0.364
0.386
0.361
0.385
0.362
0.385
ETTm2
0.252
0.308
0.248
0.309
0.248
0.307
Weather
0.238
0.273
0.242
0.279
0.239
0.275
Appendix
Table A16 : View-pairing ablation using linear probing. Default denotes the WINO-TS pairing rule: the teacher processes easy views, the student processes both easy and hard views, and every student view is matched to every teacher easy view except the identical easy-view index. Same-view additionally includes the matching easy–easy pairs. Symmetric allows both networks to process easy and hard views. Each cell reports MSE/MAE averaged over horizons {96,192,336,720} . Lower is better; best results are in bold .
Dataset
MSE
MAE
ETTh1
0.417±0.001
0.424±0.001
ETTh2
0.350±0.004
0.387±0.002
ETTm1
0.361±0.003
0.385±0.001
ETTm2
0.248±0.002
0.308±0.002
Weather
0.238±0.001
0.272±0.001
Appendix
Table A17 : Seed robustness of WINO-TS LP (DINO pre-training, TimeMixer backbone, linear probe). In-domain forecasting MSE/MAE averaged over horizons {96,192,336,720} , reported as mean ± std over 5 seeds ( {42,777,1773,2024,3407} ). Lower is better.
Dataset
WINO-TS ( Q=1,V=1 )
Q=2,V=2
Q=2,V=4
Q=2,V=6
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
ETTh1
0.416
0.423
0.433
0.433
0.429
0.431
0.429
0.430
ETTh2
0.347
0.385
0.363
0.391
0.362
0.391
0.362
0.390
ETTm1
0.364
0.386
0.361
0.385
0.362
0.385
0.362
0.386
ETTm2
0.252
0.308
0.247
0.308
0.248
0.309
0.247
0.308
Weather
0.238
0.273
0.240
0.276
0.239
0.275
0.239
0.275
Appendix
Table A18 : Multi-view ablation with a TimeMixer backbone and linear probing. The default uses Q=1 easy view and V=1 hard view; the alternatives use Q=2 and V∈{2,4,6} . Each view uses an independently sampled wavelet basis and, for hard views, independently sampled noise. Results are MSE/MAE averaged over horizons {96,192,336,720} ; lower is better and best results are in bold .
Self-supervised learning (SSL) assumes that solving pretext tasks on unlabeled data yields representations that transfer effectively across downstream applications via linear probing or fine-tuning. While this paradigm has driven major progress in vision and language, its benefits for time series remain under-investigated and often confounded by inconsistent experimental controls. To address this gap, we benchmark seven representative methods from five key SSL paradigms across anomaly detection, classification, and forecasting under parameter- and data-matched budgets. We find that transfer efficacy depends heavily on the downstream task. SSL yields substantial gains in anomaly detection and provides effective initializations for classification under fine-tuning, but offers limited to no advantage over non-pre-trained controls in forecasting. Furthermore, linear probing does not reliably predict fine-tuning performance, synthetic pre-training is often competitive with real-world corpora, and scaling encoder depth degrades forecasting accuracy. We synthesize these empirical results into practical evaluation and development guidelines for time-series SSL.
Noam Major, Kathy Razmadze, Yoli Shavit
Faculty of Engineering, Bar-Ilan University, Israel
Time series foundation models rely on large-scale pretraining over diverse datasets across domains, yet their heterogeneity in temporal patterns could hinder the effectiveness of training and learning transferable time series representations. Inspired a fundamental concept, normalized power spectral density (PSD) in signal processing, we assume harmonizing datasets via PSDs in the spectral domain could reduce mismatches and enhance pretraining. We then go beyond the direct intractable minimization optimization and innovatively reformulate it as a principled harmonization approach. Specifically, we propose Harmonizer, a module that reshapes spectral structures and implicitly harmonizing PSDs across datasets, which theoretically corresponds to a shared reparameterization of second-order temporal correlations. Our theoretical analysis further reveals token interactions with Harmonizer can be efficiently mediated by a compact set of resonators, motivating a HarmonicAttention design that performs self-attention in a low-dimensional interaction space. Then, we propose Olivia, a novel time series foundation model built upon these harmonization mechanisms. Extensive experiments on two large-scale benchmarks (TSLib and GIFT-Eval) and extra 6 datasets from GluonTS, demonstrate Olivia consistently achieves state-of-the-art performance under zero-shot, few-shot, and full-shot forecasting scenarios. Our code is available at https://github.com/TSTS13/Olivia.
Jingru Fei, Kun Yi, Alex Xing Wang +3
Beijing Institute of Technology · North China Institute of Computing Technology · State Information Center +4
Self-supervised learning for time-series representation aims to reduce reliance on labeled data while maintaining strong downstream performance, yet many existing approaches incur high computational costs or rely on assumptions that do not hold across diverse temporal dynamics. In this work, we introduce Divide and Contrast (Di-COT), an unsupervised framework that avoids data augmentation and multiple encoder passes by contrasting informative substructures within a window rather than individual timesteps. Di-COT stochastically partitions each window into a small number of overlapping sub-blocks per iteration, enabling efficient and meaningful contrast while mitigating false positives during temporal transitions. To further improve scalability, we adopt a contrastive objective whose computation depends on the batch size and the number of sub-blocks, making loss computation independent of sequence length. Extensive experiments on six large-scale real-world datasets, as well as the UCR and UEA benchmarks, demonstrate that Di-COT learns semantically structured and transferable representations, achieving state-of-the-art performance on classification, clustering, kNN, and cross-dataset transfer, while substantially reducing training time. The source code is publicly available at https://github.com/sfi-norwai/Di-COT.
Abdul-Kazeem Shamba, Kerstin Bach, Gavin Taylor
Department of Computer Science, Norwegian University of Science and Technology, Norway · Department of Computer Science, United States Naval Academy, USA.