Organizations: School of Computer Science and Engineering, Beihang University, Beijing, China · School of Computer Science and Engineering and the School of Economics and Management, Beihang University, Beijing, China · School of Computer Science and Engineering, Beihang University, Beijing, China, and are also with the MOE Engineering Research Center of Advanced Computer Application Technology, Beihang University, China · School of Economics and Management, Beihang University, Beijing 100191, China, and the Key Laboratory of Data and Decision Intelligence (Beihang University), Ministry of Industry and Information Technology, Beijing 100191, China
Time series analysis is fundamental in domains such as finance, healthcare, and meteorology. Real-world time series often exhibit multiscale characteristics shaped by diverse latent factors, resulting in intricate temporal patterns and rich frequency structures. However, existing approaches typically focus on either frequency-domain decomposition or time-domain pattern extraction in isolation, neglecting their joint structure. This decoupled modeling limits representation expressiveness and undermines performance in tasks requiring simultaneous temporal and spectral reasoning. To address this gap, we propose m-WCN, a novel end-to-end deep learning framework that neuralizes multi-wavelet decomposition for joint extraction of temporal patterns and frequency components. By approximating the classical GHM multi-wavelet transform with trainable convolutional operators and enforcing orthogonality constraints, m-WCN produces interpretable multi-resolution representations. Built on this foundation, we introduce two task-specific architectures: TFBC for time series classification, which boosts discriminative features across frequency scales, and FTB for forecasting, which ensembles frequency-aware predictors. Extensive experiments on 64 UCR datasets and seven public forecasting benchmarks demonstrate the effectiveness of our approach. Built on the neuralized m-WCN, our TFBC and FTB outperform various baseline models across diverse datasets, achieving average improvements of 19.97% in classification and 19.92% in forecasting tasks.
Figures & tables
Notations
Description
s=(s0,…,sT)
The input series for m-WCN and TSC.
c=(c1,…,cM)
The one-hot category label for TSC.
st=(st−L,…,st)
The input series for m-WCN and TSF.
zt=(st+1,…,st+L′)
The series to be predicted for TSF.
Xhn,n∈{1,…,N}
The n -th high frequency component of s .
Xln,n∈{1,…,N}
The n -th low frequency component of s .
TABLE I: Notations used in multi-wavelet decomposition and m-WCN.
Fig. 1: Comparison of single-wavelet and multi-wavelet decomposition (Pat: Pattern). A1 and A2 denote the temporal patterns associated with the morning peak and the evening peak, respectively. Single-wavelet decomposition processes the mixed patterns A1 and A2 together, whereas multi-wavelet decomposition first separates them through pre-filtering and then performs frequency decomposition on the separated pattern components.
Fig. 2: Illustration of the m-WCN framework. Dec: Decomposition. Conv: Convolution. ↓2 denotes the downsampling operation with a sample rate of 2. ∗ is the convolutional operator.
Fig. 3: The TFBC network for classification: A three-level example.
Fig. 4: The FTB model for forecasting: A three-level example. FRP: Frequency Representation Pre-training.
Dataset
miniRocket
Hivecote2
TS-Chief
OS-CNN
TimeURL
MILLET
mWDN
TFBC
Dataset
miniRocket
Hivecote2
TS-Chief
OS-CNN
TimeURL
MILLET
mWDN
TFBC
Comput/1
0.268
0.240
0.300
0.293
0.264
0.240
0.360
0.220
CrkX/4
0.179
0.172
0.179
0.145
0.267
0.154
0.216
0.168
ElecD/1
0.258
0.274
0.284
0.276
0.273
0.271
0.342
0.235
CrkY/4
0.172
0.154
0.200
0.133
0.264
0.149
0.172
0.141
LrgKA/1
0.125
0.080
0.232
0.104
0.087
0.096
0.152
0.067
CrkZ/4
0.172
0.141
0.169
0.137
0.254
0.136
0.162
0.169
RefrD/1
0.520
0.448
0.429
0.497
0.379
0.491
0.493
0.424
Haptics/4
0.471
0.445
0.471
0.490
0.451
0.513
0.461
0.427
ScrnT/1
0.539
0.429
0.501
0.474
0.467
0.405
0.480
0.461
ISkat/4
0.524
0.456
0.473
0.571
0.556
0.533
0.566
0.445
SmKA/1
0.173
0.163
0.176
0.279
0.233
0.227
0.221
0.168
ToeS1/4
0.039
0.035
0.031
0.046
0.066
0.039
0.031
0.026
TABLE II: Classification performance comparison on 64 UCR time series datasets regarding error rate. The Dataset column is in the form of “Abbr/Type ID”, where the dataset’s full name and corresponding type are in Sec. 1.1 of SM. The best result is bold, while the second-best result is underlined.
Metric
mRocket
Hivecote2
TS-Chief
OS-CNN
TimeURL
MILLET
mWDN
TFBC
Avg. Err.
0.187
0.171
0.186
0.192
0.173
0.186
0.184
0.146
Count
3
7
6
7
12
5
11
34
Rank (a)
5.11
3.53
5.00
5.34
4.17
5.00
4.63
1.88
Rank (g)
4.58
3.10
4.32
4.61
3.32
4.32
3.71
1.59
TABLE III: TSC performance summary. Avg. Err. denotes the average classification error rate. Count: winning count. Rank(a)/Rank(g): the average ranking in terms of arithmetic and geometry.
Metric
r/fp
r/wd
w/o m-WCN
w/o OR
w/o FCL
TFBC
Avg. Err.
0.162
0.170
0.294
0.156
0.154
0.146
Count
9
15
5
10
14
26
Rank(a)
3.187
3.516
5.125
2.781
2.828
2.281
Rank(g)
2.790
2.943
4.641
2.487
2.471
1.921
TABLE IV: Ablation study on the 64 UCR datasets. Avg. Err. denotes the average classification error rate. Count: winning count. Rank(a)/Rank(g): the average ranking in terms of arithmetic and geometry.
Dataset
Horizon
Metric
Auto.
Fed.
DLinear
Basis.
Path.
TimeLLM
PatchTST
FTB
ETTh1
96
MSE
0.071
0.079
0.056
0.055
0.057
0.058
0.057
0.052
MAE
0.206
0.215
0.180
0.178
0.180
0.183
0.179
0.165
192
MSE
0.114
0.104
0.071
0.072
0.075
0.072
0.076
0.064
MAE
0.262
0.245
0.204
0.204
0.208
0.202
0.209
0.180
336
MSE
0.107
0.119
0.098
0.086
0.076
0.082
0.093
0.063
MAE
0.258
0.270
0.244
0.227
0.216
0.231
0.240
0.201
TABLE V: Performance comparison of TSF with different prediction lengths. The best result is bold, while the second-best result is underlined. Auto.=Autoformer, Fed.=Fedformer, Basis.=Basisformer, and Path.=PathFormer.
Metric
Auto.
Fed.
DLinear
Basis.
Path.
TimeLLM
PatchTST
FTB
Avg. MSE
0.192
0.147
0.163
0.146
0.153
0.129
0.126
0.112
Avg. MAE
0.303
0.265
0.271
0.248
0.242
0.238
0.233
0.216
Count
0
0
1
0
5
0
4
50
Rank(a)
7.42
5.61
5.55
4.87
3.72
4.49
3.17
1.17
Rank(g)
7.26
5.35
5.22
4.65
3.21
4.26
2.86
1.12
TABLE VI: TSF performance summary over seven forecasting datasets. Count denotes the number of wins including ties over all metric-horizon comparisons. Rank(a) and Rank(g) denote the arithmetic and geometric average ranks.
Horizon
96
192
336
720
Metric
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
FTB
0.203
0.322
0.247
0.351
0.285
0.378
0.323
0.421
r/ fp
0.215
0.352
0.266
0.386
0.288
0.359
0.342
0.432
r/ wd
0.247
0.359
0.319
0.366
0.340
0.389
0.373
0.486
w/o mwcn
0.256
0.389
0.307
0.395
0.332
0.425
0.388
0.502
w/o or
0.238
0.333
0.258
0.362
0.288
0.381
0.359
0.453
TABLE VII: Ablation study of FTB on the Electricity dataset.
Method
TT (s/epoch)
TM (MB)
IT (s)
IM (MB)
Autoformer
66.27
3848.13
22.53
738.82
Fedformer
237.78
2612.12
15.32
677.35
DLinear
7.54
21.14
4.62
19.63
Basisformer
42.49
58.47
3.52
39.15
PathFormer
58.39
2227.41
6.82
367.09
PatchTST
8.91
153.49
3.00
68.23
TABLE VIII: Average computational cost on the forecasting experiments. TT, TM, IT, and IM denote training time, training memory, inference time, and inference memory, respectively.
Fig. 5: Visualization of forecasting results. Subfigures (a) and (b) show the prediction results of FTB, while subfigures (c) and (d) show the prediction results of PatchTST on the same samples.