Organizations: School of Computer Science and Engineering, Beihang University, Beijing, China · School of Computer Science and Engineering and the School of Economics and Management, Beihang University, Beijing, China · School of Computer Science and Engineering, Beihang University, Beijing, China, and are also with the MOE Engineering Research Center of Advanced Computer Application Technology, Beihang University, China · School of Economics and Management, Beihang University, Beijing 100191, China, and the Key Laboratory of Data and Decision Intelligence (Beihang University), Ministry of Industry and Information Technology, Beijing 100191, China
Time series analysis is fundamental in domains such as finance, healthcare, and meteorology. Real-world time series often exhibit multiscale characteristics shaped by diverse latent factors, resulting in intricate temporal patterns and rich frequency structures. However, existing approaches typically focus on either frequency-domain decomposition or time-domain pattern extraction in isolation, neglecting their joint structure. This decoupled modeling limits representation expressiveness and undermines performance in tasks requiring simultaneous temporal and spectral reasoning. To address this gap, we propose m-WCN, a novel end-to-end deep learning framework that neuralizes multi-wavelet decomposition for joint extraction of temporal patterns and frequency components. By approximating the classical GHM multi-wavelet transform with trainable convolutional operators and enforcing orthogonality constraints, m-WCN produces interpretable multi-resolution representations. Built on this foundation, we introduce two task-specific architectures: TFBC for time series classification, which boosts discriminative features across frequency scales, and FTB for forecasting, which ensembles frequency-aware predictors. Extensive experiments on 64 UCR datasets and seven public forecasting benchmarks demonstrate the effectiveness of our approach. Built on the neuralized m-WCN, our TFBC and FTB outperform various baseline models across diverse datasets, achieving average improvements of 19.97% in classification and 19.92% in forecasting tasks.
Figures & tables
Notations
Description
s=(s0,…,sT)
The input series for m-WCN and TSC.
c=(c1,…,cM)
The one-hot category label for TSC.
st=(st−L,…,st)
The input series for m-WCN and TSF.
zt=(st+1,…,st+L′)
The series to be predicted for TSF.
Xhn,n∈{1,…,N}
The n -th high frequency component of s .
Xln,n∈{1,…,N}
The n -th low frequency component of s .
TABLE I: Notations used in multi-wavelet decomposition and m-WCN.
Fig. 1: Comparison of single-wavelet and multi-wavelet decomposition (Pat: Pattern). A1 and A2 denote the temporal patterns associated with the morning peak and the evening peak, respectively. Single-wavelet decomposition processes the mixed patterns A1 and A2 together, whereas multi-wavelet decomposition first separates them through pre-filtering and then performs frequency decomposition on the separated pattern components.
Fig. 2: Illustration of the m-WCN framework. Dec: Decomposition. Conv: Convolution. ↓2 denotes the downsampling operation with a sample rate of 2. ∗ is the convolutional operator.
Fig. 3: The TFBC network for classification: A three-level example.
Fig. 4: The FTB model for forecasting: A three-level example. FRP: Frequency Representation Pre-training.
Dataset
miniRocket
Hivecote2
TS-Chief
OS-CNN
TimeURL
MILLET
mWDN
TFBC
Dataset
miniRocket
Hivecote2
TS-Chief
OS-CNN
TimeURL
MILLET
mWDN
TFBC
Comput/1
0.268
0.240
0.300
0.293
0.264
0.240
0.360
0.220
CrkX/4
0.179
0.172
0.179
0.145
0.267
0.154
0.216
0.168
ElecD/1
0.258
0.274
0.284
0.276
0.273
0.271
0.342
0.235
CrkY/4
0.172
0.154
0.200
0.133
0.264
0.149
0.172
0.141
LrgKA/1
0.125
0.080
0.232
0.104
0.087
0.096
0.152
0.067
CrkZ/4
0.172
0.141
0.169
0.137
0.254
0.136
0.162
0.169
RefrD/1
0.520
0.448
0.429
0.497
0.379
0.491
0.493
0.424
Haptics/4
0.471
0.445
0.471
0.490
0.451
0.513
0.461
0.427
ScrnT/1
0.539
0.429
0.501
0.474
0.467
0.405
0.480
0.461
ISkat/4
0.524
0.456
0.473
0.571
0.556
0.533
0.566
0.445
SmKA/1
0.173
0.163
0.176
0.279
0.233
0.227
0.221
0.168
ToeS1/4
0.039
0.035
0.031
0.046
0.066
0.039
0.031
0.026
TABLE II: Classification performance comparison on 64 UCR time series datasets regarding error rate. The Dataset column is in the form of “Abbr/Type ID”, where the dataset’s full name and corresponding type are in Sec. 1.1 of SM. The best result is bold, while the second-best result is underlined.
Metric
mRocket
Hivecote2
TS-Chief
OS-CNN
TimeURL
MILLET
mWDN
TFBC
Avg. Err.
0.187
0.171
0.186
0.192
0.173
0.186
0.184
0.146
Count
3
7
6
7
12
5
11
34
Rank (a)
5.11
3.53
5.00
5.34
4.17
5.00
4.63
1.88
Rank (g)
4.58
3.10
4.32
4.61
3.32
4.32
3.71
1.59
TABLE III: TSC performance summary. Avg. Err. denotes the average classification error rate. Count: winning count. Rank(a)/Rank(g): the average ranking in terms of arithmetic and geometry.
Metric
r/fp
r/wd
w/o m-WCN
w/o OR
w/o FCL
TFBC
Avg. Err.
0.162
0.170
0.294
0.156
0.154
0.146
Count
9
15
5
10
14
26
Rank(a)
3.187
3.516
5.125
2.781
2.828
2.281
Rank(g)
2.790
2.943
4.641
2.487
2.471
1.921
TABLE IV: Ablation study on the 64 UCR datasets. Avg. Err. denotes the average classification error rate. Count: winning count. Rank(a)/Rank(g): the average ranking in terms of arithmetic and geometry.
Dataset
Horizon
Metric
Auto.
Fed.
DLinear
Basis.
Path.
TimeLLM
PatchTST
FTB
ETTh1
96
MSE
0.071
0.079
0.056
0.055
0.057
0.058
0.057
0.052
MAE
0.206
0.215
0.180
0.178
0.180
0.183
0.179
0.165
192
MSE
0.114
0.104
0.071
0.072
0.075
0.072
0.076
0.064
MAE
0.262
0.245
0.204
0.204
0.208
0.202
0.209
0.180
336
MSE
0.107
0.119
0.098
0.086
0.076
0.082
0.093
0.063
MAE
0.258
0.270
0.244
0.227
0.216
0.231
0.240
0.201
TABLE V: Performance comparison of TSF with different prediction lengths. The best result is bold, while the second-best result is underlined. Auto.=Autoformer, Fed.=Fedformer, Basis.=Basisformer, and Path.=PathFormer.
Metric
Auto.
Fed.
DLinear
Basis.
Path.
TimeLLM
PatchTST
FTB
Avg. MSE
0.192
0.147
0.163
0.146
0.153
0.129
0.126
0.112
Avg. MAE
0.303
0.265
0.271
0.248
0.242
0.238
0.233
0.216
Count
0
0
1
0
5
0
4
50
Rank(a)
7.42
5.61
5.55
4.87
3.72
4.49
3.17
1.17
Rank(g)
7.26
5.35
5.22
4.65
3.21
4.26
2.86
1.12
TABLE VI: TSF performance summary over seven forecasting datasets. Count denotes the number of wins including ties over all metric-horizon comparisons. Rank(a) and Rank(g) denote the arithmetic and geometric average ranks.
Horizon
96
192
336
720
Metric
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
FTB
0.203
0.322
0.247
0.351
0.285
0.378
0.323
0.421
r/ fp
0.215
0.352
0.266
0.386
0.288
0.359
0.342
0.432
r/ wd
0.247
0.359
0.319
0.366
0.340
0.389
0.373
0.486
w/o mwcn
0.256
0.389
0.307
0.395
0.332
0.425
0.388
0.502
w/o or
0.238
0.333
0.258
0.362
0.288
0.381
0.359
0.453
TABLE VII: Ablation study of FTB on the Electricity dataset.
Method
TT (s/epoch)
TM (MB)
IT (s)
IM (MB)
Autoformer
66.27
3848.13
22.53
738.82
Fedformer
237.78
2612.12
15.32
677.35
DLinear
7.54
21.14
4.62
19.63
Basisformer
42.49
58.47
3.52
39.15
PathFormer
58.39
2227.41
6.82
367.09
PatchTST
8.91
153.49
3.00
68.23
TABLE VIII: Average computational cost on the forecasting experiments. TT, TM, IT, and IM denote training time, training memory, inference time, and inference memory, respectively.
Fig. 5: Visualization of forecasting results. Subfigures (a) and (b) show the prediction results of FTB, while subfigures (c) and (d) show the prediction results of PatchTST on the same samples.
Multivariate time series forecasting remains a challenge due to the complexity of local temporal dynamics and global dependencies across multiple variables. In this paper, we propose \textbf{N}eighboring \textbf{P}atching \textbf{Mixer} (\textbf{NPMixer}), a hierarchical architecture featuring a Learnable Stationary Wavelet Transform that adaptively learns filter coefficients to decompose signals into trend and detail components in a data-dependent manner. Our framework introduces a Neighboring Mixer Block that captures local temporal dynamics through a series of hierarchical MLP layers operating on non-overlapping patches. Specifically, the mixer block utilizes MLPs to learn temporal patterns within and across these patches, expanding the receptive field to capture multi-scale dependencies. A Channel-Mixing Encoder is applied to high-frequency components to learn channel correlations while preserving the stability of the underlying global trend. Extensive experiments on seven benchmark datasets demonstrate that NPMixer consistently outperforms state-of-the-art models, achieving better performance in 20 out of 28 (71.4%) evaluated experimental setups for MSE.
Jung Min Choi, Vijaya Krishna Yalavarthi, Lars Schmidt-Thieme
1ISMLL, University of Hildesheim, VWFS Data Analytics Research Center (VWFS-DARC), Hildesheim, Germany.
Multimodal time series forecasting is crucial in real-world applications, where decisions depend on both numerical data and contextual signals. The core challenge is to effectively combine temporal numerical patterns with the context embedded in other modalities, such as text. While most existing methods align textual features with time-series patterns one step at a time, they neglect the multiscale temporal influences of contextual information such as time-series cycles and dynamic shifts. This mismatch between local alignment and global textual context can be addressed by spectral decomposition, which separates time series into frequency components capturing both short-term changes and long-term trends. In this paper, we propose SpecTF, a simple yet effective framework that integrates the effect of textual data on time series in the frequency domain. Our method extracts textual embeddings, projects them into the frequency domain, and fuses them with the time series' spectral components using a lightweight cross-attention mechanism. This adaptively reweights frequency bands based on textual relevance before mapping the results back to the temporal domain for predictions. Experimental results demonstrate that SpecTF significantly outperforms state-of-the-art models across diverse multi-modal time series datasets while utilizing considerably fewer parameters. Code is available at https://github.com/hiepnh137/SpecTF.
Huu Hiep Nguyen, Minh Hoang Nguyen, Dung Nguyen +1
Applied Artificial Intelligence Initiative Deakin University Geelong, Australia
Multivariate Time Series Classification (MTSC) demands models that can effectively capture complex temporal patterns across multiple scales while remaining computationally efficient. However, existing approaches generally struggle to reconcile fine-grained representation learning, especially under class imbalance and real-world constraints. In this paper, we present FreSH, a Frequency-Segmented Hierarchical Multi-Expert Framework designed to address these challenges. FreSH introduces a new perspective for MTSC by enabling adaptive, multi-scale analysis of temporal signals, allowing different aspects of the data to be modeled in a complementary and coordinated manner. By combining localized specialization with holistic context modeling, FreSH achieves strong representational capacity without incurring excessive computational overhead. An adaptive fusion strategy further enhances flexibility, enabling the model to dynamically emphasize the most informative components of the input. In addition, we incorporate a more robust optimization objective that improves learning stability across varying sample difficulties and class distributions. Extensive evaluations on 30 UEA benchmark datasets and real-world vibration data demonstrate that FreSH consistently outperforms state-of-the-art methods in classification accuracy, while substantially reducing model size and efficiency.
Pingping Liu, Muyao Wang, Zijian Zhang +5
1Jilin University · 2Hong Kong Polytechnic University · 3Pengcheng Laboratory +1