Organizations: School of Information Engineering, Henan University of Science and Technology, Luoyang, 471023, Henan, China · School of Software, Henan University of Science and Technology, Luoyang, 471000, Henan, China
Multivariate time series forecasting (MTSF) is a fundamental task in many real world applications. Existing patch based forecasting methods generally fall into three categories: fixed partitioning, multi-scale partitioning, and extendable partitioning. Fixed partitioning often breaks meaningful temporal boundaries, multi-scale partitioning may introduce redundant representations across scales, and extendable partitioning improves flexibility but still lacks an explicit mechanism for organizing semantic structure and modeling interactions among heterogeneous temporal patterns. To address these limitations, we propose SCPaT, a Transformer based framework built on semantic structured partitioning. SCPaT first decomposes input sequences into semantically consistent units through adaptive semantic unit generation, then constructs a dynamic semantic graph to model directed dependencies among these units and organize them into higher order semantic blocks. Based on these structured representations, an importance aware routing mechanism adaptively dispatches different semantic blocks to different experts for customized modeling. Extensive experiments on 12 real world datasets demonstrate the effectiveness of SCPaT.
Figures & tables
Figure 1 : Comparison of four patch based schemes: fixed partitioning, multi-scale partitioning, extendable partitioning, and the proposed semantic structured partitioning.
Figure 2 : SCPaT architecture consists of: (a) Semantic Vector Encoder, which encodes time series into semantic units; (b) Transfer Entropy Graph Constructor, quantifies directed dependencies and constructs a dynamic semantic graph; and (c) Importance aware Routing, allocates customized processing strategies to different semantic blocks for effective modeling.
Task Type
Dataset
Prediction Horizons
Time Point
Dimension
Frequency
Long-term Forecasting
ETTh1
{96,192,336,720}
17420
7
Hourly
ETTh2
{96,192,336,720}
17420
7
Hourly
ETTm1
{96,192,336,720}
69680
7
15 min
ETTm2
{96,192,336,720}
69680
7
15 min
Weather
{96,192,336,720}
52603
21
10 min
Traffic
{96,192,336,720}
17451
862
Hourly
Table 1 : Summary of the 12 datasets used in our forecasting experiments, including the prediction horizons, data dimensionality, sampling frequency, and total number of time points.
Models
SCPaT ( Ours )
MSPatch
DUET
iTransformer
MSGNet
HDMixer
PatchTST
TimesNet
CrossFormer
Metric
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
ETTh1
96
0.370
0.392
0.372
0.395
0.384
0.402
0.386
0.405
0.390
0.411
0.386
0.401
0.414
0.419
0.384
0.402
0.423
0.448
192
0.418
0.422
0.428
0.426
0.437
0.429
0.441
0.436
0.442
0.442
0.441
0.430
0.460
0.445
0.446
0.439
0.471
0.474
336
0.452
0.440
0.484
0.452
0.477
0.453
0.465
0.445
0.487
0.458
0.458
0.443
0.501
0.466
0.491
0.469
0.570
0.546
720
0.449
0.458
0.480
0.473
0.486
0.475
0.503
0.491
0.494
0.488
0.512
0.484
0.500
0.488
0.521
0.500
0.653
0.621
ETTh2
96
0.288
0.338
0.292
0.344
0.298
0.351
0.297
0.349
0.328
0.371
0.293
0.338
0.302
0.348
0.340
0.374
0.745
0.584
Table 2 : Multivariate long-term forecasting results over four prediction horizons, H∈{96,192,336,720} , with the input length fixed at L=96 . The best and second-best results are marked in bold red and blue underline , respectively.
Models
SCPaT ( Ours )
MSPatch
DUET
iTransformer
MSGNet
HDMixer
TimesNet
PatchTST
CrossFormer
Metric
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
PEMS03
12
0.066
0.169
0.074
0.179
0.071
0.176
0.071
0.174
0.083
0.197
0.076
0.184
0.085
0.192
0.099
0.216
0.090
0.203
24
0.086
0.195
0.097
0.200
0.089
0.197
0.093
0.201
0.097
0.217
0.095
0.210
0.118
0.223
0.142
0.259
0.121
0.240
48
0.122
0.235
0.131
0.244
0.118
0.229
0.125
0.236
0.135
0.241
0.122
0.245
0.155
0.260
0.211
0.319
0.202
0.317
PEMS04
12
0.073
0.176
0.075
0.185
0.076
0.179
0.078
0.183
0.082
0.187
0.084
0.188
0.087
0.195
0.105
0.224
0.098
0.218
24
0.091
0.199
0.093
0.201
0.097
0.203
0.095
0.205
0.128
0.241
0.097
0.210
0.103
0.215
0.153
0.257
0.131
0.256
Table 3 : Multivariate short-term forecasting results over three prediction horizons, H∈{12,24,48} , with the input length fixed at L=96 . The best and second-best results are marked in bold red and blue underline , respectively.
Figure 3 : Sensitivity analysis of the hyperparameters α and Top-P on ETTm1, ETTh1, ETTh2, and Weather, with both the input length and prediction horizon fixed at 96.
Dataset
Expert 1
Expert 2
Expert 3
ETTh1
0.62
0.21
0.17
ETTm1
0.51
0.25
0.24
Weather
0.81
0.11
0.08
Table 4 : Average routing probability distribution across experts, with both the input length and prediction horizon fixed at 96.
Figure 4 : Performance degradation under different Gaussian noise levels on ETTh1 and Weather with both the input length and prediction horizon fixed at 96.
Figure 5 : Performance comparison under different look-back window lengths on ETTm1, ETTh1, Electricity, and Weather, with L∈{48,96,192,336,720} and the prediction horizon fixed at 96.
Missing Rate
ETTm2
ETTh1
SCPaT (Ours)
iTransformer
PatchTST
SCPaT (Ours)
iTransformer
PatchTST
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
0.00
0.173
0.258
0.180
0.264
0.181
0.259
0.370
0.392
0.386
0.405
0.414
0.419
0.05
0.213
0.297
0.249
0.441
0.252
0.327
0.379
0.401
0.404
0.413
0.436
0.428
0.10
0.248
0.326
0.295
0.472
0.328
0.376
0.390
0.412
0.433
0.424
0.455
0.441
0.15
0.287
0.354
0.379
0.559
0.411
0.423
0.404
0.424
0.462
0.432
0.467
0.460
Table 5 : Robustness analysis on the ETTm2 and ETTh1 datasets under different missing rates, with the input length fixed at L=96 and the prediction horizon fixed at H=96 . The best results are highlighted in bold.
Figure 6 : Heatmaps of the learned directed adjacency matrix Afinal on the ETTh1 and Weather datasets.
Figure 7 : Visualization of prediction results on Electricity, Traffic, and ETTh2 under the Lookback-96-Horizon-96 setting.
Dataset
ETTm2
Weather
Traffic
Prediction Length
96
192
336
720
96
192
336
720
96
192
336
720
SCPaT
MSE
0.173
0.238
0.297
0.394
0.163
0.209
0.266
0.345
0.393
0.425
0.454
0.489
MAE
0.258
0.300
0.339
0.394
0.208
0.250
0.292
0.349
0.247
0.259
0.270
0.291
w/o SV Encoder
MSE
0.196
0.259
0.317
0.412
0.204
0.248
0.302
0.373
0.414
0.448
0.478
0.510
MAE
0.278
0.318
0.354
0.408
0.254
0.287
0.323
0.367
0.274
0.281
0.294
0.316
w/o IAR
MSE
0.178
0.245
0.302
0.400
0.167
0.216
0.273
0.353
0.397
0.431
0.459
0.495
Table 6 : Ablation results of the SCPaT components on ETTm2, Weather, and Traffic over prediction horizons H∈{96,192,336,720} , with the input length fixed at L=96 . The best and second-best results are marked in bold red and blue underline , respectively.
Figure 8 : Model efficiency comparison of different methods on ETTm2 and Weather.
Time Series Forecasting (TSF) faces persistent challenges in modeling intricate temporal dependencies across different scales. Despite recent advances leveraging different decomposition operations and novel architectures based on CNN, MLP or Transformer, existing methods still struggle with static decomposition strategies, fragmented dependency modeling, and inflexible fusion mechanisms, limiting their ability to model intricate temporal dependencies. To explicitly solve the mentioned three problems respectively, we propose a novel Dynamic Multi-Scale Coordination Framework (DMSC) with Multi-Scale Patch Decomposition block (EMPD), Triad Interaction Block (TIB) and Adaptive Scale Routing MoE block (ASR-MoE). Specifically, EMPD is designed as a built-in component to dynamically segment sequences into hierarchical patches with exponentially scaled granularities, eliminating predefined scale constraints through input-adaptive patch adjustment. TIB then jointly models intra-patch, inter-patch, and cross-variable dependencies within each layer's decomposed representations. EMPD and TIB are jointly integrated into layers forming a multi-layer progressive cascade architecture, where coarse-grained representations from earlier layers adaptively guide fine-grained feature extraction in subsequent layers via gated pathways. And ASR-MoE dynamically fuses multi-scale predictions by leveraging specialized global and local experts with temporal-aware weighting. Comprehensive experiments on thirteen real-world benchmarks demonstrate that DMSC consistently maintains state-of-the-art (SOTA) performance and superior computational efficiency for TSF tasks. Code is available at https://github.com/1327679995/DMSC.
Haonan Yang, Jianchao Tang, Zhuo Li +1
National University of Defense Technology, Changsha, Hunan, China
Multivariate time series forecasting remains a challenge due to the complexity of local temporal dynamics and global dependencies across multiple variables. In this paper, we propose \textbf{N}eighboring \textbf{P}atching \textbf{Mixer} (\textbf{NPMixer}), a hierarchical architecture featuring a Learnable Stationary Wavelet Transform that adaptively learns filter coefficients to decompose signals into trend and detail components in a data-dependent manner. Our framework introduces a Neighboring Mixer Block that captures local temporal dynamics through a series of hierarchical MLP layers operating on non-overlapping patches. Specifically, the mixer block utilizes MLPs to learn temporal patterns within and across these patches, expanding the receptive field to capture multi-scale dependencies. A Channel-Mixing Encoder is applied to high-frequency components to learn channel correlations while preserving the stability of the underlying global trend. Extensive experiments on seven benchmark datasets demonstrate that NPMixer consistently outperforms state-of-the-art models, achieving better performance in 20 out of 28 (71.4%) evaluated experimental setups for MSE.
Jung Min Choi, Vijaya Krishna Yalavarthi, Lars Schmidt-Thieme
1ISMLL, University of Hildesheim, VWFS Data Analytics Research Center (VWFS-DARC), Hildesheim, Germany.
Time series foundation models (TSFMs) have recently delivered impressive zero-shot performance across diverse forecasting tasks. However, real-world decision-making frequently relies on \emph{irregular multivariate time series} (IMTS), where inconsistent inter-observation intervals and asynchronous sampling across variables coexist with informative missingness. Existing TSFMs handle such inputs either through imputation that injects spurious values or through index-based positional encodings that ignore continuous time. There is still a gap in the foundation model that follows the original IMTS patterns. In this paper, we propose a hybrid attention model that learns a unified time-aware patch representation for IMTS forecasting. We first design a \emph{time-aware patch encoding} that maps a variable number of intra-patch timestamps into a fixed-size embedding, producing a uniform format for irregular patches without resorting to imputation. We then introduce a \emph{time bias attention} mechanism that calibrates inter-patch temporal misalignment and asynchronous cross-channel dependencies as auxiliary attention offset. Finally, on top of a decoder-only Transformer backbone, we adopt a \emph{hybrid causal mask} that preserves a bidirectional full view over the historical context while keeping the forecast horizon strictly autoregressive. To support large-scale pretraining under irregular settings, we also curate VersaTSA, an archive of 30B observations that retains the native sampling sparsity of its sources. Experiments on three IMTS benchmarks and a standard regular-MTS benchmark show that our model achieves state-of-the-art zero-shot performance on IMTS and remains competitive when transferred to regular forecasting.
Li Lin, Zhihao Lin, Qi Zhang +3
Southeast University, China · Nanyang Technological University, Singapore · Timecho Ltd., China