Organizations: School of Information Engineering, Henan University of Science and Technology, Luoyang, 471023, Henan, China · School of Software, Henan University of Science and Technology, Luoyang, 471000, Henan, China
Multivariate time series forecasting (MTSF) is a fundamental task in many real world applications. Existing patch based forecasting methods generally fall into three categories: fixed partitioning, multi-scale partitioning, and extendable partitioning. Fixed partitioning often breaks meaningful temporal boundaries, multi-scale partitioning may introduce redundant representations across scales, and extendable partitioning improves flexibility but still lacks an explicit mechanism for organizing semantic structure and modeling interactions among heterogeneous temporal patterns. To address these limitations, we propose SCPaT, a Transformer based framework built on semantic structured partitioning. SCPaT first decomposes input sequences into semantically consistent units through adaptive semantic unit generation, then constructs a dynamic semantic graph to model directed dependencies among these units and organize them into higher order semantic blocks. Based on these structured representations, an importance aware routing mechanism adaptively dispatches different semantic blocks to different experts for customized modeling. Extensive experiments on 12 real world datasets demonstrate the effectiveness of SCPaT.
Figures & tables
Figure 1 : Comparison of four patch based schemes: fixed partitioning, multi-scale partitioning, extendable partitioning, and the proposed semantic structured partitioning.
Figure 2 : SCPaT architecture consists of: (a) Semantic Vector Encoder, which encodes time series into semantic units; (b) Transfer Entropy Graph Constructor, quantifies directed dependencies and constructs a dynamic semantic graph; and (c) Importance aware Routing, allocates customized processing strategies to different semantic blocks for effective modeling.
Task Type
Dataset
Prediction Horizons
Time Point
Dimension
Frequency
Long-term Forecasting
ETTh1
{96,192,336,720}
17420
7
Hourly
ETTh2
{96,192,336,720}
17420
7
Hourly
ETTm1
{96,192,336,720}
69680
7
15 min
ETTm2
{96,192,336,720}
69680
7
15 min
Weather
{96,192,336,720}
52603
21
10 min
Traffic
{96,192,336,720}
17451
862
Hourly
Table 1 : Summary of the 12 datasets used in our forecasting experiments, including the prediction horizons, data dimensionality, sampling frequency, and total number of time points.
Models
SCPaT ( Ours )
MSPatch
DUET
iTransformer
MSGNet
HDMixer
PatchTST
TimesNet
CrossFormer
Metric
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
ETTh1
96
0.370
0.392
0.372
0.395
0.384
0.402
0.386
0.405
0.390
0.411
0.386
0.401
0.414
0.419
0.384
0.402
0.423
0.448
192
0.418
0.422
0.428
0.426
0.437
0.429
0.441
0.436
0.442
0.442
0.441
0.430
0.460
0.445
0.446
0.439
0.471
0.474
336
0.452
0.440
0.484
0.452
0.477
0.453
0.465
0.445
0.487
0.458
0.458
0.443
0.501
0.466
0.491
0.469
0.570
0.546
720
0.449
0.458
0.480
0.473
0.486
0.475
0.503
0.491
0.494
0.488
0.512
0.484
0.500
0.488
0.521
0.500
0.653
0.621
ETTh2
96
0.288
0.338
0.292
0.344
0.298
0.351
0.297
0.349
0.328
0.371
0.293
0.338
0.302
0.348
0.340
0.374
0.745
0.584
Table 2 : Multivariate long-term forecasting results over four prediction horizons, H∈{96,192,336,720} , with the input length fixed at L=96 . The best and second-best results are marked in bold red and blue underline , respectively.
Models
SCPaT ( Ours )
MSPatch
DUET
iTransformer
MSGNet
HDMixer
TimesNet
PatchTST
CrossFormer
Metric
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
PEMS03
12
0.066
0.169
0.074
0.179
0.071
0.176
0.071
0.174
0.083
0.197
0.076
0.184
0.085
0.192
0.099
0.216
0.090
0.203
24
0.086
0.195
0.097
0.200
0.089
0.197
0.093
0.201
0.097
0.217
0.095
0.210
0.118
0.223
0.142
0.259
0.121
0.240
48
0.122
0.235
0.131
0.244
0.118
0.229
0.125
0.236
0.135
0.241
0.122
0.245
0.155
0.260
0.211
0.319
0.202
0.317
PEMS04
12
0.073
0.176
0.075
0.185
0.076
0.179
0.078
0.183
0.082
0.187
0.084
0.188
0.087
0.195
0.105
0.224
0.098
0.218
24
0.091
0.199
0.093
0.201
0.097
0.203
0.095
0.205
0.128
0.241
0.097
0.210
0.103
0.215
0.153
0.257
0.131
0.256
Table 3 : Multivariate short-term forecasting results over three prediction horizons, H∈{12,24,48} , with the input length fixed at L=96 . The best and second-best results are marked in bold red and blue underline , respectively.
Figure 3 : Sensitivity analysis of the hyperparameters α and Top-P on ETTm1, ETTh1, ETTh2, and Weather, with both the input length and prediction horizon fixed at 96.
Dataset
Expert 1
Expert 2
Expert 3
ETTh1
0.62
0.21
0.17
ETTm1
0.51
0.25
0.24
Weather
0.81
0.11
0.08
Table 4 : Average routing probability distribution across experts, with both the input length and prediction horizon fixed at 96.
Figure 4 : Performance degradation under different Gaussian noise levels on ETTh1 and Weather with both the input length and prediction horizon fixed at 96.
Figure 5 : Performance comparison under different look-back window lengths on ETTm1, ETTh1, Electricity, and Weather, with L∈{48,96,192,336,720} and the prediction horizon fixed at 96.
Missing Rate
ETTm2
ETTh1
SCPaT (Ours)
iTransformer
PatchTST
SCPaT (Ours)
iTransformer
PatchTST
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
0.00
0.173
0.258
0.180
0.264
0.181
0.259
0.370
0.392
0.386
0.405
0.414
0.419
0.05
0.213
0.297
0.249
0.441
0.252
0.327
0.379
0.401
0.404
0.413
0.436
0.428
0.10
0.248
0.326
0.295
0.472
0.328
0.376
0.390
0.412
0.433
0.424
0.455
0.441
0.15
0.287
0.354
0.379
0.559
0.411
0.423
0.404
0.424
0.462
0.432
0.467
0.460
Table 5 : Robustness analysis on the ETTm2 and ETTh1 datasets under different missing rates, with the input length fixed at L=96 and the prediction horizon fixed at H=96 . The best results are highlighted in bold.
Figure 6 : Heatmaps of the learned directed adjacency matrix Afinal on the ETTh1 and Weather datasets.
Figure 7 : Visualization of prediction results on Electricity, Traffic, and ETTh2 under the Lookback-96-Horizon-96 setting.
Dataset
ETTm2
Weather
Traffic
Prediction Length
96
192
336
720
96
192
336
720
96
192
336
720
SCPaT
MSE
0.173
0.238
0.297
0.394
0.163
0.209
0.266
0.345
0.393
0.425
0.454
0.489
MAE
0.258
0.300
0.339
0.394
0.208
0.250
0.292
0.349
0.247
0.259
0.270
0.291
w/o SV Encoder
MSE
0.196
0.259
0.317
0.412
0.204
0.248
0.302
0.373
0.414
0.448
0.478
0.510
MAE
0.278
0.318
0.354
0.408
0.254
0.287
0.323
0.367
0.274
0.281
0.294
0.316
w/o IAR
MSE
0.178
0.245
0.302
0.400
0.167
0.216
0.273
0.353
0.397
0.431
0.459
0.495
Table 6 : Ablation results of the SCPaT components on ETTm2, Weather, and Traffic over prediction horizons H∈{96,192,336,720} , with the input length fixed at L=96 . The best and second-best results are marked in bold red and blue underline , respectively.
Figure 8 : Model efficiency comparison of different methods on ETTm2 and Weather.