Organizations: Key Laboratory of Smart Manufacturing in Energy Chemical Process, Ministry of Education, East China University of Science and Technology, Shanghai 200237, China · Faculty of Informatics, Università della Svizzera italiana, 69000 Lugano, Switzerland · Department of Electronics, Information and Bioengineering, Politecnico di Milano, 20133 Milan, Italy
Spatio-temporal forecasting remains challenging under non-stationary environments because both data distributions and spatial relations evolve over time. Temporal normalization and de-normalization are widely used to mitigate distribution shifts, but they may distort inter-node relationships and thereby impair spatial dependency modeling. To address these issues, we propose the Distribution and Relation Adaptive Network (DRAN) for spatio-temporal forecasting. DRAN incorporates a Spatial Factor Learner (SFL) module, which enables effective normalization and de-normalization while preserving spatial dependencies in spatio-temporal systems. To model evolving spatial interactions, DRAN further proposes the Dynamic-Static Fusion Learner (DSFL) module. DSFL decomposes features into static and dynamic components and adaptively fuses them according to input variability. Experiments on six benchmark datasets show that DRAN outperforms state-of-the-art baselines. Additional analyses demonstrate that SFL consistently reduces spatial-relation distortion across multiple normalization schemes, whereas DSFL captures complementary static and dynamic dependencies and adjusts their contributions according to temporal variability.
Figures & tables
Fig. 1: Time variance in spatio-temporal systems and corresponding distribution adaptation strategies. (a) Temporal variations of a spatio-temporal system, including shifts in temporal distributions and changes in node connectivity. The color of each node represents the mean value of the node variables (consistent with the colorbar), while the vertical size of each ellipse indicates the variance. (b) Comparison between conventional normalization and de-normalization strategies for distribution adaptation, which may disrupt the effectiveness of spatial relational learning, and the proposed SFL module for preserving spatial distribution structures. “Norm.” and “De-norm.” denote normalization and de-normalization, respectively.
Fig. 2: The overall architecture of DRAN. DRAN follows a temporal-spatial learning paradigm. SFL facilitates the normalization and de-normalization processes, denoted as “Norm.” and “De-norm.”, respectively. DSFL decomposes the representations into dynamic and static components, and models spatial dependencies through attention mechanisms together with the adaptive adjacency matrix ASt constructed from learnable node embeddings.
Symbol
Description
L,H
Lookback and forecasting length.
N,C
Number of nodes and input variables.
C′
Latent feature dimension.
X , X^
Observed and predicted spatio-temporal sequence.
μX,σX
Temporal mean and standard deviation.
Hraw
Features extracted from unnormalized inputs.
TABLE I: Summary of main notation and hyperparameters.
Hyper -parameter
Weather
NYCBike1
NYCBike2
NYCTaxi
PeMS04
PeMS08
α
0.05
0.05
0.05
1
0.001
0.1
β
0.5
0.5
0.5
0.5
5
0.05
TABLE II: Balanced hyperparameter selection
Attributes
Duration time
Freq.
Node number
Length (In → Out)
Weather
01/01/2012 ∼ 31/12/2022
1 h
263
24 →
12
NYCBike1
01/04/2014 ∼ 30/09/2014
30 min
128
19 →
1
NYCBike2
01/07/2016 ∼ 29/08/2016
30 min
200
35 →
1
NYCTaxi
01/01/2015 ∼ 01/03/2015
30 min
200
35 →
1
PeMS04
01/01/2018 ∼ 28/02/2018
5 min
307
12 →
12
PeMS08
01/07/2016 ∼ 31/08/2016
5 min
170
12 →
12
TABLE III: Dataset details
Task type
Methods
Task adaptive
Dynamic adaptive
Time series forecasting
DA-RNN [ 27 ]
✗
✗
InfoTS [ 54 ]
✓
✓
AutoTCL [ 55 ]
✓
✓
Spatio-temporal forecasting
TGCN [ 15 ]
✗
✗
STGCN [ 16 ]
✗
✗
GCGRU [ 17 ]
✗
✗
TABLE IV: Baseline methods.
Model
Weather
NYCBike1
NYCBike2
MAE ↓
MAPE(%) ↓
WD ↓
MAE ↓
MAPE(%) ↓
WD ↓
MAE ↓
MAPE(%) ↓
WD ↓
DA-RNN [ 27 ]
5.492 ( ± 0.725)
1.874 ( ± 0.276)
3.758 ( ± 0.670)
15.773 ( ± 2.406)
61.948 ( ± 7.968)
5.970 ( ± 0.955)
15.159 ( ± 3.591)
63.687 ( ± 9.591)
3.567 ( ± 1.106)
InfoTS [ 54 ]
1.274 ( ± 0.137)
0.435 ( ± 0.053)
0.492 ( ± 0.035)
6.526 ( ± 0.336)
33.681 ( ± 1.831)
0.923 ( ± 0.044)
6.259 ( ± 0.368)
30.628 ( ± 1.626)
0.455 ( ± 0.044)
AutoTCL [ 55 ]
1.194 ( ± 0.022)
0.408 ( ± 0.008)
0.420 ( ± 0.019)
6.213 ( ± 0.213)
28.824 ( ± 0.808)
0.978 ( ± 0.024)
5.772 ( ± 0.246)
28.639 ( ± 0.993)
0.476 ( ± 0.026)
STGCN [ 16 ]
2.074 ( ± 1.004)
0.709 ( ± 0.396)
1.960 ( ± 0.918)
17.141 ( ± 0.142)
58.498 ( ± 1.610)
11.634 ( ± 1.827)
17.297 ( ± 0.244)
55.595 ( ± 0.805)
7.990 ( ± 0.046)
TGCN [ 15 ]
1.864 ( ± 0.797)
0.635 ( ± 0.311)
1.769 ( ± 0.704)
7.544 ( ± 0.311)
34.848 ( ± 1.287)
3.166 ( ± 0.713)
11.488 ( ± 8.575)
36.789 ( ± 7.720)
3.402 ( ± 3.165)
TABLE V: The prediction results on weather, NYCBike1 and NYCBike2 datasets.
Model
NYCTaxi
PeMS04
PeMS08
MAE ↓
MAPE(%) ↓
WD ↓
MAE ↓
MAPE(%) ↓
WD ↓
MAE ↓
MAPE(%) ↓
WD ↓
DA-RNN [ 27 ]
26.682 ( ± 8.176)
69.028 ( ± 12.085)
11.912 ( ± 9.342)
138.741 ( ± 19.896)
192.640 ( ± 37.494)
80.935 ( ± 12.140)
109.276 ( ± 13.097)
123.719 ( ± 33.187)
69.984 ( ± 11.873)
InfoTS [ 54 ]
13.286 ( ± 1.695)
21.505 ( ± 0.562)
1.670 ( ± 0.038)
25.801 ( ± 0.408)
20.058 ( ± 1.193)
11.188 ( ± 0.644)
23.604 ( ± 1.243)
14.642 ( ± 0.583)
12.469 ( ± 1.371)
AutoTCL [ 55 ]
13.119 ( ± 1.729)
21.601 ( ± 0.509)
1.768 ( ± 0.060)
23.814 ( ± 0.043)
17.355 ( ± 0.200)
10.374 ( ± 0.180)
20.879 ( ± 0.088)
13.006 ( ± 0.121)
10.074 ( ± 0.094)
STGCN [ 16 ]
25.227 ( ± 4.222)
29.533 ( ± 5.694)
8.563 ( ± 10.891)
29.114 ( ± 5.555)
23.466 ( ± 7.692)
26.608 ( ± 4.168)
27.173 ( ± 10.017)
16.585 ( ± 7.589)
13.060 ( ± 11.726)
TGCN [ 15 ]
23.227 ( ± 10.239)
44.356 ( ± 10.694)
5.563 ( ± 6.891)
34.853 ( ± 0.263)
28.078 ( ± 0.743)
30.670 ( ± 0.337)
37.969 ( ± 0.324)
30.431 ( ± 0.690)
23.658 ( ± 0.523)
TABLE VI: The prediction results on NYCTaxi, Pems04 and Pems08 datasets.
Fig. 3: Horizon-wise MAE comparison of DRAN and two strong baselines on Weather, PeMS04, and PeMS08. The curves and shaded regions denote the mean and standard deviation over five random seeds, respectively.
Fig. 4: Visualization of prediction results. (a)–(c) Absolute prediction errors of DRAN, RGSL, and DST-Mamba on the Weather dataset, respectively. (d)–(h) Prediction results on NYCBike1, where (d) shows the ground truth, (e) and (f) show the predictions of DRAN and ST-SSL, respectively, and (g) and (h) present their corresponding absolute prediction errors.
Fig. 5: Computational cost analysis on the NYCTaxi dataset. The x-axis denotes inference time (in seconds), and the y-axis represents MAE. Bubble size indicates the number of model parameters. DRAN achieves an accuracy-efficiency trade-off, obtaining lower MAE with competitive inference time compared to existing methods.
Dataset
Htem vs. X
Hspa vs. X
PDD ↓
WDrel↓
PDD ↓
WDrel↓
Weather
0.491 ( ± 0.205)
0.177 ( ± 0.056)
0.180 ( ± 0.054)
0.086 ( ± 0.029)
NYCBike1
0.481 ( ± 0.075)
0.470 ( ± 0.087)
0.215 ( ± 0.071)
0.129 ( ± 0.078)
NYCBike2
0.177 ( ± 0.034)
0.143 ( ± 0.036)
0.168 ( ± 0.071)
0.103 ( ± 0.072)
NYCTaxi
0.130 ( ± 0.024)
0.109 ( ± 0.029)
0.118 ( ± 0.040)
0.079 ( ± 0.035)
PeMS04
0.198 ( ± 0.020)
0.437 ( ± 0.233)
0.171 ( ± 0.202)
0.137 ( ± 0.024)
TABLE VII: Quantitative comparison of spatial-distribution preservation.
Dataset
Model
Clean
Gaussian Noise
Missing Sensors
High-dynamic
σ=0.1
σ=0.2
σ=0.3
ρ=10%
ρ=20%
ρ=30%
Top 10%
Weather
DRAN
0.676 ( ± 0.005)
1.369 ( ± 0.060)
2.039 ( ± 0.037)
2.557 ( ± 0.031)
1.227 ( ± 0.032)
1.328 ( ± 0.062)
1.397 ( ± 0.051)
0.948 ( ± 0.009)
RGSL
0.727 ( ± 0.003)
1.481 ( ± 0.004)
2.161 ( ± 0.006)
2.871 ( ± 0.004)
1.315 ( ± 0.004)
1.385 ( ± 0.016)
1.415 ( ± 0.046)
0.973 ( ± 0.003)
PeMS08
DRAN
13.366 ( ± 0.117)
14.366 ( ± 0.040)
16.061 ( ± 0.061)
18.483 ( ± 0.162)
21.919 ( ± 1.719)
29.591 ( ± 2.030)
39.017 ( ± 1.771)
19.325 ( ± 0.035)
STAEformer
13.538 ( ± 0.034)
14.854 ( ± 0.110)
17.168 ( ± 0.442)
20.344 ( ± 0.961)
22.412 ( ± 0.707)
32.739 ( ± 1.328)
41.851 ( ± 1.805)
19.434 ( ± 0.095)
TABLE VIII: Robustness comparison under test-time perturbations and naturally high-dynamic periods.
Strategies
Weather
NYCBike1
NYCBike2
NYCTaxi
PeMS04
PeMS08
MAE ↓
MAE ↓
MAE ↓
MAE ↓
MAE ↓
MAE ↓
+RevIN
0.844 ( ± 0.029)
5.419 ( ± 0.178)
5.386 ( ± 0.187)
11.656 ( ± 1.486)
18.875 ( ± 0.196)
13.602 ( ± 0.282)
+DAIN
1.004 ( ± 0.453)
5.541 ( ± 0.189)
5.424 ( ± 0.174)
11.867 ( ± 1.401)
18.695 ( ± 0.231)
13.704 ( ± 0.070)
+DAIN+SFL
0.840 ( ± 0.069)
5.291 ( ± 0.182)
5.276 ( ± 0.017)
11.517 ( ± 1.307)
18.305 ( ± 0.346)
13.366 ( ± 0.116)
+Non-st
1.194 ( ± 0.002)
5.502 ( ± 0.163)
5.266 ( ± 0.229)
12.426 ( ± 1.016)
18.642 ( ± 0.064)
13.995 ( ± 0.555)
+Non-st+SFL
0.676 ( ± 0.005)
5.046 ( ± 0.141)
4.845 ( ± 0.203)
10.721 ( ± 0.980)
18.132 ( ± 0.008)
13.366 ( ± 0.117)
TABLE IX: The effectiveness of SFL on various temporal normalization methods.
Fig. 6: Comparison of spatial distributions before and after SFL. Panels (a)–(d) correspond to the Weather, NYCBike1, NYCBike2, and NYCTaxi datasets, respectively. Each panel shows Gaussian kernel density estimates [ 62 ] for the raw inputs X , the latent representations before SFL Htem , and the representations after SFL Hspa , denoted as “Lookback”, “Before SFL”, and “After SFL”, respectively.
Fig. 7: Visualization of static and dynamic relations learned by DSFL on the Weather dataset. Panels (a) and (b) are the adjacency matrices of static and dynamic branches. Panels (c)–(h) show the relation strengths for three nodes located in different areas of Weather dataset. Panels (c), (e), and (g) show the static relations of the nodes, while panels (d), (f) and (h) display the dynamic relations. The orange triangles represent the selected target nodes. The darker color indicates a closer relationship between nodes.
Fig. 8: A case study of the gate fusion mechanism in the DSFL module. A time window from 21:00 on 25/09/2022 to 21:00 on 26/09/2022 in the Weather dataset is selected. Panels (a) and (b) show the learned static component HSt and dynamic component HDy , respectively, while (c) presents the corresponding gate signal. The highlighted regions (black and purple boxes) indicate periods of significant dynamic variation.
Strategies
Weather
NYCBike1
NYCBike2
NYCTaxi
PeMS04
PeMS08
MAE ↓
WD ↓
MAE ↓
WD ↓
MAE ↓
WD ↓
MAE ↓
WD ↓
MAE ↓
WD ↓
MAE ↓
WD ↓
DRAN
0.676 ( ± 0.005)
0.392 ( ± 0.011)
5.046 ( ± 0.141)
0.415 ( ± 0.148)
4.845 ( ± 0.203)
0.437 ( ± 0.039)
10.721 ( ± 0.980)
0.750 ( ± 0.146)
18.132 ( ± 0.008)
4.687 ( ± 0.187)
13.366 ( ± 0.117)
4.180 ( ± 0.075)
w/o SFL & Lspa
1.194 ( ± 0.002)
2.000 ( ± 0.045)
5.502 ( ± 0.163)
0.720 ( ± 0.061)
5.266 ( ± 0.229)
1.495 ( ± 0.021)
12.426 ( ± 1.016)
1.520 ( ± 0.076)
18.642 ( ± 0.064)
5.712 ( ± 0.099)
13.995 ( ± 0.555)
4.822 ( ± 0.265)
w/o Lspa
0.887 ( ± 0.004)
1.168 ( ± 0.069)
5.403 ( ± 0.185)
0.688 ( ± 0.080)
5.148 ( ± 0.182)
0.696 ( ± 0.007)
11.451 ( ± 1.231)
1.371 ( ± 0.118)
18.281 ( ± 0.046)
5.485 ( ± 0.186)
13.719 ( ± 0.301)
4.536 ( ± 0.076)
w/o DSFL
1.193 ( ± 0.002)
1.542 ( ± 0.072)
5.529 ( ± 0.170)
0.730 ( ± 0.080)
5.536 ( ± 0.221)
1.577 ( ± 0.058)
12.276 ( ± 1.404)
1.612 ( ± 0.143)
18.695 ( ± 0.131)
5.586 ( ± 0.157)
13.663 ( ± 0.142)
4.613 ( ± 0.126)
w/o Decomposition
0.794 ( ± 0.013)
0.669 ( ± 0.038)
5.404 ( ± 0.170)
0.655 ( ± 0.071)
5.258 ( ± 0.193)
0.600 ( ± 0.058)
11.375 ( ± 1.348)
1.345 ( ± 0.098)
18.275 ( ± 0.070)
4.908 ( ± 0.145)
13.537 ( ± 0.084)
4.534 ( ± 0.069)
TABLE X: The ablation results.
Figure 19Figure 20Figure 21Figure 22Figure 23
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Attributes
Training set
Test set
Validation set
Node number
Feature number
Weather
5,623
1,607
803
263
1
NYCBike1
3,023
864
431
128
2
NYCBike2
1,912
546
274
200
2
NYCTaxi
1,912
546
274
200
2
PeMS04
10,181
3,394
3,394
307
1
PeMS08
10,700
3,566
3,566
170
1
Appendix
TABLE A.1: Additional datasets details
Weather
NYCBike1
NYCBike2
NYCTaxi
PeMS04
PeMS08
α
MAE
α
MAE
α
MAE
α
MAE
α
MAE
α
MAE
0.001
0.856
0.001
5.102
0.001
4.958
0.001
11.375
0.001
18.186
0.001
13.604
0.010
0.895
0.010
5.075
0.010
4.942
0.010
11.423
0.010
18.249
0.010
13.682
0.050
0.667
0.050
5.037
0.050
4.855
0.050
11.191
0.050
18.343
0.050
13.560
0.100
0.726
0.100
5.082
0.100
4.972
0.100
11.085
0.100
18.240
0.100
13.499
0.500
0.758
0.500
5.101
0.500
5.040
0.500
10.892
0.500
18.299
0.500
13.526
Appendix
TABLE A.2: The selection of balanced hyperparameters.
Fig. 9: Additional node-level temporal prediction results. Panels (a) and (b) show temperature forecasts for two Weather nodes, while panels (c) and (d) show inflow and outflow forecasts for one NYCBike1 node.
Distribution shift severely degrades the performance of deep forecasting models. While this issue is well-studied for individual time series, it remains a significant challenge in the spatio-temporal domain. Effective solutions like instance normalization and its variants can mitigate temporal shifts by standardizing statistics. However, distribution shift on a graph is far more complex, involving not only the drift of individual node series but also heterogeneity across the spatial network where different nodes exhibit distinct statistical properties. To tackle this problem, we propose Reversible Residual Normalization (RRN), a novel framework that performs spatially-aware invertible transformations to address distribution shift in both spatial and temporal dimensions. Our approach integrates graph convolutional operations within invertible residual blocks, enabling adaptive normalization that respects the underlying graph structure while maintaining reversibility. By combining Center Normalization with spectral-constrained graph neural networks, our method captures and normalizes complex Spatio-Temporal relationships in a data-driven manner. The bidirectional nature of our framework allows models to learn in a normalized latent space and recover original distributional properties through inverse transformation, offering a robust and model-agnostic solution for forecasting on dynamic spatio-temporal systems.
Zhaobo Hu, Vincent Gauthier, Mehdi Naima
SAMOVAR, Télécom SudParis · Institut Polytechnique de Paris · CNRS – LIP6 +1
Spatial-temporal (ST) forecasting underpins many real-world systems such as traffic, climate, and energy networks. While existing methods implicitly assume strong spatiotemporal coupling, we observe that real-world ST data exhibits distinct coupling regimes, ranging from temporal-dominated and spatial-dominated to strongly coupled patterns. This mismatch causes current models to suffer from spurious dependencies and degraded performance when one correlation dominates. To overcome this limitation, we aim to dynamically modulate spatial and temporal modeling based on the data's inherent coupling structure. However, three key challenges exist: unknown coupling structure, heterogeneous coupling dynamics, and suboptimal spatial modeling. We propose AdaST, an adaptive ST forecasting framework that tackles these challenges through a decompose-recompose paradigm. AdaST factorizes inputs into components capturing different coupling patterns using heterogeneity-aware experts. Each component is processed by role-aligned modules, and a correlation-informed adaptive recomposer integrates them for final prediction. Extensive experiments confirm that AdaST significantly outperforms state-of-the-art baselines, validating the necessity of an adaptive approach.
Zhenyu Lei, Chenghao Liu, Yushun Dong +2
University of Virginia Charlottesville, VA, USA · Datadog Paris, France · Florida State University Tallahassee, FL, USA +1
Real-world traffic data exhibit heterogeneous spatial correlations and nonlinear temporal dynamics, posing substantial challenges for accurate spatio-temporal forecasting. Existing approaches have developed increasingly sophisticated graph, attention, and decomposition architectures, while the influence of the underlying nonlinear function approximator has received comparatively less attention. In this work, we propose STKAN, a spatio-temporal forecasting architecture that introduces Taylor-polynomial Kolmogorov--Arnold Network modules into spatial and temporal token mixing. STKAN first constructs high-level spatial representations through a learnable soft node-group assignment mechanism, applies group-wise spatial mixing, and subsequently models temporal dependencies over the compressed sequence. Spatial and temporal self-attention layers are further employed to capture long-range interactions. Experiments on five traffic forecasting benchmarks show that STKAN achieves competitive performance and performs better than the evaluated MLP-based variant in the tested settings. These results suggest that the design of nonlinear function approximators can serve as a useful complement to architectural design in spatio-temporal forecasting.
Sicong Lai, Yuehong Hu, Siru Zhong +3
The Hong Kong University of Science and Technology (Guangzhou) · Chang’an University · China University of Geosciences