AdaST: Adaptive Coupling for Spatial-Temporal Forecasting
Authors: Zhenyu Lei, Chenghao Liu, Yushun Dong, Qi R. Wang, Jundong Li
Organizations: University of Virginia Charlottesville, VA, USA · Datadog Paris, France · Florida State University Tallahassee, FL, USA · Northeastern University Boston, MA, USA
Spatial-temporal (ST) forecasting underpins many real-world systems such as traffic, climate, and energy networks. While existing methods implicitly assume strong spatiotemporal coupling, we observe that real-world ST data exhibits distinct coupling regimes, ranging from temporal-dominated and spatial-dominated to strongly coupled patterns. This mismatch causes current models to suffer from spurious dependencies and degraded performance when one correlation dominates. To overcome this limitation, we aim to dynamically modulate spatial and temporal modeling based on the data's inherent coupling structure. However, three key challenges exist: unknown coupling structure, heterogeneous coupling dynamics, and suboptimal spatial modeling. We propose AdaST, an adaptive ST forecasting framework that tackles these challenges through a decompose-recompose paradigm. AdaST factorizes inputs into components capturing different coupling patterns using heterogeneity-aware experts. Each component is processed by role-aligned modules, and a correlation-informed adaptive recomposer integrates them for final prediction. Extensive experiments confirm that AdaST significantly outperforms state-of-the-art baselines, validating the necessity of an adaptive approach.
Figures & tables
Correlation
TD
SD
SC
Temporal
0.63±0.07
0.25±0.01
0.54±0.05
Spatial
−0.05±0.06
0.33±0.12
0.21±0.02
Table 1: Average temporal and spatial correlation coefficient in three synthetic datasets. Bold values indicate dominant correlations matching the intended coupling regime.
Figure 1 : Performance of different architectures on synthetic datasets with known coupling structures. Darker colors indicate lower normalized MAE.
Figure 2 : The overall framework of AdaST. Different colors denote different components.
Dataset
PurpleAir
PEMS04
PEMS07
PEMS08
MAE
RMSE
MAPE
MAE
RMSE
MAPE
MAE
RMSE
MAPE
MAE
RMSE
MAPE
HI
3.430
5.983
72.52%
42.35
61.66
29.92%
49.03
71.18
22.75%
36.66
50.45
21.63%
DeepAR
0.994
1.817
32.48%
20.64
32.35
14.28%
22.00
35.44
9.31%
16.80
26.38
10.66%
NBeats
0.511
1.100
23.30%
25.30
39.65
17.66%
26.14
42.72
11.37%
18.90
31.39
12.11%
GWNet
0.514
1.012
23.51%
18.80
30.14
13.19%
20.47
33.47
8.61%
14.67
23.55
9.46%
DCRNN
0.656
1.268
26.99%
19.63
31.26
13.59%
21.16
34.14
9.02%
15.22
24.17
10.21%
Table 2 : Main results on PurpleAir and PEMS benchmarks. The best and second-best scores are highlighted in bold and underlined . AdaST achieves the best performance across all datasets.
Dataset
PurpleAir
PEMS07
MAE
RMSE
MAPE
MAE
RMSE
MAPE
w/o En
0.519
1.072
23.79%
19.44
32.84
8.49%
w/o Eh
0.495
0.999
22.50%
20.39
34.14
8.68%
w/o Ew
0.498
1.023
22.85%
19.21
32.63
8.07%
w/o Ea
0.516
1.075
23.47%
20.16
33.23
11.26%
SpaAtt
0.518
1.072
23.42%
19.35
32.93
8.11%
Table 3: Ablation study on PurpleAir and PEMS07. Removing each component results in performance drop.
Figure 3 : Average gate scores of different datasets, revealing data-specific coupling structures.
Figure 4 : Gate scores across different time periods and locations, illustrating dynamic coupling.
Figure 5 : T-SNE visualization of learned representations for temporal (T), spatial (S), and spatial-temporal-coupling (ST) components, demonstrating clear separation and effective disentanglement.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6 : Visualization of three synthetic datasets with distinct coupling structures: (a) Temporal-Dominated (TD), (b) Spatial-Dominated (SD), and (c) Strongly-Coupled (SC).
Figure 7 : Temporal correlation (ACF) analysis: TD data exhibits strong autocorrelation while SD data shows weak temporal dependencies.
Figure 8 : Spatial correlation (Pearson) analysis: SD data exhibits the strongest spatial dependencies while TD data shows negligible spatial correlation.
Model
PEMS04
PEMS07
PEMS08
PurpleAir
SpaMixer
233 s
1154 s
141 s
66 s
SpaAtt
341 s
2160 s
220 s
134 s
Speedup
1.5×
1.9×
1.6×
2.0×
Appendix
Table 4 : Average training time per epoch (seconds) comparing spatial mixer and spatial attention.
Method
ExchangeRate
ETTh1
METR-LA
PurpleAir
PEMS08
Avg. Rank
MAE / Rank
MAE / Rank
MAE / Rank
MAE / Rank
MAE / Rank
PatchTST
0.073 / 1
0.461 / 4
4.762 / 5
0.505 / 2
22.07 / 5
3.4
DLinear
0.077 / 3
0.444 / 2
4.820 / 6
0.525 / 3
22.51 / 6
4.0
STID
0.075 / 2
0.466 / 5
3.146 / 3
0.563 / 4
14.21 / 4
3.6
STNorm
0.108 / 5
0.492 / 6
3.153 / 4
0.574 / 5
15.41 / 3
4.6
HimNet
0.124 / 6
0.458 / 3
3.131 / 2
0.574 / 5
13.52 / 2
3.6
Appendix
Table 5 : Results on additional benchmarks. Best and second-best scores are in bold and underlined .
Variant
PEMS04
PEMS08
MAE
RMSE
MAPE
MAE
RMSE
MAPE
w/o En
18.45
29.93
12.76%
13.68
23.37
9.64%
w/o Eh
18.58
30.15
12.42%
14.65
23.94
9.51%
w/o Ew
18.35
29.88
12.68%
13.60
23.36
9.25%
w/o Ea
18.36
29.96
12.23%
13.70
23.70
9.03%
SpaAtt
18.28
30.15
12.28%
13.51
23.26
9.03%
Appendix
Table 6 : Ablation study on PEMS04 and PEMS08.
Figure 9 : Normalized correlation measures across datasets, further validating the adaptive recomposition mechanism. Strong correspondence with gate scores in Figure 3 confirms that correlation measures reliably reflect each component’s information content.
Figure 10 : Fine-grained temporal visualization (1-day span) of predictions and gate scores for 2 locations in PEMS07. More accurate predictions correspond to more dynamic gate score trajectories.
Figure 11 : Fine-grained temporal visualization (10-hour span) of predictions and gate scores for 2 locations in PEMS07. Location 0 achieves better accuracy alongside more adaptive gate score dynamics than location 1.
Figure 12 : Heatmap of gate scores across 160 spatial locations, showing a globally stable coupling structure with local variations that motivate the use of spatial heterogeneity experts.
Real-world traffic data exhibit heterogeneous spatial correlations and nonlinear temporal dynamics, posing substantial challenges for accurate spatio-temporal forecasting. Existing approaches have developed increasingly sophisticated graph, attention, and decomposition architectures, while the influence of the underlying nonlinear function approximator has received comparatively less attention. In this work, we propose STKAN, a spatio-temporal forecasting architecture that introduces Taylor-polynomial Kolmogorov--Arnold Network modules into spatial and temporal token mixing. STKAN first constructs high-level spatial representations through a learnable soft node-group assignment mechanism, applies group-wise spatial mixing, and subsequently models temporal dependencies over the compressed sequence. Spatial and temporal self-attention layers are further employed to capture long-range interactions. Experiments on five traffic forecasting benchmarks show that STKAN achieves competitive performance and performs better than the evaluated MLP-based variant in the tested settings. These results suggest that the design of nonlinear function approximators can serve as a useful complement to architectural design in spatio-temporal forecasting.
Sicong Lai, Yuehong Hu, Siru Zhong +3
The Hong Kong University of Science and Technology (Guangzhou) · Chang’an University · China University of Geosciences
Accurate traffic forecasting is essential for intelligent transportation systems, supporting a wide range of real-world applications. However, it remains challenging due to two key factors:(1) Traffic series contain heterogeneous temporal patterns, where stable periodic regularities coexist with event-driven fluctuations. Existing methods often treat them within a unified representation, limiting their ability to capture fine-grained temporal dynamics.(2)Spatial dependencies among nodes are inherently dynamic and sparse, while dense all-pairs attention often introduces redundant interactions and amplifies noise. To address these issues, we propose ADMFormer, an Adaptive-Decomposition Transformer with Time-Varying Masked Spatial Attention. Specifically, ADMFormer first employs a time-node adaptive gating mechanism to decouple traffic signals into dominant regularities and residual fluctuations that vary across time and nodes. A dual-branch temporal module is then designed to separately capture global periodic dependencies and high-frequency irregular variations from these two decomposed components. Furthermore, ADMFormer introduces a time-varying masked spatial attention that sparsifies spatial interactions based on real-time traffic states, thereby effectively preserving dynamic and informative dependencies. Extensive experiments on four real-world datasets demonstrate that ADMFormer achieves state-of-the-art performance.
Ruiwen Gu, Qitai Tan, Yahao Liu +1
Shenzhen Ubiquitous Data Enabling Key Lab Shenzhen International Graduate School, Tsinghua University, Shenzhen, China · School of Computer Science and Engineering University of Electronic Science and Technology of China, Chengdu, China
Multi-modal spatio-temporal forecasting (MM-STF) supports weather nowcasting, traffic prediction, and earth-system modeling by combining heterogeneous sources such as physical fields, satellite imagery, and in-situ sensors. Three obstacles persist: (i) modalities have different spatio-temporal sampling rates, forcing lossy interpolation onto a unified grid; (ii) modalities are frequently missing at deployment due to sensor outages or revisit gaps, while most methods train with full availability; and (iii) autoregressive decoders accumulate errors over long horizons, amplified by multi-modal conditioning. We propose AsyncCouple-Flow to address these issues jointly. A Modality-Aware Token Sparsification (MATS) module performs scale-aware tokenization and uses a shared importance scorer to select top-k tokens per timestep, producing equal-length sequences. An Asynchronous Cross-Modal Coupling Graph (ACCG) replaces fixed cross-attention with a learnable graph whose edges encode time offsets, semantic similarity, and modality-specific physical priors, enabling fusion under arbitrary asynchrony and missingness. A Flow-Matching Forecasting Head models multi-step prediction as a conditional ODE, trained with stochastic modality dropout and integrated jointly to avoid autoregressive drift. Experiments on ERA5+GOES+ISD weather forecasting and PEMS-BAY traffic prediction with multi-source side information show that AsyncCouple-Flow outperforms state-of-the-art baselines and remains robust with up to two missing modalities. The code will be released upon acceptance.
Zhixiang Wu, Yining Liu, Bo Zhao +4
Institute of Computing Technology, Chinese Academy of Sciences, China · Emory University, USA · University of California, Berkeley, USA +3