Station-based weather forecasting supports daily life and economic activity, yet accurate forecasts require modeling complex spatial dependencies among stations. Recent clustering-based selective modeling offers a promising alternative to dense inter-station interactions. However, a grouping shared across an observation window may obscure local changes in station relationships, while intra-cluster interactions alone may miss important global context. The theoretical advantages of selective interactions over dense connectivity also remain insufficiently understood. We therefore propose STCFormer, an adaptive spatio-temporal Transformer that dynamically groups stations according to their local evolution within each temporal patch. Its Cluster-Guided Attention Block combines fine-grained local attention within clusters and global attention over regional state summaries, allowing each station to access information beyond its own cluster. We further show that a derived Lipschitz upper bound for cluster-conditioned local attention is no larger than its fully connected counterpart, explaining a potential robustness benefit and motivating the design of InfoLoss. Experiments on three real-world weather datasets spanning eight temperature and wind forecasting tasks show that STCFormer achieves the lowest 24-hour mean squared error on all eight tasks and ranks first or second in 47 of 48 comparisons across metrics and forecasting horizons. Ablations and case studies further confirm the benefits of locally adaptive grouping and complementary local-global interactions. Our code can be obtained at https://github.com/hnu-vis/STCFormer.
Figures & tables
Figure 1: Motivation for locally adaptive station grouping. (a) A single grouping is shared across an observation window. (b) A real 72-hour case from the French temperature dataset shows the strongest relationship shifting from A–B to B–C before all three stations become highly correlated. Insets report Pearson correlations within each 24-hour segment.
Figure 2: Overview of STCFormer. (a) The forecasting pipeline stacks CGAB encoders over patch embeddings and dual-view features. (b) Dual-view representation separates temporal state and local evolution. (c) Spatio-temporal patch embedding forms joint station–time tokens. (d) CGAB combines dynamic clustering, local attention within clusters, and global attention over regional summaries.
Method
French
Hunan
Global
Temp.
U-Wind
V-Wind
Temp.
U-Wind
V-Wind
Temp.
Wind
Spatio-Temporal Methods
DCRNN
6.428±0.062
6.111±0.159
8.091±0.133
4.921 ±2.835
1.947±0.007
2.815±0.039
8.401±0.138
3.675±0.005
Corrformer
7.911±0.161
4.700±0.068
6.411±0.001
5.594±0.723
2.044±0.006
2.640±0.030
7.709±0.009
3.889±0.004
EasyST
7.823±0.697
4.735±0.092
5.884±0.047
6.673±1.390
2.026±0.034
2.680±0.039
7.333±0.028
3.715±0.007
CDPNet
7.592±0.136
4.478±0.049
4.998±0.065
5.432±0.006
1.979±0.000
2.590±0.047
7.029 ±0.018
3.638±0.001
Table 1: 24-hour forecasting results (MSE, ↓ ) with a 48-hour look-back window. Values are mean ± standard deviation over three runs. Bold and underlined means indicate the best and second-best reported values in each column; ties share the same marking.
Variant
French
Hunan
Global
Temp.
U-Wind
V-Wind
Temp.
U-Wind
V-Wind
Temp.
Wind
Window-shared
6.516±0.093
4.181±0.041
4.996±0.081
4.891±0.078
1.913±0.024
2.428±0.032
6.954±0.027
3.758±0.012
Full Attention
6.334±0.059
4.197±0.118
5.044±0.183
4.930±0.118
1.932±0.004
2.518±0.012
6.976±0.019
3.787±0.004
Static Clustering
6.503±0.175
4.204±0.062
5.079±0.130
4.904±0.401
1.905±0.013
2.437±0.018
6.932±0.016
3.796±0.006
w/o InfoLoss
6.634±0.310
4.269±0.044
4.891±0.027
5.119±0.018
1.902±0.017
2.439±0.025
6.930±0.006
3.765±0.008
w/o Global Attn.
6.363±0.065
4.193±0.019
4.926±0.048
5.185±0.083
1.902±0.002
2.421±0.007
6.966±0.012
3.760±0.003
Table 2: 24-hour ablation results (MSE, ↓ ) with a 48-hour look-back window. Values are mean ± standard deviation over three runs. Bold means indicate the best reported values in each column; ties are both bold.
Figure 3: Hyperparameter sensitivity on three temperature tasks (24-hour MSE). Red markers indicate the lowest reported MSE within each dataset’s sweep.
Figure 4: Efficiency on Hunan temperature (48-to-24-hour forecasting, batch size 8).
Figure 5: Dynamic, fixed, and random grouping on all French-temperature test windows ( 48→24 ). Parentheses report MSE.
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Operator compared
Time complexity
DCRNN
Diffusion recurrence over T steps
O(Tq(N+∣E∣))
Corrformer
Multi-correlation
O(NTlogT)
PatchTST
Per-channel temporal attention
O(NPt2)
Crossformer
Temporal attention and routers
O(NPt2+NPtr)
CCM (DLinear)
Cluster assignment and prototype cross-attention
O(NKCCM)
DUET
Masked channel-attention fusion
O(N2)
Appendix
Table 3: Computational complexity of representative dependency-modeling operators among the current baselines. N,T denote stations and history steps. Pt is each baseline’s actual temporal-token count (accounting for its patch length and stride), r is Crossformer’s router count, KCCM is CCM’s cluster count, and q,∣E∣ are DCRNN’s diffusion order and graph-edge count. For STCFormer, n=⌈N/Ls⌉ , m=⌈T/Lt⌉ , with spatial and temporal patch sizes Ls,Lt . Widths are fixed; the stated scope excludes embeddings, output heads, and training losses.
Statistic
French
Hunan
Global
Stations
234
3,570
3,850
Time steps
8,784
4,368
17,544
Start
2016-01-01 00:00
2023-04-01 09:00
2019-01-01 00:00
End
2016-12-31 23:00
2023-09-30 08:00
2020-12-31 23:00
Interval
1 hour
1 hour
1 hour
Variables
Temp., U-wind, V-wind
Temp., U-wind, V-wind
Temp., wind speed
Appendix
Table 4: Statistics of the processed weather datasets. Each variable defines a forecasting task.
Method
French
Hunan
Global
Temp.
U-Wind
V-Wind
Temp.
U-Wind
V-Wind
Temp.
Wind
Spatio-Temporal Methods
DCRNN
1.894 ±0.017
1.791±0.040
2.157±0.025
1.558±0.456
0.654±0.005
0.783±0.009
1.941±0.023
1.282 ±0.000
Corrformer
2.121±0.018
1.540±0.016
1.794±0.005
1.662±0.126
0.675±0.001
0.756±0.002
1.888±0.004
1.304±0.011
EasyST
2.138±0.100
1.554±0.023
1.673±0.000
1.829±0.230
0.683±0.012
0.762±0.012
1.843±0.003
1.289±0.002
CDPNet
2.083±0.022
1.502±0.004
1.535±0.007
1.696±0.009
0.691±0.000
0.789±0.004
1.843±0.001
1.290±0.002
Appendix
Table 5: 24-hour forecasting results (MAE, ↓ ) with a 48-hour look-back window. Values are mean ± standard deviation over three runs. Bold and underlined means indicate the best and second-best reported values in each column; ties share the same marking.
Method
French
Hunan
Global
Temp.
U-Wind
V-Wind
Temp.
U-Wind
V-Wind
Temp.
Wind
Spatio-Temporal Methods
DCRNN
0.556±0.061
0.676±0.006
0.607±0.053
0.583±0.091
0.561±0.030
0.613±0.052
0.920 ±0.003
0.648±0.008
Corrformer
0.503±0.012
0.645±0.019
0.616 ±0.015
0.649±0.058
0.532±0.001
0.627±0.001
0.887±0.008
0.595±0.002
EasyST
0.632±0.004
0.598±0.000
0.483±0.003
0.701±0.134
0.541±0.010
0.551±0.007
0.901±0.001
0.631±0.001
CDPNet
0.625±0.002
0.671±0.004
0.607±0.004
0.735±0.021
0.522±0.006
0.629 ±0.015
0.916±0.001
0.639±0.008
Appendix
Table 6: 24-hour forecasting results (SEDI, ↑ ) with a 48-hour look-back window. Values are mean ± standard deviation over three runs. Bold and underlined means indicate the best and second-best reported values in each column; ties share the same marking.
Method
French
Hunan
Global
Temp.
U-Wind
V-Wind
Temp.
U-Wind
V-Wind
Temp.
Wind
Spatio-Temporal Methods
DCRNN
3.302±0.016
1.987±0.126
2.103±0.127
2.388±0.034
0.733±0.003
0.891±0.026
2.924±0.340
1.492±0.003
Corrformer
2.711±0.030
2.100±0.034
1.996±0.008
2.813±0.245
0.721±0.002
0.893±0.005
2.784±0.005
1.465±0.002
EasyST
2.878±0.063
1.992±0.016
1.957±0.049
4.431±0.660
0.710±0.002
0.855±0.010
2.502 ±0.018
1.498±0.000
CDPNet
2.862±0.007
1.973±0.005
1.966±0.010
2.418±0.028
0.723±0.003
0.869±0.001
2.511±0.001
1.499±0.002
Appendix
Table 7: 72-hour forecasting results (MAE, ↓ ) with a 48-hour look-back window. Values are mean ± standard deviation over three runs. Bold and underlined means indicate the best and second-best reported values in each column; ties share the same marking.
Method
French
Hunan
Global
Temp.
U-Wind
V-Wind
Temp.
U-Wind
V-Wind
Temp.
Wind
Spatio-Temporal Methods
DCRNN
19.294±0.154
7.819±0.883
9.671±1.655
11.447±0.063
2.369±0.052
3.567±0.109
17.582±4.554
4.852±0.018
Corrformer
12.053±0.319
9.043±0.302
8.401±0.023
14.779±0.658
2.223±0.027
3.735±0.001
13.125±0.010
4.826±0.006
EasyST
13.579±0.614
7.361±0.078
7.437±0.387
31.283±7.987
2.216 ±0.021
3.263 ±0.074
12.895±0.239
4.892±0.017
CDPNet
13.508±0.064
7.838±0.022
8.214±0.105
11.485±0.351
2.371±0.026
3.494±0.001
12.985±0.077
5.011±0.008
Appendix
Table 8: 72-hour forecasting results (MSE, ↓ ) with a 48-hour look-back window. Values are mean ± standard deviation over three runs. Bold and underlined means indicate the best and second-best reported values in each column; ties share the same marking.
Method
French
Hunan
Global
Temp.
U-Wind
V-Wind
Temp.
U-Wind
V-Wind
Temp.
Wind
Spatio-Temporal Methods
DCRNN
0.394±0.173
0.456±0.004
0.406±0.010
0.548±0.025
0.416±0.006
0.441±0.025
0.867±0.010
0.521±0.003
Corrformer
0.342±0.015
0.472±0.010
0.370±0.014
0.328±0.046
0.481±0.007
0.473±0.016
0.849±0.004
0.450±0.003
EasyST
0.499±0.003
0.463±0.003
0.371±0.010
0.447±0.034
0.424±0.009
0.463±0.005
0.874±0.001
0.494±0.002
CDPNet
0.498±0.002
0.513±0.003
0.395±0.000
0.550 ±0.006
0.456±0.004
0.460±0.000
0.880±0.005
0.520±0.002
Appendix
Table 9: 72-hour forecasting results (SEDI, ↑ ) with a 48-hour look-back window. Values are mean ± standard deviation over three runs. Bold and underlined means indicate the best and second-best reported values in each column; ties share the same marking.
Variant
French
Hunan
Global
Temp.
U-Wind
V-Wind
Temp.
U-Wind
V-Wind
Temp.
Wind
Window-shared
1.885±0.015
1.411±0.008
1.497±0.016
1.522±0.020
0.652±0.008
0.721±0.006
1.766±0.003
1.325±0.010
Full Attention
1.869±0.010
1.412±0.019
1.490±0.032
1.518±0.023
0.649±0.000
0.722±0.001
1.766±0.001
1.298±0.001
Static Clustering
1.886±0.023
1.414±0.013
1.492±0.009
1.508±0.043
0.649±0.002
0.725±0.002
1.764±0.003
1.298±0.001
w/o InfoLoss
1.906±0.046
1.412±0.009
1.479±0.005
1.549±0.009
0.659±0.002
0.719±0.001
1.767±0.000
1.295±0.001
w/o Global Attn.
1.877±0.009
1.410±0.003
1.477±0.011
1.572±0.009
0.653±0.000
0.728±0.002
1.773±0.001
1.294±0.001
Appendix
Table 10: 24-hour ablation results (MAE, ↓ ) with a 48-hour look-back window. Values are mean ± standard deviation over three runs. Bold means indicate the best reported values in each column; ties are both bold.
Variant
French
Hunan
Global
Temp.
U-Wind
V-Wind
Temp.
U-Wind
V-Wind
Temp.
Wind
Window-shared
0.665±0.001
0.664±0.000
0.600±0.018
0.718±0.013
0.544±0.005
0.611±0.006
0.910±0.001
0.638±0.001
Full Attention
0.670±0.004
0.667±0.003
0.611±0.000
0.723±0.015
0.546±0.003
0.628±0.006
0.909±0.000
0.633±0.000
Static Clustering
0.668±0.013
0.656±0.004
0.609±0.031
0.724±0.036
0.546±0.001
0.625±0.009
0.914±0.002
0.649±0.001
w/o InfoLoss
0.670±0.003
0.655±0.002
0.596±0.007
0.713±0.010
0.553±0.007
0.632±0.004
0.911±0.000
0.645±0.000
w/o Global Attn.
0.665±0.000
0.658±0.006
0.607±0.017
0.712±0.017
0.553±0.001
0.627±0.013
0.912±0.001
0.646±0.001
Appendix
Table 11: 24-hour ablation results (SEDI, ↑ ) with a 48-hour look-back window. Values are mean ± standard deviation over three runs. Bold means indicate the best reported values in each column; ties are both bold.
Method
French
Hunan
Global
Temp.
U-Wind
V-Wind
Temp.
U-Wind
V-Wind
Temp.
Wind
Time Series Foundation Models
Timer
7.304±0.014
4.831±0.008
5.694±0.033
5.041±0.012
1.987±0.008
2.769±0.020
8.002±0.008
4.873±0.015
Chronos-Bolt-Tiny
6.706±0.012
4.538±0.002
5.437±0.002
4.990±0.003
1.924±0.003
2.537±0.000
7.894±0.004
4.691±0.005
Time-MoE
6.427±0.014
4.231±0.017
5.026±0.009
5.138±0.008
1.947±0.011
2.429±0.014
7.267±0.007
3.798±0.006
Grid-Based Predictive Models
Appendix
Table 12: Additional 24-hour forecasting comparisons (MSE, ↓ ) with a 48-hour input window. Values are mean ± standard deviation over three runs. Bold means indicate the best reported values in each column; ties are both bold.
Figure 6: Look-back sensitivity on French temperature with a fixed 72-hour forecasting horizon. The horizontal axis is the input-window length in hours.
Figure 7: Joint sensitivity to temporal and spatial patch sizes (24-hour MSE). Each bar represents one (Lt,Ls) configuration; colors distinguish spatial patch lengths.
Figure 8: Efficiency on French temperature (48-to-24-hour forecasting, batch size 8). Bubble area indicates peak training GPU memory; labels report time per epoch and memory. The time axis uses a logarithmic scale.
Method
Time
Memory
MSE
French temperature
Full Attention
12.478
0.315
6.334
CGAB
15.314
0.378
6.291
Global temperature
Full Attention
63.194
10.023
6.976
CGAB
70.427
8.913
6.917
Appendix
Table 13: CGAB versus full attention. Time: s/epoch; memory: GiB. All values are three-run means; lower is better. Best values per dataset are bold.
Construction
French
Hunan
Random
6.381±0.043
4.982±0.284
Hilbert curve
6.152±0.018
4.808±0.044
Capacity-constrained geographic clustering
6.259±0.140
4.782±0.155
Serpentine (default)
6.291±0.090
4.743±0.149
Appendix
Table 14: Spatial patch construction (temperature MSE). Values are mean ± standard deviation over three runs; best means are bold.
Metric
Preferred
Full-1L
NoInfo-1L
Δ
Mutual information I
↑
2.2878
2.0283
+0.2594
Largest-cluster ratio m(H)/n
↓
0.1248
0.2454
−0.1206
Cluster-size CV
↓
0.1513
0.6951
−0.5438
Empty-cluster rate
↓
0.0000
0.0225
−0.0225
Bound ratio B(m(H))/B(n)
↓
0.4076
0.5337
−0.1261
Appendix
Table 15: French Temperature one-layer assignment and bound diagnostics, averaged over 13,480 test-window/patch records per model. Δ is Full-1L minus NoInfo-1L.
Figure 9: Numerical diagnostics on French temperature. (a) Normalized bound versus largest-cluster ratio; open markers denote test-record medians. (b) Maximum sampled fixed-partition amplification versus median B(q) across five partition families.
Figure 10: Prediction visualization on French temperature with a 48-hour input and a 72-hour forecast horizon. All methods are shown for the same station and window. Gray: observed history; black: ground truth; red dashed: prediction. The shaded region denotes the forecast horizon.
Figure 11: Prediction visualization on Hunan temperature with a 48-hour input and a 72-hour forecast horizon. The plotting conventions follow Figure 10 . Values are shown in dataset units.
Figure 12: Prediction visualization on Global wind speed with a 48-hour input and a 72-hour forecast horizon. The plotting conventions follow Figure 10 . Stored values are divided by 10 for display, consistently for observations and predictions.
Spatial-temporal (ST) forecasting underpins many real-world systems such as traffic, climate, and energy networks. While existing methods implicitly assume strong spatiotemporal coupling, we observe that real-world ST data exhibits distinct coupling regimes, ranging from temporal-dominated and spatial-dominated to strongly coupled patterns. This mismatch causes current models to suffer from spurious dependencies and degraded performance when one correlation dominates. To overcome this limitation, we aim to dynamically modulate spatial and temporal modeling based on the data's inherent coupling structure. However, three key challenges exist: unknown coupling structure, heterogeneous coupling dynamics, and suboptimal spatial modeling. We propose AdaST, an adaptive ST forecasting framework that tackles these challenges through a decompose-recompose paradigm. AdaST factorizes inputs into components capturing different coupling patterns using heterogeneity-aware experts. Each component is processed by role-aligned modules, and a correlation-informed adaptive recomposer integrates them for final prediction. Extensive experiments confirm that AdaST significantly outperforms state-of-the-art baselines, validating the necessity of an adaptive approach.
Zhenyu Lei, Chenghao Liu, Yushun Dong +2
University of Virginia Charlottesville, VA, USA · Datadog Paris, France · Florida State University Tallahassee, FL, USA +1
Accurate traffic forecasting is essential for intelligent transportation systems, supporting a wide range of real-world applications. However, it remains challenging due to two key factors:(1) Traffic series contain heterogeneous temporal patterns, where stable periodic regularities coexist with event-driven fluctuations. Existing methods often treat them within a unified representation, limiting their ability to capture fine-grained temporal dynamics.(2)Spatial dependencies among nodes are inherently dynamic and sparse, while dense all-pairs attention often introduces redundant interactions and amplifies noise. To address these issues, we propose ADMFormer, an Adaptive-Decomposition Transformer with Time-Varying Masked Spatial Attention. Specifically, ADMFormer first employs a time-node adaptive gating mechanism to decouple traffic signals into dominant regularities and residual fluctuations that vary across time and nodes. A dual-branch temporal module is then designed to separately capture global periodic dependencies and high-frequency irregular variations from these two decomposed components. Furthermore, ADMFormer introduces a time-varying masked spatial attention that sparsifies spatial interactions based on real-time traffic states, thereby effectively preserving dynamic and informative dependencies. Extensive experiments on four real-world datasets demonstrate that ADMFormer achieves state-of-the-art performance.
Ruiwen Gu, Qitai Tan, Yahao Liu +1
Shenzhen Ubiquitous Data Enabling Key Lab Shenzhen International Graduate School, Tsinghua University, Shenzhen, China · School of Computer Science and Engineering University of Electronic Science and Technology of China, Chengdu, China
Station weather forecasting is fundamentally shaped by both complex spatial dependencies across stations and strong physical coupling among weather variables. However, existing studies often consider these relationships separately and use different datasets and experimental settings, hindering systematic assessment of their individual and joint contributions. In this paper, we introduce M2Weather, a benchmark for joint multi-station and multi-variable weather forecasting. Through multi-criteria quality control and station stratification, we collect 2,809 high-quality stations with 5 physically coupled weather variables across three spatial scales: France, Europe, and Global. This multi-scale design lets us examine whether conclusions persist from national to global station networks. We also introduce unified training and evaluation protocols to enable fair comparison of different station-variable modeling paradigms. To further examine the benefits of modeling station-variable relationships, we design a lightweight, plug-and-play adapter. With a trained weather forecasting model, this adapter can introduce missing station or variable relationships without retraining the model. This enables fair and efficient investigation of station-variable relationships. Systematic evaluation of 16 representative models shows the benefits of jointly modeling station and variable relationships. Completing missing relationships further reduces MSE for all adapted models on all three datasets. Together, these results identify the complementary information across stations and variables as an important resource for improving station weather forecasting. Our code can be obtained at https://github.com/hnu-vis/M2-Weather.
Rongwen Li, Xiao Wang, Mingyang Wang +4
Hunan University · China Meteorological Administration