Organizations: Department of Computer & Information Science & Engineering, University of Florida · Department of Computer Science, University of Kentucky
Traffic forecasting uses recent measurements by sensors installed at chosen locations to forecast the future road traffic. Existing work either assumes all locations are equipped with sensors or focuses on short-term forecast. This paper studies partial sensing forecast of long-term traffic, assuming sensors are available only at some locations. The problem is challenging due to the unknown data distribution at unsensed locations, the intricate spatio-temporal correlation in long-term forecasting, as well as noise to traffic patterns. We propose a Spatio-temporal Long-term Partial sensing Forecast model (SLPF) for traffic prediction, with several novel contributions, including a rank-based embedding technique to reduce the impact of noise in data, a spatial transfer matrix to overcome the spatial distribution shift from sensed locations to unsensed locations, and a multi-step training process that utilizes all available data to successively refine the model parameters for better accuracy. Extensive experiments on several real-world traffic datasets demonstrate its superior performance. Our source code is at https://github.com/zbliu98/SLPF
Figures & tables
Figure 1: Flow rates of three locations for two days
Figure 2: Noise has the greatest impact on GinAR for long-term partial sensing than for long-term full sensing, which is in turn more sensitive to noise than traditional short-term full sensing.
Figure 3: The model training phase consists of three steps above. After the model is trained and deployed, only XM,T is collected for input, and XM′,T′ is output.
Dataset
PEMS03
PEMS04
PEMS08
PEMS-BAY
METR-LA
Area
in CA, USA
Time Span
9/1/2018 -
1/1/2018 -
7/1/2016 -
1/1/2017 -
3/1/2012 -
11/30/2018
2/28/2018
8/31/2016
5/31/2017
6/30/2012
Time Interval
5 min
Number of Locations, n
358
307
170
325
207
Number of Time Intervals
26,208
16,992
17,856
52,116
34,272
Table 1: Basic statistics of the datasets used in our experiments.
Models
PEMS03
PEMS04
PEMS08
PEMSBAY
METRLA
m′=250,m′/n=69%
m′=250,m′/n=81%
m′=150,m′/n=88.2%
m′=250,m′/n=76.9%
m′=150,m′/n=72.4%
MAE
RMSE
MAPE
MAE
RMSE
MAPE
MAE
RMSE
MAPE
MAE
RMSE
MAPE
MAE
RMSE
MAPE
Matrix Factorization ( Lee and Seung, 2000 )
69.01
110.17
135.41
91.38
150.09
129.65
76.01
132.64
85.13
6.25
15.63
20.21
16.72
34.21
49.83
PatchTST* ( Nie et al., 2022 )
57.71
86.33
103.23
58.28
86.15
54.99
43.90
64.68
28.59
4.58
8.66
12.50
11.67
21.81
26.86
iTransformer* ( Liu et al., 2023b )
67.31
103.07
127.88
87.44
125.04
99.1
69.81
91.03
61.24
5.13
9.60
15.62
13.65
22.58
27.51
D2STGNN* ( Shao et al., 2022c )
37.31
61.1
45.69
47.29
75.41
48.98
47.37
67.4
35.2
6.62
10.94
15.52
14.93
27.25
42.33
Table 2: Performance comparison over five datasets, where n is the total number of locations and m′ is the number of unsensed locations. A baseline model with suffix * is the model adapted for the partial sensing task. Our models are in bold at the bottom.
Method
MAE
RMSE
MAPE
SLPF
16.58
29.22
17.85
LPF (1 step, see Fig. 3 , from XM,T to XM′,T′ )
18.47
32.08
18.83
2 Step (from XM,T to XM′,T , then from XM,T , XM′,T to XM′,T′ )
17.24
30.18
18.06
Plain Spatial Transfer Matrix
17.80
32.35
17.87
No Spatial Transfer Matrix
17.76
32.35
17.85
No Rank-based Node Embedding
17.39
31.95
17.89
Table 3: Different ablation methods under weighted selection and m′=50 condition on PEMS08 dataset over the average horizon.
Figure 4: Visualization on different locations when our model with or without ranking embedding.
Figure 5: Performance comparison with varying γ under Gaussian noise on PEMS08 dataset.
Figure 6: Performance comparison with varying γ under Uniform noise on PEMS08 dataset.
Figure 7: Performance comparison with varying γ under Laplace noise on PEMS08 dataset.
Figure 8: Comparing our model SLPF with the baselines in training efficient, in the context of accuracy-efficiency tradeoff, on dataset PEMS08, with weighted selection and m′=50 .
Figure 9: Improvement of SLPF over its variant without rank in bins of 5% forecasts with descending MAE, on dataset PEMS08 dataset with weighted selection and m′=50 .
Figure 10: Accuracy comparison in terms of RMSE with respect to the number of unsensed locations under different selection methods on PEMS08 dataset.
Models
PEMS08
MAE
RMSE
MAPE
STID ( Shao et al., 2022a )
19.59
32.20
13.05
PDFormer ( Jiang et al., 2023a )
18.58
31.86
12.64
STAEFormer ( Liu et al., 2023a )
18.77
32.73
12.10
TESTAM ( Lee and Ko, 2024 )
18.90
32.81
12.45
MegaCRN ( Jiang et al., 2023b )
19.99
33.67
13.09
Table 4: Performance comparison on PEMS08 dataset among the full-sensing baselines with the hypothetical assumption that they have the knowledge of the unsensed locations. The best results are in bold.
Figure 11: Accuracy comparison in terms of RMSE with respect to different forecasting lengths on PEMS08 dataset with the random selection method and the number of unsensed location being m′=50 .
Figure 12: Impact of number of unsensed locations on forecast accuracy (RMSE) of SLPF when it is with or without rank-based node embedding, on PEMS08 dataset with weighted selection.
Figure 13: Parameter sensitivity of the proposed SLPF on the average horizon of PEMS08 with weighted selection and m′=150 .