RideBench: A Large-Scale Exogenous-Aware Benchmark for Ride-Hailing Time Series Forecasting
Organizations: South China University of Technology, China · Didi Chuxing, China · The University of Hong Kong, China
Abstract
We release Ride-Hailing, a large-scale ride-hailing time series dataset synthesized from DiDi's marketplace data across 200 spatial areas. Ride-Hailing spans four consecutive years at half-hourly granularity and covers three representative exogenous scenarios: Weather Disturbance, Holiday Effect, and Large-scale Event Impact. Built upon Ride-Hailing, we introduce RideBench, a comprehensive benchmark for exogenous-aware ride-hailing forecasting, covering both regular week-ahead forecasting and long-horizon 8-week-ahead forecasting with up to 2,688 prediction steps. RideBench evaluates over 30 representative forecasting methods, including endogenous-only models, exogenous-aware models, and time series foundation models. Our results show that future-known exogenous variables provide clear benefits in regular week-ahead forecasting, especially under weather, holiday, and large-scale event (e.g., major sporting events and concerts) scenarios. However, current exogenous-aware models still struggle to fully capture disturbance-induced pattern changes under complex external contexts. For long-horizon forecasting, existing models cannot simultaneously achieve low pointwise errors, accurate broad trends, and reliable near-term forecasts. These findings reveal a clear mismatch between existing forecasting models and real-world ride-hailing requirements, highlighting the need for models that can better exploit future-known exogenous information, scale across heterogeneous areas, and support long-horizon planning. By introducing Ride-Hailing and RideBench, we aim to encourage the community to study these practical challenges in real-world ride-hailing forecasting.
Figures & tables
| Benchmark | Future Exo. Con./Dis. | Scenario Eval. | Multi-Area Modeling | Max Eval. Horizon |
| Monash ( Godahewa et al., 2021 ) | ✗/✗ | ✗ | ✗ | 168 |
| TSlib ( Wang et al., 2024a ) | ✗/✗ | ✗ | ✗ | 720 |
| BasicTS+ ( Shao et al., 2024 ) | ✗/✗ | ✗ | ✗ | 720 |
| TFB ( Qiu et al., 2024 ) | ✓/✗ | ✗ | ✗ | 720 |
| ProbTS ( Zhang et al., 2024 ) | ✗/✗ | ✗ | ✗ | 720 |
| GIFT-Eval ( Aksu et al., 2024 ) | ✗/✗ | ✗ | ✗ | 900 |
| Scenarios | Overall | Weather Disturbance | Holiday Effect | Large-scale Event Impact | Rank | |||||||||
| Metrics | WMAPE | MAE | RMSE | WMAPE | MAE | RMSE | WMAPE | MAE | RMSE | WMAPE | MAE | RMSE | Average | |
| SeasonalMean | 0.1285 | 4.916 | 14.37 | 0.1509 | 6.311 | 20.90 | 0.1701 | 7.052 | 20.85 | 0.1224 | 5.395 | 16.32 | 29 | |
| Sim. | SeasonalNaive | 0.1490 | 5.693 | 17.75 | 0.1755 | 7.338 | 24.41 | 0.1981 | 8.203 | 24.57 | 0.1422 | 6.261 | 19.10 | 32 |
| PatchTST | 0.1079 | 4.132 | 13.38 | 0.1391 | 5.821 | 20.84 | 0.1507 | 6.249 | 20.37 | 0.1130 | 4.983 | 16.56 | 10 | |
| PETformer | 0.1078 | 4.124 | 13.35 | 0.1396 | 5.839 | 20.86 | 0.1520 | 6.303 | 20.42 | 0.1130 | 4.982 | 16.48 | 11 | |
| TimeMixer | 0.1113 | 4.261 | 13.47 | 0.1411 | 5.902 | 20.90 | 0.1548 | 6.420 | 20.42 | 0.1159 | 5.113 | 16.73 | 19 | |
| Setting | Overall | Trend | First Week | |||||||||
| Metrics | WMAPE | MAE | RMSE | Rank | Acc.( ) | Corr.( ) | Rank | WMAPE | MAE | RMSE | Rank | |
| SeasonalMean | 0.1485 | 5.661 | 14.94 | 28 | 0.5875 | 0.2059 | 28 | 0.1391 | 5.349 | 14.86 | 31 | |
| Sim. | SeasonalNaive | 0.1731 | 6.582 | 19.63 | 31 | 0.5518 | 0.1246 | 30 | 0.1613 | 6.182 | 19.36 | 32 |
| PETformer | 0.1278 | 4.879 | 14.58 | 2 | 0.6046 | 0.2549 | 12 | 0.1162 | 4.463 | 14.33 | 2 | |
| TimeMixer | 0.1294 | 4.940 | 14.57 | 3 | 0.6061 | 0.2679 | 9 | 0.1182 | 4.542 | 14.43 | 3 | |
| SegRNN | 0.1273 | 4.862 | 14.66 | 4 | 0.6059 | 0.2470 | 13 | 0.1173 | 4.508 | 14.48 | 6 | |