We release Ride-Hailing, a large-scale ride-hailing time series dataset synthesized from DiDi's marketplace data across 200 spatial areas. Ride-Hailing spans four consecutive years at half-hourly granularity and covers three representative exogenous scenarios: Weather Disturbance, Holiday Effect, and Large-scale Event Impact. Built upon Ride-Hailing, we introduce RideBench, a comprehensive benchmark for exogenous-aware ride-hailing forecasting, covering both regular week-ahead forecasting and long-horizon 8-week-ahead forecasting with up to 2,688 prediction steps. RideBench evaluates over 30 representative forecasting methods, including endogenous-only models, exogenous-aware models, and time series foundation models. Our results show that future-known exogenous variables provide clear benefits in regular week-ahead forecasting, especially under weather, holiday, and large-scale event (e.g., major sporting events and concerts) scenarios. However, current exogenous-aware models still struggle to fully capture disturbance-induced pattern changes under complex external contexts. For long-horizon forecasting, existing models cannot simultaneously achieve low pointwise errors, accurate broad trends, and reliable near-term forecasts. These findings reveal a clear mismatch between existing forecasting models and real-world ride-hailing requirements, highlighting the need for models that can better exploit future-known exogenous information, scale across heterogeneous areas, and support long-horizon planning. By introducing Ride-Hailing and RideBench, we aim to encourage the community to study these practical challenges in real-world ride-hailing forecasting.
Figures & tables
Figure 1. Three key challenges of real-world ride-hailing time series forecasting.
Benchmark
Future Exo. Con./Dis.
Scenario Eval.
Multi-Area Modeling
Max Eval. Horizon
Monash ( Godahewa et al., 2021 )
✗/✗
✗
✗
168
TSlib ( Wang et al., 2024a )
✗/✗
✗
✗
720
BasicTS+ ( Shao et al., 2024 )
✗/✗
✗
✗
720
TFB ( Qiu et al., 2024 )
✓/✗
✗
✗
720
ProbTS ( Zhang et al., 2024 )
✗/✗
✗
✗
720
GIFT-Eval ( Aksu et al., 2024 )
✗/✗
✗
✗
900
Table 1. Overview of existing benchmarks and RideBench.
Figure 2. Statistics and variable information of Ride-Hailing.
Figure 3. Illustration of the key characteristics of Ride-Hailing. (a) Multi-scale periodicity at daily, weekly, and yearly levels. (b) Exogenous effects under weather disturbance, holiday effect, and large-scale event impact. (c) Heterogeneous patterns across different areas, time periods and exogenous variables, where t-SNE ( Van der Maaten and Hinton, 2008 ) is used to visualize spatial and temporal heterogeneity.
Figure 4. Overall architecture of RideBench. The left illustrates the general benchmark workflow, and the right demonstrates the detailed implementation details.
Scenarios
Overall
Weather Disturbance
Holiday Effect
Large-scale Event Impact
Rank
Metrics
WMAPE
MAE
RMSE
WMAPE
MAE
RMSE
WMAPE
MAE
RMSE
WMAPE
MAE
RMSE
Average
SeasonalMean
0.1285
4.916
14.37
0.1509
6.311
20.90
0.1701
7.052
20.85
0.1224
5.395
16.32
29
Sim.
SeasonalNaive
0.1490
5.693
17.75
0.1755
7.338
24.41
0.1981
8.203
24.57
0.1422
6.261
19.10
32
PatchTST
0.1079
4.132
13.38
0.1391
5.821
20.84
0.1507
6.249
20.37
0.1130
4.983
16.56
10
PETformer
0.1078
4.124
13.35
0.1396
5.839
20.86
0.1520
6.303
20.42
0.1130
4.982
16.48
11
TimeMixer
0.1113
4.261
13.47
0.1411
5.902
20.90
0.1548
6.420
20.42
0.1159
5.113
16.73
19
Table 2. Regular week-ahead forecasting performance under the overall evaluation and the three exogenous scenarios. Darker shading indicates better performance. Average rank is obtained via arithmetic mean of ranks across all evaluation metrics.
Setting
Overall
Trend
First Week
Metrics
WMAPE
MAE
RMSE
Rank
Acc.( ↑ )
Corr.( ↑ )
Rank
WMAPE
MAE
RMSE
Rank
SeasonalMean
0.1485
5.661
14.94
28
0.5875
0.2059
28
0.1391
5.349
14.86
31
Sim.
SeasonalNaive
0.1731
6.582
19.63
31
0.5518
0.1246
30
0.1613
6.182
19.36
32
PETformer
0.1278
4.879
14.58
2
0.6046
0.2549
12
0.1162
4.463
14.33
2
TimeMixer
0.1294
4.940
14.57
3
0.6061
0.2679
9
0.1182
4.542
14.43
3
SegRNN
0.1273
4.862
14.66
4
0.6059
0.2470
13
0.1173
4.508
14.48
6
Table 3. Eight-week-ahead forecasting performance, reporting long-horizon point-wise errors, day-level trend quality (i.e., directional accuracy of day and correlation of day), and first-week accuracy. Rank is calculated as the arithmetic mean of metric ranks across all experimental settings.
Figure 5. Efficiency of representative forecasting models under regular and long-horizon settings.
Figure 6. Effect of incorporating exogenous variables under three scenario-specific evaluations.
Figure 7. Studies on key benchmark design factors. (a) Effect of increasing the number of training areas while keeping the 10 target areas fixed. (b) Effect of varying the input lookback length. (c) Impact of different training losses.
Figure 8. Forecasting visualization under different evaluation scenarios.
Train delay prediction is an important problem for both passengers and railway operators, yet progress in the field remains difficult to assess due to the lack of standardized datasets, prediction targets, and evaluation protocols. To address this gap, we introduce RIDE, an open dataset and benchmark for train delay prediction built at nationwide scale over the Belgian railway network. RIDE covers 94.5M train events, 3.6M journeys, and 35.7M weather records from 2023 to 2025. It is organized as a layered data pipeline from raw railway and weather sources to two public releases: a reusable intermediate relational dataset and model-ready benchmark datasets. The benchmark standardizes the prediction task and the training and testing data. It also provides a unified evaluation protocol that supports direct comparison across models. Using this framework, we provide the first comprehensive comparative evaluation of non-learning, statistical learning, and deep learning models. We show that learning-based methods clearly outperform non-learning models, with graph neural networks achieving the best mean performance, while the strongest learning-based models remain relatively close to one another. Beyond aggregate mean absolute error (MAE) and root mean squared error (RMSE), the framework also provides breakdowns by prediction horizon and delay change, enabling more detailed analysis of model behavior across forecasting regimes.
Clément Elliker, Mathis Le Bail, Clément Mantoux +2
LIX, École Polytechnique, IP Paris, France · e.SNCF Solutions, France
In large-scale ride-hailing, hold control is a critical mechanism for improving passenger-driver experience. By selectively delaying certain driver-order pairs, the system waits for better opportunities, reduces cancellations, and mitigates wasted driver effort. However, existing industrial hold strategies often rely on heuristic thresholding over multiple predictive models, which can be brittle under non-stationary traffic and hard to optimize for multi-objective experience signals. We propose EXHOLD, a deployable two-stage framework decoupling experience-aware pair assessment from hold-time execution. In Stage I, we learn a decision model assigning each driver-order pair to discrete, interpretable experience tiers by optimizing a unified objective that aggregates satisfaction signals across the matching funnel. In Stage II, we solve for a monotone hold-time schedule via constrained optimization over empirical quantiles. This explicitly enforces service guardrails bounding the unnecessary holding of promising matches while maximizing overall experience improvement. We evaluate EXHOLD through randomized A/B experiments in DiDi's production system in Brazil. Results show consistent gains in marketplace efficiency and experience: EXHOLD increases trip completion and driver income, significantly reduces passenger cancellations, and improves funnel efficiency. Ablations and behavioral analyses confirm both stages are essential and that the policy makes calibrated decisions under spatiotemporal heterogeneity. EXHOLD is currently deployed, serving production traffic in Brazil.
Accurate station-level demand forecasting is essential for the efficient operation of bike-sharing systems, yet it remains challenging due to complex spatio-temporal dependencies and the large scale of urban networks. This paper presents STAGformer, a Spatio-Temporal Agent Graph Transformer that achieves efficient global modeling with linear computational complexity. The model introduces a two-step agent attention mechanism, where a small set of learnable spatial and temporal agent tokens first aggregate global information and then broadcast it back to individual stations and time steps, effectively capturing long-range interactions while reducing the quadratic cost of standard self-attention to O(NT). STAGformer integrates four core modules: a spatio-temporal encoder that fuses dynamic node features with external contextual factors (weather, time, points of interest), a graph propagation module for spatial neighbor aggregation, a temporal convolution module for local pattern extraction, and the agent attention module for global dependency modeling. Extensive experiments on two real-world datasets -- NYC Citi-Bike and Chicago Divvy-Bike -- demonstrate that STAGformer consistently outperforms state-of-the-art baselines across multiple prediction horizons, achieving significant improvements in both RMSE and MAE. Ablation studies validate the contribution of each component, with the agent attention mechanism proving critical for modeling global spatio-temporal dependencies.
Ye Zihao
Department of System Engineering City University of Hong Kong