On-demand delivery platforms pay riders through incentive activities whose tiers are set from recent completions of riders with a similar history. Operators request such plans for changing periods, rider populations, payment rules and budgets, often for holidays or bad weather, where randomized trials are scarce and take months to collect. We present a request-driven framework that composes four stages (conditional prediction, population reduction, trajectory integration and budget allocation) through seven replaceable modules that exchange conditional trajectory laws, whose award probabilities and award-marked moments give payment and uplift for any activity rule. A response-correction step reweights trajectories from abundant no-offer history to match the moments of a short pilot. We prove that, on a fixed plan menu and given the stage errors, the end-to-end value loss is bounded by the sum of four stage terms, and that for every stage there are instances on which omitting it leaves an error floor the others cannot remove. On 3,000 riders over 45 weekly origins, all 127 windows of a week are answered 11.04x faster with identical scenarios and at most 0.92% value lost by the allocation. On 24 new controlled response laws, the response correction with a one-week pilot lowers regret by 51.2% relative to a trial with the same nominal randomized rider-weeks, and a four-week pilot with exact summation comes within +0.007 of an 18-week trial. In registered studies where windows, populations, rules and binding budgets change from request to request, the framework's regret is below that of a trial with the same nominal rider-weeks and below dose interpolation of the same pilot data, and reusing its one-off preparation answers 60 requests 14.1x and 2.70x faster with identical answers. Against a nine-offer trial fitted with the framework's own dose curve, one-week regret is 0.055 lower.
Figures & tables
Figure 1: Request-driven framework: four stages and seven modules. Dashed chips are inputs (request fields, candidate plans, activity rule, budget band, short pilot); arrows show data flow between modules. Predict (M1–M3) gives every eligible rider a conditional completion law; Reduce (M4) merges riders into activity layers with plan menus; Integrate (M5–M6) tilts the no-offer law to short-pilot moments and scores every plan by a weighted sum over stored samples of the corrected law, an exact sum when the trajectory types can be enumerated; Allocate (M7) chooses one plan per layer inside the two-sided budget band and returns the rewards that attain it. The bar names the four terms of the end-to-end bound (Proposition 1 ); the table marks what a new request recomputes.
Alternative per request
Accuracy the framework gives up
Time: alternative → framework
Speed-up
Rider records: 3,000 riders, 45 weekly origins, five activity layers
Re-draw trajectories and solve exactly
none in the scenarios (1,431/1,431 requests identical); allocation loses value in 3/405 instances, at most 0.92%
117.8 s → 3.07 s per week-long request
38.6 × [30.8, 43.0]; 11.0 × over 127 requests with the one-off pass
One plan per rider, exact
not measurable: the exact plan did not finish within 3 hours and 190 GiB of memory
relative regret (planned and scored within 20% of budget) 0.113 → 0.121 ( + 0.007 [ − 0.012, + 0.027]) with a 4-week pilot; a trial with the pilot’s rider-weeks: 0.421 with 1 week (framework 0.205) and 0.225 with 4
9,216 → 2,048 nominal randomized rider-weeks per group (18 → 4 weeks)
4.5 × fewer
Public retail (M5) and three payment rules (Section 6 )
Table 1: Accuracy given up and time gained against the most accurate way to answer the same request without the framework, on the same inputs. Speed-up: median of paired ratios (mixed-integer program: ratio of geometric-mean times); brackets: interquartile range or 95% interval over laws.
Stage
Full framework
Cheaper alternative
Accuracy change
Prediction
per rider-day distribution
point forecast
plan-cost error × 3.08 [2.48, 3.81]
Reduction
activity layers
equal-width buckets
regret 0.37% → 0.78%
Integration
joint law across days
independent days
plan-cost error × 1.51 [1.33, 1.74]
simplified normal
× 2.46 [2.02, 3.01]; 5–6-day award -10.31 pp
plug-in mean
× 1.85 [1.58, 2.19]
Allocation
budget-grid program
equal budget split
value − 15.7% [13.8, 17.4]
Table 2: Rider records, 45 weekly origins × 3 seeds, five activity layers. Each row replaces one stage of the full framework by a cheaper alternative. Ratios are of origin means with 95% intervals over origins. Regret at budget fraction 0.5. ∗ 45 origins × 3 budget fractions.
Within 20% of budget
Within 10% (pre-specified)
Method
1 wk
2 wk
4 wk
1 wk
2 wk
4 wk
Long RCT, 18 weeks (business alternative)
0.113
0.130
Same-calendar RCT
0.421
0.313
0.225
0.542
0.383
0.264
Framework (tilt, exact sum)
0.205
0.159
0.121
0.252
0.186
0.136
sampled integration ( S=2,048 )
0.245
0.197
0.163
0.286
0.223
0.189
without history (pilot frequencies)
0.603
0.560
0.527
0.658
0.611
0.569
Table 3: Relative regret (lower is better), 18 response laws × 8 replicates × 33 budgets, pilot of 1/2/4 weeks; planning, optimum and scoring within 20% of the budget or 10% (pre-specified). Indented rows change one stage (others: Table 9 ). Validity (%): issued plans with true expected spend in the 10% band, all 24 laws.
A: integrated study
B: new generator
Pilot
1 wk
2 wk
4 wk
1 wk
2 wk
4 wk
Same-calendar RCT
0.391
0.296
0.216
0.477
0.326
0.240
Framework
0.220
0.172
0.148
0.188
0.150
0.116
Dose interpolation (same information)
0.227
0.177
0.154
0.210
0.171
0.131
Independent days (least costly replacement)
0.311
0.292
0.296
0.198
0.158
0.124
Budgets that bind (%, all 48 laws)
72.3
75.0
Table 4: Two registered studies with changing requests and binding budgets (registered primary condition, 36 response laws each; relative regret within 20% of the budget). A: fresh laws from the generator of Table 3 ; B: a new rider-level generator with a rule-dependent response. Speed-up: 60 requests, one-off preparation charged, against re-running the pipeline per request.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Step
Framework
Re-draw per request
Model fitting (once per origin) a
221.3 once
221.3 once
Per-day draws, 7 days
108.7 once
108.7
Cross-day rank template
3.5 once
reused
Cross-day coupling
1.6 once
1.8
Window reduction
0.54
0.54
Integration, 540 plans
2.52
2.52
Appendix
Table 5: Where the time of rider requests goes (seconds; medians over origins, one CPU core). Top: one week-long request, by step. Bottom: K requests per origin in the fixed order (whole week, seven single days, remaining subsets); the framework total includes its once-per-origin sampling pass; speed-ups are medians of per-origin ratios. a Gradient-boosted node quantiles (29.8 s) plus share and concentration models (189.1 s), needed equally by both approaches. b Speed-up when model fitting is charged once to both sides.
Table 6: Cheaper shortcuts the framework declines, on rider records (five layers, main plan grid). Accuracy: plan-cost error of the shortcut over the framework’s, or value lost (95% intervals over origins). Time: the framework’s stage time over the shortcut’s.
Origins /
Riders
Plan-cost error
Within
Condition
runs
per run
Held-out
Reference
Change
margin
Unseen cities (four folds)
45 / 180
311
0.343
0.340
+ 0.003 [ − 0.018, + 0.020]
yes
Appendix
Table 7: Cities left out of training, rider records (five-band system, budget fraction 1). Plan-cost error is the median absolute relative error of predicted plan cost; change = held-out minus reference on the same riders and outcomes, 95% interval over origins. Pre-set margin: upper bound ≤ 0.045.
Replacement
Endpoint
Fit once
Unseen cities
Point forecast
plan-cost error
+ 0.416 [ + 0.359, + 0.472]
+ 0.380 [ + 0.292, + 0.467]
Plug-in integration
plan-cost error
+ 0.317 [ + 0.235, + 0.382]
+ 0.123 [ + 0.097, + 0.148]
Two-moment normal
plan-cost error
+ 0.459 [ + 0.382, + 0.523]
+ 0.205 [ + 0.167, + 0.240]
Independent days
plan-cost error
+ 0.202 [ + 0.142, + 0.250]
+ 0.058 [ + 0.035, + 0.080]
Equal-width buckets
regret
+ 0.0017 [ + 0.0006, + 0.0028]
+ 0.0046 [ + 0.0024, + 0.0070]
Equal budget split
regret
+ 0.048 [ + 0.021, + 0.080]
+ 0.122 [ + 0.099, + 0.145]
Appendix
Table 8: Stage replacements on held-out riders: replacement minus full framework, budget fraction 1, 95% interval over origins. Prediction and integration are scored by plan-cost error, reduction and allocation by selection regret. Unseen cities: the condition of Table 7 ; fit once: models fitted at the origin of 2026-04-21 and used for the next 16 weekly origins.
Replacement
1 wk
2 wk
4 wk
Without moment matching (history only)
0.656
0.656
0.656
Without joint integration (plug-in mean)
0.583
0.599
0.616
Without budget allocation (equal split, exact tables)
0.861
0.860
0.855
Appendix
Table 9: Further stage replacements of the framework (penalized relative regret, pre-specified 10% scoring, 18 response laws; the first two arms were not rescored at 20%, and the equal split on the exact tables is also given at 20% in the description of the methods above). Their differences from the framework are in Section 5 .
A Pilot length
1 week
2 weeks
4 weeks
Synthetic term εsynv / εsync
5.92 / 5.53
5.92 / 5.53
5.92 / 5.53
Moment term εmmv / εmmc
8.82 / 7.67
6.29 / 5.31
4.48 / 3.69
Integration term εintv / εintc ( S=2,048 )
3.43 / 3.00
3.44 / 3.00
3.51 / 3.00
Sum of stage terms
18.16 / 16.20
15.65 / 13.83
13.90 / 12.22
Direct error Eˉv / Eˉc (truth vs. final tables)
11.79 / 10.88
10.06 / 9.35
8.89 / 8.24
Requests with a comparator in B2Eˉc
26/6336
140/6336
499/6336
Appendix
Table 10: Empirical check of Proposition 1 . A : controlled response laws with known truth (24 laws × 8 replicates, 33 budgets each), sampled variant ( S=2,048 ), which carries the synthetic, moment-matching and integration terms; under exact summation the integration term is zero and the other terms are unchanged; allocation enumerates exactly, so εdec=0 , and the last two rows of A are an allocation diagnostic comparing the budget-grid program with enumeration; value errors are uplift in completions and cost errors are currency units, both summed over five layers (the median optimal uplift of a request is 14.1). B : rider records, 45 weekly origins × 3 seeds, five activity layers; each row replaces one stage of the framework and is compared with the framework on the realized records; intervals resample origins.
Within 20% of budget
Within 10%
Method
1 wk
2 wk
4 wk
1 wk
2 wk
4 wk
Same-calendar RCT
0.391
0.296
0.216
0.527
0.388
0.269
Framework
0.220
0.172
0.148
0.258
0.198
0.164
without moment matching
0.592
0.592
0.592
0.626
0.626
0.626
without history
0.471
0.433
0.404
0.561
0.509
0.470
plug-in integration
0.592
0.604
0.631
0.669
0.672
0.690
Appendix
Table 11: Integrated study: relative regret (lower is better) on 36 response laws × 8 replicates × 180 budgeted requests, with each request’s window, population, rule and budget drawn afresh. Indented rows change one stage of the framework; every difference from the framework has a 95% interval above zero. Validity: share of the framework’s issued plans whose true expected spend lies in the band, over all 48 laws.
1-week pilot
2-week pilot
Level
Laws
Framework
minus same-calendar
Framework
minus same-calendar
Rule
fixed award
36
0.225
− 0.175 [ − 0.208, − 0.142]
0.176
− 0.127 [ − 0.166, − 0.089]
per completion
36
0.229
− 0.172 [ − 0.206, − 0.138]
0.180
− 0.125 [ − 0.165, − 0.087]
attendance
36
0.205
− 0.165 [ − 0.192, − 0.139]
0.162
− 0.120 [ − 0.151, − 0.089]
Window days
1–2
36
0.221
− 0.193 [ − 0.226, − 0.159]
0.173
− 0.146 [ − 0.184, − 0.110]
3–5
36
0.219
− 0.171 [ − 0.200, − 0.141]
0.171
− 0.123 [ − 0.159, − 0.088]
Appendix
Table 12: Integrated study by request type (within 20% of budget, descriptive): framework regret and framework minus same-calendar RCT with 95% intervals over the laws in which the subgroup is defined.
Pilot
Preparation
Request (reuse)
Request (re-run)
K=8
K=32
K=60
1 week
7.27
0.22
7.57
5.8 ×
11.9 ×
14.1 × [13.8, 14.5]
2 weeks
7.40
0.22
7.57
5.8 ×
11.9 ×
14.2 × [13.8, 14.5]
4 weeks
6.25
0.21
6.61
5.9 ×
12.0 ×
14.3 × [14.0, 14.7]
Appendix
Table 13: Integrated study, preparation charged (milliseconds on one core; medians over replicates for the one-off preparation and over requests otherwise). Speed-up: time to answer the first K requests of a replicate by re-running the whole pipeline for each, over the preparation plus the reusing requests; geometric mean over the 36 response laws (replicates averaged on the log scale) with 95% intervals. Answers are identical on all 51,840 timed requests of these laws and on all 69,120 of all 48 laws.
Within 20% of budget
Within 10%
Method
1 wk
2 wk
4 wk
1 wk
2 wk
4 wk
Same-calendar RCT
0.477
0.326
0.240
0.765
0.601
0.415
Framework
0.188
0.150
0.116
0.270
0.197
0.140
without moment matching
0.433
0.433
0.433
0.781
0.781
0.781
without history
0.415
0.354
0.309
0.610
0.517
0.448
plug-in integration
0.447
0.479
0.510
0.576
0.577
0.591
Appendix
Table 14: New generator: relative regret (lower is better) on 36 response laws × 8 replicates × 180 budgeted requests. Indented rows change one stage of the framework; every difference from the framework, and that of dose interpolation, has a 95% interval above zero. Validity: share of the framework’s issued plans whose true expected spend lies in the band, over all 48 laws.
1-week pilot
4-week pilot
Level
Laws
Framework
minus same-calendar
Framework
minus same-calendar
Rule
fixed award
36
0.179
− 0.318 [ − 0.343, − 0.293]
0.120
− 0.127 [ − 0.143, − 0.112]
per completion
36
0.196
− 0.322 [ − 0.343, − 0.300]
0.111
− 0.127 [ − 0.144, − 0.111]
attendance
36
0.202
− 0.183 [ − 0.213, − 0.155]
0.132
− 0.114 [ − 0.137, − 0.093]
Window days
1–2
36
0.186
− 0.340 [ − 0.365, − 0.315]
0.119
− 0.154 [ − 0.175, − 0.133]
3–5
36
0.188
− 0.285 [ − 0.308, − 0.263]
0.116
− 0.121 [ − 0.138, − 0.106]
Appendix
Table 15: New generator by request type (within 20% of budget, descriptive): framework regret and framework minus same-calendar RCT with 95% intervals over laws.
Condition
Method
1 wk
2 wk
4 wk
Independent
Framework
0.196
0.154
0.126
Dose interpolation (same information)
0.219
0.171
0.138
Normal law, exact truncated moments
0.250
0.208
0.179
History-anchored tilt of the trial
0.405
0.291
0.211
Same-calendar RCT
0.480
0.340
0.241
Calendar
Framework , concurrent anchor
0.223
0.215
0.207
Appendix
Table 16: Study C: relative regret within 20% of the budget, 288 response laws × 8 replicates × 60 requests × 3 budgets. Every arm uses the same stored weeks, data and exact enumeration. Calendar: repeated riders with one shock per group and calendar week.
Method
1 wk
2 wk
4 wk
Framework (pilot at three offers)
0.198
0.154
0.123
Joint fit of the trial, framework’s dose curve
0.253
0.206
0.183
History-anchored tilt of the trial, per offer
0.405
0.289
0.211
Same-calendar RCT
0.479
0.337
0.241
Placement share
0.265 [0.252, 0.279]
0.382 [0.361, 0.405]
0.683 [0.623, 0.745]
Appendix
Table 17: Study E: relative regret within 20% of the budget, 288 response laws × 8 replicates × 60 requests × 3 budgets; share of the difference to the tilted trial due to placing the observations at three offers, with 95% law-bootstrap intervals.
Method
Below
Above
Excursion
Valid-plan loss
Framework
0.034
0.041
0.035
0.136
Joint fit of the trial
0.065
0.044
0.046
0.165
Dose interpolation (same information)
0.044
0.045
0.041
0.149
History-anchored tilt of the trial, per offer
0.179
0.010
0.070
0.266
Same-calendar RCT
0.272
0.010
0.082
0.288
Appendix
Table 18: Risk components at one week within 20% of the budget (study E, 288 response laws; law means). Every arm issues a plan for more than 99.99% of the requests. Below/above: share of requests whose plan has true expected spend below/above the band. Excursion: mean distance outside the band, in units of the budget, over the plans that leave it. Valid-plan loss: value lost by the plans inside the band relative to the band optimum.
Method
1 wk
2 wk
4 wk
Selected, 1 wk
Request-aware reuse
0.380
0.329
0.297
0.395
Direct reuse of the full-week calibration
0.576
0.584
0.596
0.742
Appendix
Table 19: Study D: relative regret within 20% of the budget on the 216 window-dependent laws (8 replicates × 60 requests × 3 budgets). Selected: requests that target a subpopulation chosen on the previous week.
Layers per group
All laws
Both channels
Intensity
Threshold
1 (group-level plan)
0.825 [0.820, 0.831]
0.938
0.944
0.594
2
0.900 [0.897, 0.904]
0.966
0.971
0.764
3
0.926 [0.923, 0.928]
0.974
0.978
0.825
4
0.938 [0.936, 0.940]
0.977
0.982
0.856
5
0.946 [0.944, 0.948]
0.979
0.984
0.875
Appendix
Table 20: Study F: share of the individual-level optimum retained by L history-quantile layers per group, within 20% of the budget; 3,072 laws, 60 requests × 3 budgets each; 95% law-bootstrap intervals. By family: saturating response on both channels, intensity only, threshold response near the goal.
Stage
Module
Input → output
Delivery instance
Predict
1. Response ranking (Rank)
pre-decision context and past offers → ranking by incremental activity
doubly robust learner with a 5% no-offer group ( Dudík et al., 2011 )
2. Roster (Recall)
ranking and eligibility → population fixed before attendance, incl. zero-activity riders
at least one delivery in the preceding 182 days
3. Baseline forecast (Forecast)
histories and context → no-offer outcome law for the requested window
baseline capacity → groups and non-dominated plan menus
ordered buckets merged to balance dispersion and spending stability
Integrate
5. Response correction (Correct)
no-offer law and pilot or context-matched outcomes → offer-conditioned joint trajectories
exponential tilt, Eq. ( 5 ); residual matching on incentive, time and weather
6. Settlement (Score)
trajectories, payment rule, attainable rewards → cost, award probability, score and reward witness
Eqs. ( 6 )–( 7 )
Appendix
Table 21: Four stages and seven modules. The last column gives the instance used on the delivery platform.
Held-out condition
Change in plan-cost error
Change in regret
Met
Unseen period (fit once, 16 origins)
+ 0.051 [ + 0.002, + 0.098]
+ 0.000 [ + 0.000, + 0.000]
no
Unseen cities (four folds, 45 origins)
+ 0.003 [ − 0.018, + 0.020]
+ 0.000 [ + 0.000, + 0.000]
yes
Pilot rule
1-week pilot
2-week pilot
4-week pilot
Near-lossless vs. 18-week trial (upper ≤0.05 )
+ 0.156 [ + 0.114, + 0.201] (no)
+ 0.093 [ + 0.061, + 0.126] (no)
+ 0.059 [ + 0.039, + 0.077] (no)
Beats same-calendar trial (upper <0 )
− 0.256 [ − 0.290, − 0.222] (yes)
− 0.160 [ − 0.186, − 0.134] (yes)
− 0.076 [ − 0.109, − 0.043] (yes)
Without history is worse (lower >0 )
+ 0.371 [ + 0.266, + 0.488] (yes)
+ 0.388 [ + 0.261, + 0.525] (yes)
+ 0.380 [ + 0.258, + 0.514] (yes)
Appendix
Table 22: Pre-specified decision rules and their outcomes (rules fixed before the runs, and except for the integrated study also the analysis scripts). Rider held-out validation (Appendix A.1 ): paired change, held-out minus reference, of plan-cost error and selection regret at budget fraction 1, 95% intervals over origins; the rule requires upper bounds ≤0.045 and ≤0.01 . Synthetic trajectories and a short pilot (Section 5 ): the pre-specified primary arm is the framework with sampled integration ( S=2,048 ); differences of penalized relative regret at 10% scoring on the 18 response laws, 95% intervals over laws. Integrated study (Appendix C ): the registered confirmatory family on 36 response laws; H1 is the framework minus the same-calendar trial in relative regret (band within 20% of the budget), H2 the geometric-mean speed-up of answering all 60 requests of an origin by reuse, one-off preparation charged, over re-running the whole pipeline per request; 95% intervals over laws. Its analysis script was completed after the run; H2 is computed on the registered 36 response laws. New generator (Appendix D ): the same family on its 36 response laws, with the rules and the analysis script fixed before the run. Stronger controls, behaviour change, joint fit and population reduction (Appendix E , studies C–F): registered families of seven, five, five and three one-sided tests, Holm step-down at level 0.025 within each family; estimates with 95% law-bootstrap intervals, relative regret within 20% of the budget (study F: retained share of the individual-level optimum). Study E: full re-run after the correction described in Appendix E .
Ride-hailing platforms like DiDi Chuxing operate in highly dynamic environments where balancing driver supply and passenger demand is critical. Although driver-side subsidies serve as a primary lever to align these forces and improve key KPIs like completed rides (\texttt{Rides}) and gross merchandise value (\texttt{GMV}), optimizing them in production requires simultaneously meeting three constraints: (i) responsiveness to stochastic shocks, (ii) strict subsidy-rate caps, and (iii) low-latency execution at city scale. These requirements rule out expensive per-order optimization, calling for a forward-looking, constraint-aware city-level controller for online sequential decision making. To meet these requirements, we introduce D3-Subsidy (Dynamic Driver-side Diffusion-based Subsidy), a hierarchical diffusion-based framework for deployable city-wide subsidy control. To bridge the train-inference gap, D3-Subsidy employs a prefix-conditioned diffusion model that samples plausible future trajectories from immutable historical observations, ensuring the training protocol aligns with the fixed-history nature of online deployment. These generated plans are then decoded by a context-conditioned inverse module into low-dimensional city-level control signals. For scalable execution, we bridge the gap between city-level planning and fine-grained dispatch via a Lagrangian-dual-derived mapping, which embeds subsidy-rate caps directly into order-driver incentives without iterative optimization. Additionally, a multi-city pretraining strategy with parameter-efficient fine-tuning enables robust transfer across heterogeneous cities. Extensive offline evaluations demonstrate that D3-Subsidy improves \texttt{Rides} and \texttt{GMV} while enhancing cap compliance, and a real-world A/B test confirms significant uplift while keeping budget-related violation metrics within operational thresholds.
Taijie Chen, Rui Su, Siyuan Feng +6
University of Hong Kong · Harbin Institute of Technology · Hong Kong Polytechnic University +1
We study budget pacing in repeated first-price auctions when an advertiser's private-value distributions change over time and the stationary competing-bid distribution is unknown. We ask how a feasible expenditure plan should enter online bid shading, learning, and hard budget control. We establish a plan-to-performance decomposition for a plan-driven projected-dual policy. The policy uses any feasible expenditure plan as a soft target, learns an unknown stationary competing-bid CDF from thresholds revealed after each auction, and enforces the campaign budget on every sample path. Against a distribution-informed expected-budget fluid benchmark, the uniform-plan reward gap is O(T)+O(WT), where WT measures heterogeneity in private-value distributions. With a supplied feasible plan, the global gap decomposes into a one-sided O(T) fixed-plan execution term and a plan-mismatch term bounded by (b/2a)PlanError. The same analysis provides guarantees for strict and relaxed period-cap comparators, exact recovery of the global benchmark under a specific allowance vector, and separate lower bounds establishing the necessity of the temporal-heterogeneity and Plan Error terms. An upstream planner can translate forecasts or managerial priorities into a feasible spending trajectory, while the online controller adapts bids using realized thresholds and expenditures. The guarantee is modular: it evaluates the final normalized or projected plan through PlanError. A specific forecasting model can be linked to the guarantee by establishing how its primitive estimation errors propagate to this plan-quality metric.
Yige Wang, Jiashuo Jiang
Department of Industrial Engineering & Decision Analytics, Hong Kong University of Science and Technology
Real-time bidding (RTB) ad exchanges typically forward nearly all incoming requests to demand-side platforms (DSPs), even though only a small fraction receive bids. This over-distribution weakens auction outcomes: DSPs throttle participation under compute and budget constraints, reducing the effective use of limited bidding capacity. We present a competition-aware request dispatch framework that uses distributional bid prediction and probabilistic forwarding to decide whether each request should be sent to each DSP. The system adapts per-DSP thresholds over time through lightweight policy optimization to track non-stationary market conditions. We evaluate the framework through four sequential online experiments on a production platform serving over 20 billion daily requests. A full multi-DSP deployment reduces DSP request volume under the policy by 34.2% while increasing net revenue by 4.6% (p<0.001) in a recent 14-day window after an initial DSP adaptation period. Further analysis highlights strong heterogeneity across traffic segments and reveals that aggregate metrics can be misleading. Segment-level and per-DSP analyses suggest that the policy surfaces comparative advantages among DSPs, improving monetized outcomes without increasing overall request volume.
Jonaid Shianifar, Blaz Mramor, Fangda Zou +5
Huawei Ireland Research Center, Dublin, Ireland · Huawei, Nanjing, China