A General Framework for Budgeted Threshold Incentives on Request
Organizations: Meituan · National Taiwan University · Tsinghua University · Tongji University
Abstract
On-demand delivery platforms pay riders through incentive activities whose tiers are set from recent completions of riders with a similar history. Operators request such plans for changing periods, rider populations, payment rules and budgets, often for holidays or bad weather, where randomized trials are scarce and take months to collect. We present a request-driven framework that composes four stages (conditional prediction, population reduction, trajectory integration and budget allocation) through seven replaceable modules that exchange conditional trajectory laws, whose award probabilities and award-marked moments give payment and uplift for any activity rule. A response-correction step reweights trajectories from abundant no-offer history to match the moments of a short pilot. We prove that, on a fixed plan menu and given the stage errors, the end-to-end value loss is bounded by the sum of four stage terms, and that for every stage there are instances on which omitting it leaves an error floor the others cannot remove. On 3,000 riders over 45 weekly origins, all 127 windows of a week are answered 11.04x faster with identical scenarios and at most 0.92% value lost by the allocation. On 24 new controlled response laws, the response correction with a one-week pilot lowers regret by 51.2% relative to a trial with the same nominal randomized rider-weeks, and a four-week pilot with exact summation comes within +0.007 of an 18-week trial. In registered studies where windows, populations, rules and binding budgets change from request to request, the framework's regret is below that of a trial with the same nominal rider-weeks and below dose interpolation of the same pilot data, and reusing its one-off preparation answers 60 requests 14.1x and 2.70x faster with identical answers. Against a nine-offer trial fitted with the framework's own dose curve, one-week regret is 0.055 lower.
Figures & tables
| Alternative per request | Accuracy the framework gives up | Time: alternative framework | Speed-up |
|---|---|---|---|
| Rider records: 3,000 riders, 45 weekly origins, five activity layers | |||
| Re-draw trajectories and solve exactly | none in the scenarios (1,431/1,431 requests identical); allocation loses value in 3/405 instances, at most 0.92% | 117.8 s 3.07 s per week-long request | 38.6 [30.8, 43.0]; 11.0 over 127 requests with the one-off pass |
| One plan per rider, exact | not measurable: the exact plan did not finish within 3 hours and 190 GiB of memory | unfinished 4.05 s per table | — |
| Controlled response laws: 18 laws 8 replicates, 33 budgets each (Section 5 ) | |||
| Randomized trial over all offers | relative regret (planned and scored within 20% of budget) 0.113 0.121 ( 0.007 [ 0.012, 0.027]) with a 4-week pilot; a trial with the pilot’s rider-weeks: 0.421 with 1 week (framework 0.205) and 0.225 with 4 | 9,216 2,048 nominal randomized rider-weeks per group (18 4 weeks) | 4.5 fewer |
| Public retail (M5) and three payment rules (Section 6 ) | |||
| Stage | Full framework | Cheaper alternative | Accuracy change |
|---|---|---|---|
| Prediction | per rider-day distribution | point forecast | plan-cost error 3.08 [2.48, 3.81] |
| Reduction | activity layers | equal-width buckets | regret 0.37% 0.78% |
| Integration | joint law across days | independent days | plan-cost error 1.51 [1.33, 1.74] |
| simplified normal | 2.46 [2.02, 3.01]; 5–6-day award -10.31 pp | ||
| plug-in mean | 1.85 [1.58, 2.19] | ||
| Allocation | budget-grid program | equal budget split | value 15.7% [13.8, 17.4] |
| Within 20% of budget | Within 10% (pre-specified) | |||||
|---|---|---|---|---|---|---|
| Method | 1 wk | 2 wk | 4 wk | 1 wk | 2 wk | 4 wk |
| Long RCT, 18 weeks (business alternative) | 0.113 | 0.130 | ||||
| Same-calendar RCT | 0.421 | 0.313 | 0.225 | 0.542 | 0.383 | 0.264 |
| Framework (tilt, exact sum) | 0.205 | 0.159 | 0.121 | 0.252 | 0.186 | 0.136 |
| sampled integration ( ) | 0.245 | 0.197 | 0.163 | 0.286 | 0.223 | 0.189 |
| without history (pilot frequencies) | 0.603 | 0.560 | 0.527 | 0.658 | 0.611 | 0.569 |
| A: integrated study | B: new generator | |||||
|---|---|---|---|---|---|---|
| Pilot | 1 wk | 2 wk | 4 wk | 1 wk | 2 wk | 4 wk |
| Same-calendar RCT | 0.391 | 0.296 | 0.216 | 0.477 | 0.326 | 0.240 |
| Framework | 0.220 | 0.172 | 0.148 | 0.188 | 0.150 | 0.116 |
| Dose interpolation (same information) | 0.227 | 0.177 | 0.154 | 0.210 | 0.171 | 0.131 |
| Independent days (least costly replacement) | 0.311 | 0.292 | 0.296 | 0.198 | 0.158 | 0.124 |
| Budgets that bind (%, all 48 laws) | 72.3 | 75.0 | ||||
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
| Step | Framework | Re-draw per request |
|---|---|---|
| Model fitting (once per origin) a | 221.3 once | 221.3 once |
| Per-day draws, 7 days | 108.7 once | 108.7 |
| Cross-day rank template | 3.5 once | reused |
| Cross-day coupling | 1.6 once | 1.8 |
| Window reduction | 0.54 | 0.54 |
| Integration, 540 plans | 2.52 | 2.52 |
| Stage | Shortcut | Accuracy cost of the shortcut | Time the shortcut saves |
|---|---|---|---|
| Prediction | point forecast | plan-cost error 3.08 [2.48, 3.81] | not timed separately |
| Integration | independent days | plan-cost error 1.51 [1.33, 1.74] | 1.4% of sampling; integration 1.001 |
| Integration | simplified normal | plan-cost error 2.46 [2.02, 3.01]; 5–6-day plans 3.29 | integration 85.8 (3.86 s per table) |
| Allocation | equal budget split | value 15.7% [13.8, 17.4] | 0.51 0.01 ms |
| Origins / | Riders | Plan-cost error | Within | |||
| Condition | runs | per run | Held-out | Reference | Change | margin |
| Unseen cities (four folds) | 45 / 180 | 311 | 0.343 | 0.340 | 0.003 [ 0.018, 0.020] | yes |
| Replacement | Endpoint | Fit once | Unseen cities |
|---|---|---|---|
| Point forecast | plan-cost error | 0.416 [ 0.359, 0.472] | 0.380 [ 0.292, 0.467] |
| Plug-in integration | plan-cost error | 0.317 [ 0.235, 0.382] | 0.123 [ 0.097, 0.148] |
| Two-moment normal | plan-cost error | 0.459 [ 0.382, 0.523] | 0.205 [ 0.167, 0.240] |
| Independent days | plan-cost error | 0.202 [ 0.142, 0.250] | 0.058 [ 0.035, 0.080] |
| Equal-width buckets | regret | 0.0017 [ 0.0006, 0.0028] | 0.0046 [ 0.0024, 0.0070] |
| Equal budget split | regret | 0.048 [ 0.021, 0.080] | 0.122 [ 0.099, 0.145] |
| Replacement | 1 wk | 2 wk | 4 wk |
|---|---|---|---|
| Without moment matching (history only) | 0.656 | 0.656 | 0.656 |
| Without joint integration (plug-in mean) | 0.583 | 0.599 | 0.616 |
| Without budget allocation (equal split, exact tables) | 0.861 | 0.860 | 0.855 |
| A Pilot length | 1 week | 2 weeks | 4 weeks |
|---|---|---|---|
| Synthetic term / | 5.92 / 5.53 | 5.92 / 5.53 | 5.92 / 5.53 |
| Moment term / | 8.82 / 7.67 | 6.29 / 5.31 | 4.48 / 3.69 |
| Integration term / ( ) | 3.43 / 3.00 | 3.44 / 3.00 | 3.51 / 3.00 |
| Sum of stage terms | 18.16 / 16.20 | 15.65 / 13.83 | 13.90 / 12.22 |
| Direct error / (truth vs. final tables) | 11.79 / 10.88 | 10.06 / 9.35 | 8.89 / 8.24 |
| Requests with a comparator in | 26/6336 | 140/6336 | 499/6336 |
| Within 20% of budget | Within 10% | |||||
|---|---|---|---|---|---|---|
| Method | 1 wk | 2 wk | 4 wk | 1 wk | 2 wk | 4 wk |
| Same-calendar RCT | 0.391 | 0.296 | 0.216 | 0.527 | 0.388 | 0.269 |
| Framework | 0.220 | 0.172 | 0.148 | 0.258 | 0.198 | 0.164 |
| without moment matching | 0.592 | 0.592 | 0.592 | 0.626 | 0.626 | 0.626 |
| without history | 0.471 | 0.433 | 0.404 | 0.561 | 0.509 | 0.470 |
| plug-in integration | 0.592 | 0.604 | 0.631 | 0.669 | 0.672 | 0.690 |
| 1-week pilot | 2-week pilot | |||||
| Level | Laws | Framework | minus same-calendar | Framework | minus same-calendar | |
| Rule | fixed award | 36 | 0.225 | 0.175 [ 0.208, 0.142] | 0.176 | 0.127 [ 0.166, 0.089] |
| per completion | 36 | 0.229 | 0.172 [ 0.206, 0.138] | 0.180 | 0.125 [ 0.165, 0.087] | |
| attendance | 36 | 0.205 | 0.165 [ 0.192, 0.139] | 0.162 | 0.120 [ 0.151, 0.089] | |
| Window days | 1–2 | 36 | 0.221 | 0.193 [ 0.226, 0.159] | 0.173 | 0.146 [ 0.184, 0.110] |
| 3–5 | 36 | 0.219 | 0.171 [ 0.200, 0.141] | 0.171 | 0.123 [ 0.159, 0.088] | |
| Pilot | Preparation | Request (reuse) | Request (re-run) | |||
|---|---|---|---|---|---|---|
| 1 week | 7.27 | 0.22 | 7.57 | 5.8 | 11.9 | 14.1 [13.8, 14.5] |
| 2 weeks | 7.40 | 0.22 | 7.57 | 5.8 | 11.9 | 14.2 [13.8, 14.5] |
| 4 weeks | 6.25 | 0.21 | 6.61 | 5.9 | 12.0 | 14.3 [14.0, 14.7] |
| Within 20% of budget | Within 10% | |||||
|---|---|---|---|---|---|---|
| Method | 1 wk | 2 wk | 4 wk | 1 wk | 2 wk | 4 wk |
| Same-calendar RCT | 0.477 | 0.326 | 0.240 | 0.765 | 0.601 | 0.415 |
| Framework | 0.188 | 0.150 | 0.116 | 0.270 | 0.197 | 0.140 |
| without moment matching | 0.433 | 0.433 | 0.433 | 0.781 | 0.781 | 0.781 |
| without history | 0.415 | 0.354 | 0.309 | 0.610 | 0.517 | 0.448 |
| plug-in integration | 0.447 | 0.479 | 0.510 | 0.576 | 0.577 | 0.591 |
| 1-week pilot | 4-week pilot | |||||
|---|---|---|---|---|---|---|
| Level | Laws | Framework | minus same-calendar | Framework | minus same-calendar | |
| Rule | fixed award | 36 | 0.179 | 0.318 [ 0.343, 0.293] | 0.120 | 0.127 [ 0.143, 0.112] |
| per completion | 36 | 0.196 | 0.322 [ 0.343, 0.300] | 0.111 | 0.127 [ 0.144, 0.111] | |
| attendance | 36 | 0.202 | 0.183 [ 0.213, 0.155] | 0.132 | 0.114 [ 0.137, 0.093] | |
| Window days | 1–2 | 36 | 0.186 | 0.340 [ 0.365, 0.315] | 0.119 | 0.154 [ 0.175, 0.133] |
| 3–5 | 36 | 0.188 | 0.285 [ 0.308, 0.263] | 0.116 | 0.121 [ 0.138, 0.106] | |
| Condition | Method | 1 wk | 2 wk | 4 wk |
| Independent | Framework | 0.196 | 0.154 | 0.126 |
| Dose interpolation (same information) | 0.219 | 0.171 | 0.138 | |
| Normal law, exact truncated moments | 0.250 | 0.208 | 0.179 | |
| History-anchored tilt of the trial | 0.405 | 0.291 | 0.211 | |
| Same-calendar RCT | 0.480 | 0.340 | 0.241 | |
| Calendar | Framework , concurrent anchor | 0.223 | 0.215 | 0.207 |
| Method | 1 wk | 2 wk | 4 wk |
|---|---|---|---|
| Framework (pilot at three offers) | 0.198 | 0.154 | 0.123 |
| Joint fit of the trial, framework’s dose curve | 0.253 | 0.206 | 0.183 |
| History-anchored tilt of the trial, per offer | 0.405 | 0.289 | 0.211 |
| Same-calendar RCT | 0.479 | 0.337 | 0.241 |
| Placement share | 0.265 [0.252, 0.279] | 0.382 [0.361, 0.405] | 0.683 [0.623, 0.745] |
| Method | Below | Above | Excursion | Valid-plan loss |
|---|---|---|---|---|
| Framework | 0.034 | 0.041 | 0.035 | 0.136 |
| Joint fit of the trial | 0.065 | 0.044 | 0.046 | 0.165 |
| Dose interpolation (same information) | 0.044 | 0.045 | 0.041 | 0.149 |
| History-anchored tilt of the trial, per offer | 0.179 | 0.010 | 0.070 | 0.266 |
| Same-calendar RCT | 0.272 | 0.010 | 0.082 | 0.288 |
| Method | 1 wk | 2 wk | 4 wk | Selected, 1 wk |
|---|---|---|---|---|
| Request-aware reuse | 0.380 | 0.329 | 0.297 | 0.395 |
| Direct reuse of the full-week calibration | 0.576 | 0.584 | 0.596 | 0.742 |
| Layers per group | All laws | Both channels | Intensity | Threshold |
|---|---|---|---|---|
| 1 (group-level plan) | 0.825 [0.820, 0.831] | 0.938 | 0.944 | 0.594 |
| 2 | 0.900 [0.897, 0.904] | 0.966 | 0.971 | 0.764 |
| 3 | 0.926 [0.923, 0.928] | 0.974 | 0.978 | 0.825 |
| 4 | 0.938 [0.936, 0.940] | 0.977 | 0.982 | 0.856 |
| 5 | 0.946 [0.944, 0.948] | 0.979 | 0.984 | 0.875 |
| Stage | Module | Input output | Delivery instance |
|---|---|---|---|
| Predict | 1. Response ranking (Rank) | pre-decision context and past offers ranking by incremental activity | doubly robust learner with a 5% no-offer group ( Dudík et al., 2011 ) |
| 2. Roster (Recall) | ranking and eligibility population fixed before attendance, incl. zero-activity riders | at least one delivery in the preceding 182 days | |
| 3. Baseline forecast (Forecast) | histories and context no-offer outcome law for the requested window | recurrent network, 60-day history, five-day zero-inflated Beta shares | |
| Reduce | 4. Grouping (Activity layers) | baseline capacity groups and non-dominated plan menus | ordered buckets merged to balance dispersion and spending stability |
| Integrate | 5. Response correction (Correct) | no-offer law and pilot or context-matched outcomes offer-conditioned joint trajectories | exponential tilt, Eq. ( 5 ); residual matching on incentive, time and weather |
| 6. Settlement (Score) | trajectories, payment rule, attainable rewards cost, award probability, score and reward witness | Eqs. ( 6 )–( 7 ) |
| Held-out condition | Change in plan-cost error | Change in regret | Met |
| Unseen period (fit once, 16 origins) | 0.051 [ 0.002, 0.098] | 0.000 [ 0.000, 0.000] | no |
| Unseen cities (four folds, 45 origins) | 0.003 [ 0.018, 0.020] | 0.000 [ 0.000, 0.000] | yes |
| Pilot rule | 1-week pilot | 2-week pilot | 4-week pilot |
| Near-lossless vs. 18-week trial (upper ) | 0.156 [ 0.114, 0.201] (no) | 0.093 [ 0.061, 0.126] (no) | 0.059 [ 0.039, 0.077] (no) |
| Beats same-calendar trial (upper ) | 0.256 [ 0.290, 0.222] (yes) | 0.160 [ 0.186, 0.134] (yes) | 0.076 [ 0.109, 0.043] (yes) |
| Without history is worse (lower ) | 0.371 [ 0.266, 0.488] (yes) | 0.388 [ 0.261, 0.525] (yes) | 0.380 [ 0.258, 0.514] (yes) |