Synergizing Drone Delivery Order Pooling and Road Network Monitoring through Monitoring-Task Orderization
Authors: Yulong Hu, Meng Xu, Sen Li, Nikolas Geroliminis
Organizations: Department of Civil and Environmental Engineering, The Hong Kong University of Science and Technology, Hong Kong, China · School of Architecture, Civil and Environmental Engineering (ENAC), École Polytechnique Fédérale de Lausanne (EPFL), Lausanne, Switzerland
This paper investigates the real-time dispatch of a shared drone fleet for on-demand food delivery and urban road network monitoring. We consider a courier-drone collaborative setting in which couriers transport orders to launchpads and drones complete the final delivery leg to kiosks. Drones may consolidate multiple origin-destination orders within one flight and make monitoring-aware route adjustments to collect real-time traffic information subject to delivery-time constraints. This yields a joint decision problem coupling dynamic order-to-drone matching, multi-order pooling, routing, and time-varying monitoring under fleet-level competition and uncertainty. We propose monitoring-task orderization, which periodically converts road-network nodes with high congestion and stale information into virtual monitoring orders. Pooling these virtual tasks with food-delivery orders creates a unified heterogeneous task set and transforms the coupled matching-and-routing problem into an order-level decision process. Building on this abstraction, we formulate a decentralized graph-interdependent Multi-Agent Markov Decision Process and develop Graph Multi-Agent Q-Learning (Graph-MAQL), which captures localized inter-agent dependencies through bipartite match coordination graphs. Agent-task value estimates are then used as edge weights in a dynamic heterogeneous bipartite matching program for globally feasible execution. Experiments using real-world data reveal strong operational synergy between delivery and monitoring. Monitoring-task orderization improves monitoring performance by 25.1% with less than a 1% reduction in delivery performance, while Graph-MAQL improves the aggregate objective by up to 20.8%, reduces deadline violations by over 40%, and transfers zero-shot to higher demand intensity without retraining.
Figures & tables
Figure 1: An illustrative example of dispatching drone fleets for on-demand delivery pooling and road network monitoring.
Figure 2: Overview of our proposed framework for dispatch of drone fleets for on-demand delivery and road network monitoring.
α
Orders Picked Up
Deadline Violation Rate (%)
Link Coverage Rate (%)
Bottleneck Roads Monitor Frequency
Delivery Score (HKD)
Monitoring Score
Average Occupied Seats
0
4188 ± 12.2
5.0 ± 0.90
63.9 ± 1.07
3.90 ± 0.18
79639 ± 657
5555 ± 145
2.25 ± 0.03
3
4191 ± 12.3
4.3 ± 0.80
69.7 ± 0.69
4.29 ± 0.14
80129 ± 557
5926 ± 97.3
2.31 ± 0.03
6
4189 ± 14.4
6.4 ± 1.16
71.1 ± 1.04
4.59 ± 0.15
79057 ± 774
6269 ± 132.3
2.33 ± 0.04
8
4161 ± 16.2
6.6 ± 1.15
72.3 ± 0.70
4.80 ± 0.13
78129 ± 652
6550 ± 74.3
2.40 ± 0.03
10
4138 ± 13.3
9.0 ± 0.84
72.7 ± 0.61
4.90 ± 0.15
75747 ± 1027
6558 ± 104.3
2.45 ± 0.03
Table 1: Performance under varying monitoring priority weight α without monitoring-task orderization.
k
Orders Picked Up
Deadline Violation Rate (%)
Link Coverage Rate (%)
Bottleneck Roads Monitor Frequency
Delivery Score (HKD)
Monitoring Score
Overall Score (HKD)
0
4191 ± 12.3
4.3 ± 0.80
69.7 ± 0.69
4.29 ± 0.14
80129 ± 557
5926 ± 97
97887 ± 710
25
4203 ± 14.4
5.6 ± 1.19
74.2 ± 0.75
5.13 ± 0.20
80105 ± 1016
6936 ± 111
100894 ± 902
50
4197 ± 15.4
5.9 ± 0.73
75.9 ± 0.63
5.68 ± 0.19
79834 ± 888
7248 ± 78
101556 ± 881
75
4194 ± 15.1
6.3 ± 1.14
75.2 ± 0.53
5.94 ± 0.18
80099 ± 904
7361 ± 70
102159 ± 893
100
4168 ± 15.8
6.8 ± 0.68
75.2 ± 1.00
6.09 ± 0.16
79362 ± 642
7413 ± 78
101576 ± 708
Table 2: Performance under varying numbers k of virtual monitoring orders generated every five minutes, with monitoring-task orderization enabled and α fixed at 3. The row k=0 coincides with the α=3 configuration of Table 1 .
Method
Orders Picked Up
Viol. Rate (%)
Link Cov. (%)
Bttlnk. Freq.
Delivery Score
Monitor. Score
Overall Score
Greedy
3702 ± 15.8
19.0 ± 1.26
76.8 ± 0.79
6.27 ± 0.23
63238 ± 753
7617 ± 81
86059 ± 906
Independent-MAQL
4168 ± 15.8
6.8 ± 0.68
75.2 ± 1.00
6.09 ± 0.16
79362 ± 642
7413 ± 78
101576 ± 708
Graph-MAQL
4199 ± 6.8
3.9 ± 0.39
74.0 ± 0.75
5.84 ± 0.16
82332 ± 225
7231 ± 103
104000 ± 378
Table 3: Comparison of Graph-MAQL against baselines. Both learning-based policies are trained at the 1.0× demand level; panel (b) reports zero-shot evaluation of those same policies under approximately 5,300 orders, with no retraining. The Independent-MAQL row in panel (a) coincides with the k=100 configuration of Table 2 . Abbreviations: Viol. Rate = deadline violation rate; Link Cov. = link coverage rate; Bttlnk. Freq. = bottleneck roads monitoring frequency.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Visualization of the Kowloon food-delivery dataset and road-network speed map.
A central question in deploying teams of mobile robots for persistent monitoring is how task performance scales with fleet size, and whether this scaling holds once sensing drives downstream action rather than mere observation. We study this question for a team of drones performing traffic-jam detection and prediction in a simulated road network, whose reports drive an adaptive traffic-signal controller in closed loop. We build a multi-agent simulation, with vehicles following Nagel-Schreckenberg cellular-automaton dynamics and drones patrolling junctions via a round-robin policy, and sweep fleet size, traffic level, and network size to evaluate detection rate, detection delay, and prediction rate. We show how performance plateaus for fleet size approximating the number of junctions being monitored, and offer a general fleet-provisioning rule for persistent-monitoring deployments. More significantly, adapting the signal on a predicted jam, rather than a detected one, roughly doubles the resulting reduction in jam duration, showing that the value of onboard prediction in a sensing-to-action pipeline can exceed the value of adding more robots. Prediction accuracy, not sensing coverage, is now the binding constraint on further improvement, pointing to onboard inference, not fleet size, as the more promising direction for future work.
Urban last-mile parcel delivery increasingly relies on heterogeneous fleets whose performance depends on timely coordination, reliable communication, and scalable control. Truck-drone collaboration has emerged as a networked cyber-physical delivery paradigm that combines the payload capacity and range efficiency of trucks with the agility of drones in congested or access-limited urban environments. This paper proposes a layered planning and coordination framework that structures truck-drone collaborative delivery (TDCD) from a systems and control perspective. The framework consists of five interrelated layers: spatial-demand alignment, collaborative delivery configuration, resource and workflow orchestration, performance evaluation, and scalability analysis, providing a unified view of coordination, control, and system-level performance in networked delivery operations. The proposed framework is evaluated using a realistic urban last-mile delivery scenario derived from the 2021 Amazon Last Mile Routing Research Challenge dataset. The case study demonstrates how coordinated truck-drone operation, enabled by structured task orchestration and inter-agent synchronization, improves end-to-end system efficiency under operational constraints. Results show a 42.4% reduction in total delivery time and a 44.2% reduction in energy consumption compared to a conventional truck-only delivery model. The scalability analysis further highlights how coordination gains persist as system size increases, and shows the importance of efficient control and communication in heterogeneous delivery networks.
Didem Cicek, Burak Kantarci
School of Electrical Engineering and Computer Science at the University of Ottawa, Ottawa, ON, K1N 6N5, Canada
Truck-drone delivery is an emerging last-mile logistics mode combining the long-haul capacity of trucks with the flexible service capability of drones. In locker-based operations, smart lockers serve not only as temporary parcel storage facilities but also as automated drone docking and service nodes. These automated nodes support drone takeoff, landing, parcel handover, and battery replacement, thereby significantly extending the service range and operational flexibility of drone-assisted delivery networks. However, practical locker-based delivery systems face complex real-world challenges, requiring the integrated coordination of not only parcel delivery, return pickup, battery-constrained and load-dependent drone flights, but also necessary detours around restricted airspace. To address this practical and multifaceted challenge, this paper introduces a locker-based truck-drone routing problem with integrated considerations of pickups, deliveries, and no-fly zones (LTDRP-PDNF), with the objective of minimizing the total operational cost of a fleet of drone-equipped trucks. We formulate the route construction process as a Markov Decision Process and develop a two-stage deep reinforcement learning-based neural heuristic. The first stage utilizes an attention-based encoder and a Bidirectional Gated Recurrent Unit decoder to solve the truck-only routing problem, formulated as a capacitated vehicle routing problem. The second stage combines a policy-transfer strategy with a hybrid dispatch assignment heuristic to construct fully coordinated truck and drone routes for LTDRP-PDNF. Experiments on instances of different scales demonstrate that the proposed method outperforms metaheuristic and neural heuristic baselines in most cases while maintaining exceptionally short computation times, offering an effective, scalable solution framework under practical operational constraints.
Xuanyu Liu, Hui Hu, Jiao Zhao +2
College of Transportation Engineering, Chang’an University, China · Faculty of Science and Technology, University of Nottingham Ningbo China