Depot-Closed Multi-Component Construction for Neural Vehicle Routing
Organizations: Core Technology, R&D Division Panasonic Connect Co., Ltd. · Graduate School of Informatics Kyoto University
Abstract
Most neural constructive solvers for the vehicle routing problem (VRP) use route-by-route construction, extending one route until completion before starting the next. This commits route membership early and hinders global coordination across routes. We propose multi-component construction, which maintains many route components simultaneously and merges them in an arbitrary order. This removes the depot-return cue that route-by-route construction obtains from the remaining capacity; to compensate, we introduce an interpretation in which every component is treated as an implicitly depot-closed route. Under this depot-closed interpretation, every intermediate state of standard CVRP construction is a complete feasible solution, and the exact cost reduction of a merge is the Clarke-Wright saving. The neural policy combines this CW-saving signal with the evolving component state to learn what to connect and when to connect. A policy trained only on CVRP100 outperforms the reported results of representative neural solvers on CVRP100-500 with greedy inference and, reused for ruin-and-reconstruct, performs strongly at all evaluated sizes up to CVRP1000. In a zero-shot Constraint Tightness evaluation with capacities from to , it outperforms the reported neural solvers at every capacity. Controlled analyses show that robustness persists without CW grounding and point to learned route-closing behavior as a plausible contributor to the tight-regime degradation of learned route-by-route solvers.
Figures & tables
| Method | CVRP100 | CVRP200 | CVRP500 | CVRP1000 | Mean | Source |
|---|---|---|---|---|---|---|
| BQ | 2.726 | 2.972 | 3.248 | 5.892 | 3.710 | ReLD Table 3 / BQ original |
| LEHD | 3.648 | 3.312 | 3.178 | 4.912 | 3.763 | LEHD Table 1; ReLD Table 3 |
| MnLP | 3.509 | 3.206 | 2.928 | 6.195 | 3.960 | MnLP Table 1 |
| INViT | 9.964 | 12.160 | 13.772 | 15.548 | 12.861 | MnLP Table 1 |
| DGL | 6.441 | 8.368 | 9.356 | 16.727 | 10.223 | MnLP Table 1 |
| Classical CW | 5.479 | 7.310 | 6.525 | 11.999 | 7.828 | This work |
| Method | Inference | CVRP100 | CVRP200 | CVRP500 | CVRP1000 | Mean | Source |
| MDAM | bs50 | 2.211 | 4.304 | 10.498 | 27.814 | 11.207 | ReLD Table 3 |
| POMO | aug 8 | 1.004 | 3.403 | 11.135 | 110.632 | 31.544 | ReLD Table 3 |
| ELG | aug 8 | 1.207 | 2.553 | 5.472 | 10.760 | 4.998 | ReLD Table 3 |
| ReLD | aug 8 | 0.960 | 1.654 | 2.975 | 6.757 | 3.087 | ReLD Table 3 |
| BQ | bs16 | 0.611 | 1.141 | 2.991 | 7.784 | 3.132 | LEHD Table 1 |
| BQ | bs16, later rerun | 1.020 | 0.940 | 1.010 | 2.880 | 1.463 | DRHG Table 3 |
| Stage | Subset | 100 | 200 | 500 | 1000 | Mean |
|---|---|---|---|---|---|---|
| S2 | No | 1.217 | 1.833 | 3.615 | 9.952 | 4.154 |
| S2 | Yes | 1.383 | 1.589 | 2.049 | 7.495 | 3.129 |
| S3 | No | 0.889 | 1.159 | 2.746 | 8.057 | 3.213 |
| S3 | Yes | 1.019 | 0.828 | 1.407 | 7.544 | 2.700 |
| Stage | Subset | 100 | 200 | 500 | 1000 | Mean |
|---|---|---|---|---|---|---|
| S2 | No | 1.217 | 1.833 | 3.615 | 9.952 | 4.154 |
| S2 | Yes | 1.383 | 1.589 | 2.049 | 7.495 | 3.129 |
| S3 | No | 0.889 | 1.159 | 2.746 | 8.057 | 3.213 |
| S3 | Yes | 1.019 | 0.828 | 1.407 | 7.544 | 2.700 |
| Training | Inference | Gap |
|---|---|---|
| no CW grounding | no CW bias | 5.822 |
| no CW grounding | +CW bias | 3.883 |
| CW fine-tuning | +CW bias | 3.787 |
| Model | Mean | |||||||
|---|---|---|---|---|---|---|---|---|
| AM | 45.16 | 7.66 | 12.85 | 21.63 | 26.26 | 25.35 | 27.23 | 23.73 |
| POMO | 20.26 | 3.66 | 11.17 | 22.45 | 29.56 | 30.80 | 34.12 | 21.72 |
| MDAM | 7.47 | 5.38 | 12.47 | 21.13 | 24.31 | 22.55 | 23.91 | 16.75 |
| BQ | 18.02 | 3.25 | 3.48 | 4.53 | 6.53 | 6.54 | 8.16 | 7.22 |
| LEHD | 36.56 | 4.22 | 4.87 | 4.18 | 5.78 | 4.92 | 6.82 | 9.62 |
| ELG | 10.44 | 5.24 | 10.91 | 18.70 | 22.81 | 21.84 | 24.05 | 16.28 |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Stage | Trajectory; initial state | Objective | Role |
|---|---|---|---|
| S0 | shuffle-guided; construction prefix | set-valued SL | Shuffling suppresses dependence on a particular merge ordering; focuses on valid merge selection (what to connect). |
| S1 | policy-guided; construction prefix | set-valued SL | Learns what to connect on expert-compatible, policy-generated construction trajectories. |
| S2 | policy-guided; ruined solution | set-valued SL | Learns what to connect on expert-compatible, policy-generated repair trajectories. |
| S3 | free policy rollout; ruined solution | terminal route-cost RL | Removes the teacher constraint and jointly learns what and when to connect. |
| Item | Setting |
|---|---|
| Training problem | CVRP100, |
| Training / validation data | 1,000,000 HGS-solved instances / 1,000 instances |
| Embedding dimension | 128 |
| Light encoder | 1 attention layer |
| Heavy decoder | 6 attention layers |
| Attention | 8 heads, Q/K/V dimension 16 per head |
| Setting | S0 | S1 | S2 | S3 |
| Learning objective | set-valued SL | set-valued SL | set-valued SL | REINFORCE |
| State generation | shuffle-guided | policy-guided | policy-guided | free rollout |
| Initial state | constr. prefix | constr. prefix | polar ruin | polar ruin |
| Route-subset sampling | Yes | Yes | Yes | Yes |
| Optimizer | Adam | Adam | Adam | Adam |
| Learning rate |
| Method | Inference | CVRP100 | CVRP200 | CVRP500 | CVRP1000 |
|---|---|---|---|---|---|
| LEHD (rerun) | greedy | 0.0012 | 0.0116 | 0.0628 | 0.3845 |
| S3 | greedy | 0.0024 | 0.0156 | 0.0611 | 0.3115 |
| S3 | RRC1000 | 1.176 | 8.354 | 37.783 | 204.096 |
| Inference | CVRP100 | CVRP200 | CVRP500 | CVRP1000 | Mean |
|---|---|---|---|---|---|
| size-calibrated | 1.019 | 0.828 | 1.407 | 7.544 | 2.700 |
| fixed | 1.019 | 0.828 | 1.358 | 7.694 | 2.725 |
| Method | Inference | CVRP100 | CVRP200 | CVRP500 | CVRP1000 | Mean | Source |
| MDAM | bs50 | 2.211 | 4.304 | 10.498 | 27.814 | 11.207 | ReLD Table 3 |
| POMO | aug 8 | 1.004 | 3.403 | 11.135 | 110.632 | 31.544 | ReLD Table 3 |
| ELG | aug 8 | 1.207 | 2.553 | 5.472 | 10.760 | 4.998 | ReLD Table 3 |
| ReLD | aug 8 | 0.960 | 1.654 | 2.975 | 6.757 | 3.087 | ReLD Table 3 |
| BQ | bs16 | 0.611 | 1.141 | 2.991 | 7.784 | 3.132 | LEHD Table 1 |
| BQ | bs16, later rerun | 1.020 | 0.940 | 1.010 | 2.880 | 1.463 | DRHG Table 3 |