End-to-end planners based on waypoint regression achieve strong open-loop accuracy, but they primarily learn to mimic expert geometry and remain difficult to adapt to deployment-time safety constraints. We propose a query-based cost-learning framework that estimates bounded costs for dynamically reachable ego trajectory queries, rather than dense BEV cells or a small regressed trajectory set. Compact joint scene tokens capture coherent multimodal agent futures, while contingency-aware cost aggregation and cost-guided intra-cluster MPPI mixing convert the learned cost topology into feasible ego plans. On nuScenes, our method improves over prior cost-estimation planners such as ST-P3 and NMP, outperforms most regression baselines in collision rate, while remaining competitive in L2, and retaining an interpretable cost interface. On real-world driving logs, the proposed planner reduces collision rates compared with SparseDrive and Alpamayo without fine-tuning, while maintaining a diverse set of candidate trajectories.
Figures & tables
Figure 1 : Multi-view images feed a sparse perception encoder with temporal instance memory. The proposed planner combines joint scene representations and reachable ego queries to estimate bounded costs and refine ego plans through intra-cluster mixing.
Figure 2 : Detailed view of the proposed planner. Joint scene tokens encode multimodal futures, while the ego branch samples and encodes reachable trajectory queries. A lightweight head predicts scene-conditioned costs with contingency aggregation, optional inference-time external-cost repair, and cost-guided intra-cluster mixing.
Method
L2 (m) ↓
Collision (%) ↓
1s
2s
3s
Avg.
1s
2s
3s
Avg.
Cost-based methods
NMP [ 31 ] †
–
–
2.31
–
–
–
1.92
–
SA-NMP [ 31 ] †
–
–
2.05
–
–
–
1.59
–
ST-P3 [ 9 ]
1.33
2.11
2.90
2.11
0.23
0.62
1.27
0.71
Ours
0.28
0.65
1.21
0.71
0.00
0.02
0.18
0.07
Table 1 : Open-loop planning performance on nuScenes under the UniAD metrics. † denotes LiDAR-based methods.
Method
L2 (m) ↓
Collision (%) ↓
Diversity (m) ↑
Params.
Latency (ms) ↓
1s
2s
3s
Avg.
1s
2s
3s
Avg.
Avg.
Endpoint
SparseDrive-S [ 24 ]
2.38
4.17
6.11
4.22
0.53
2.57
5.08
2.73
1.77
3.41
85.9M
115
Alpamayo 1.5 [ 27 ]
0.64
1.41
2.39
1.48
0.03
0.90
1.79
0.87
1.47
3.25
10.0B
1878
Ours
0.85
2.12
3.98
2.31
0.01
0.40
1.50
0.68
5.54
11.88
86.2M
104
Table 2 : Zero-shot real-world transfer. No method is fine-tuned on the real-world logs. Latency is measured on a workstation with a RTX 3090 GPU.
Collision supervision
L2 (m) ↓
Collision (%) ↓
1s
2s
3s
Avg.
1s
2s
3s
Avg.
Ground-truth collision
0.27
0.63
1.16
0.69
0.00
0.06
0.25
0.10
Top-1 predicted collision
0.25
0.59
1.10
0.64
0.02
0.07
0.26
0.12
Contingency collision risk
0.28
0.65
1.21
0.71
0.00
0.02
0.18
0.07
Table 3 : Ablation of collision supervision for the safety margin on nuScenes.
Objective Lrank
nuScenes
Real-world
L2 ↓
Coll. ↓
Cost AUC ↑
Cost Gap ↑
L2 ↓
Coll. ↓
Cost AUC ↑
Cost Gap ↑
NMP-style loss [ 31 ]
0.76
0.33
0.70
6.60
2.29
0.84
0.59
3.99
ST-P3-style loss [ 9 ]
0.61
0.12
0.79
7.80
3.26
0.83
0.61
4.24
Proposed safety margin
0.71
0.07
0.86
45.70
2.31
0.68
0.63
16.18
Table 4 : Comparison of cost-supervision objectives.
Margin source
nuScenes
Real-world
L2 ↓
Coll. ↓
Cost AUC ↑
Cost Gap ↑
L2 ↓
Coll. ↓
Cost AUC ↑
Cost Gap ↑
Distance
0.72
0.20
0.71
10.51
2.36
1.71
0.56
4.89
Collision + off-road
0.88
0.11
0.87
51.25
2.34
1.48
0.66
23.57
All
0.71
0.07
0.86
45.70
2.31
0.68
0.63
16.18
Table 5 : Ablation of safety-margin sources.
Figure 3 : Qualitative comparison using the conflict area of a red traffic light received from a V2X unit. SparseDrive’s low-diversity proposals already violate the rule, so post-hoc rescoring cannot recover a rule-compliant trajectory. Our planner applies the external penalty to the aggregated query-cost profile, producing safe low-cost candidates.
Figure 4 : Qualitative example under a temporally premature right-turn command. Panel (a) shows the final selected trajectory candidates projected into the camera views, with ours in green and SparseDrive in red. SparseDrive follows the premature command into the lane divider, while our planner selects the locally safe left bend.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5 : Command-conditioned reachable query sampling. The kinematic sampler generates clustered, dynamically executable ego trajectories that cover the diverse intentions implied by the navigation command while preserving local control variations within each cluster.
Tb (s)
L2 (m) ↓
Collision (%) ↓
1s
2s
3s
Avg.
1s
2s
3s
Avg.
0
0.33
0.80
1.50
0.88
0.02
0.14
0.35
0.17
1
0.28
0.65
1.21
0.71
0.00
0.02
0.18
0.07
2
0.30
0.69
1.28
0.76
0.02
0.07
0.29
0.13
3
0.32
0.78
1.47
0.86
0.03
0.16
0.37
0.19
Appendix
Table 6 : Ablation of Branching-time on nuScenes.
Variant
L2 (m) ↓
Collision (%) ↓
1s
2s
3s
Avg.
1s
2s
3s
Avg.
Global selection
0.29
0.70
1.30
0.77
0.02
0.07
0.31
0.13
Best per cluster
0.26
0.63
1.14
0.68
0.01
0.05
0.28
0.11
MPPI
0.28
0.65
1.21
0.71
0.00
0.02
0.18
0.07
MPPI + rescoring
0.28
0.67
1.22
0.73
0.01
0.12
0.31
0.15
Appendix
Table 7 : Ablation of Plan selection and rescoring on nuScenes.
Variant
L2 (m) ↓
Collision (%) ↓
Cost AUC ↑
Cost Gap ↑
1s
2s
3s
Avg.
1s
2s
3s
Avg.
Safety margin
0.35
0.89
1.72
0.99
0.00
0.151
0.41
0.19
0.81
36.17
+ candidate classification
0.34
0.82
1.54
0.90
0.04
0.19
0.44
0.22
0.81
38.64
+ candidate regression
0.38
1.00
1.91
1.10
0.04
0.11
0.31
0.15
0.82
35.76
+ classification + regression
0.28
0.65
1.21
0.71
0.00
0.02
0.18
0.07
0.86
45.70
Appendix
Table 8 : Ablation of planning loss components on nuScenes.
Collision signal
Negative agg.
L2 (m) ↓
Collision (%) ↓
Cost AUC ↑
Cost Gap ↑
1s
2s
3s
Avg.
1s
2s
3s
Avg.
Ground-truth collision
Max
0.29
0.69
1.30
0.76
0.01
0.06
0.27
0.11
0.85
41.67
Ground-truth collision
Top- k
0.27
0.63
1.16
0.69
0.00
0.06
0.25
0.10
0.86
46.66
Top-1 predicted collision
Max
0.27
0.62
1.16
0.68
0.00
0.04
0.19
0.08
0.88
42.56
Top-1 predicted collision
Top- k
0.27
0.61
1.12
0.67
0.04
0.12
0.23
0.13
0.87
45.71
Contingency collision risk
Max
0.28
0.66
1.23
0.72
0.01
0.10
0.28
0.13
0.84
41.22
Appendix
Table 9 : Ablation of hard-negative aggregation for safety-margin training on nuScenes.
Mixing score
L2 (m) ↓
Collision (%) ↓
Cost AUC ↑
Cost Gap ↑
1s
2s
3s
Avg.
1s
2s
3s
Avg.
Mean trajectory cost
0.28
0.65
1.21
0.71
0.00
0.02
0.18
0.07
0.86
45.70
Per-timestep cost
0.30
0.71
1.32
0.77
0.02
0.10
0.30
0.14
0.86
42.99
Appendix
Table 10 : Ablation of cost profiles for cost-guided cluster mixing on nuScenes.
Aggregation
L2 (m) ↓
Collision (%) ↓
Cost AUC ↑
Cost Gap ↑
1s
2s
3s
Avg.
1s
2s
3s
Avg.
Top-1
0.33
0.79
1.50
0.87
0.03
0.17
0.37
0.19
0.81
33.99
Worst Case
0.32
0.78
1.47
0.86
0.03
0.16
0.37
0.19
0.81
37.81
Mean
0.32
0.77
1.45
0.85
0.00
0.09
0.30
0.13
0.82
38.97
Probability-weighted
0.33
0.80
1.50
0.88
0.02
0.14
0.35
0.17
0.81
37.97
Contingency
0.28
0.65
1.21
0.71
0.00
0.02
0.18
0.07
0.86
45.70
Appendix
Table 11 : Ablation of scene-cost aggregation on nuScenes.
Figure 6 : Qualitative nuScenes examples. The camera views and BEV visualizations show representative interactive scenes in which the learned cost topology guides selection among reachable ego candidates.
Figure 7 : Zero-shot qualitative examples on real-world driving logs. Each row shows surround-view camera predictions, the BEV cost visualization with the selected trajectory, and ground-truth boxes with ego trajectory and semantic map support.
Learning-based driving planners are usually trained and evaluated in open loop against logged trajectories. In closed loop, a trajectory with small displacement error can still stall the vehicle, steer it into a conflict with surrounding agents, or be executed with abrupt braking. We introduce Closed-Loop Refinement and Execution (CLRE), a hierarchical receding-horizon control framework designed to mitigate these failure modes while leaving the upstream planner frozen and adding no new learned model. The upper layer treats the nominal trajectory as a reference and solves a finite-horizon optimal control problem that trades route progress against interaction with predicted agents. Solving it from several initializations gives a candidate set, and a prediction-conditioned oriented-bounding-box (OBB) feasibility test retains only candidates whose minimum predicted OBB clearance over the horizon meets a threshold. The lower layer executes the lowest-cost survivor, or a route-centerline backup when none remains, through the tracking controller supplied with the planner, augmented by a range-based speed bound and a saturated proportional braking law. In closed-loop simulation on 126 Bench2Drive routes with VAD as the upstream planner, CLRE raises the driving score from 43.41 to 56.42 and route completion from 57.27 to 72.23, and reduces collision events from 70 to 53.
Huaijin Hu, Shanting Wang, Zhongyu Mo +1
Systems Engineering Program, Cornell University, Ithaca, NY 14853 USA
Safe and explainable motion planning remains a central challenge in autonomous driving. While rule-based planners offer predictable and explainable behavior, they often fail to grasp the complexity and uncertainty of real-world traffic. Conversely, learned planners exhibit strong adaptability but suffer from reduced transparency and occasional safety violations. We introduce Mosaic, a framework for structured decision-making that integrates both paradigms through arbitration graphs. By decoupling trajectory verification and selection from the generation of trajectories by individual planners, every decision becomes transparent and traceable. This separation lets verification and trajectory selection contribute independently: centralized verification acts as a safety floor, reducing at-fault collisions from 25 for each standalone planner to 16. In contrast, per-step trajectory selection acts as a performance ceiling, combining the complementary strengths of a rule-based and a learned planner. In experimental evaluation on nuPlan, Mosaic achieves 95.56 CLS-NR and 94.18 CLS-R on the Val14 closed-loop benchmark, setting a new state of the art. On the interPlan benchmark, focused on highly interactive and out-of-distribution scenarios, Mosaic scores 54.10 CLS-R, outperforming its best constituent planner by 22.8% -- all without retraining or requiring additional data. The code is available at github.com/KIT-MRT/mosaic.
Nick Le Large, Marlon Steiner, Lingguang Wang +4
Institute of Measurement and Control Systems, Karlsruhe Institute of Technology (KIT), Karlsruhe, Germany · FZI Research Center for Information Technology, Karlsruhe, Germany
End-to-end autonomous driving planners typically generate trajectories from current observations alone. However, real-world driving is highly dynamic, and such reactive planning cannot anticipate future scene evolution, often leading to myopic decisions and safety-critical failures. We propose ProDrive, a world-model-based proactive planning framework that enables ego-environment co-evolution for autonomous driving. ProDrive jointly trains a query-centric trajectory planner and a bird's-eye-view (BEV) world model end-to-end: the planner generates diverse candidate trajectories and planning-aware ego tokens, while the world model predicts future scene evolution conditioned on them. By injecting planner features into the world model and evaluating all candidates in parallel, ProDrive preserves end-to-end gradient flow and allows future outcome assessment to directly shape planning. This bidirectional coupling enables proactive planning beyond current-observation-driven decision-making. Experiments on NAVSIM v1 show that ProDrive outperforms strong baselines in both safety and planning efficiency, while ablations validate the effectiveness of the proposed ego-environment coupling design.
Chuyao Fu, Shengzhe Gan, Zhuoli Ouyang +5
1Southern University of Science and Technology · 2Hong Kong University of Science and Technology