End-to-end planners based on waypoint regression achieve strong open-loop accuracy, but they primarily learn to mimic expert geometry and remain difficult to adapt to deployment-time safety constraints. We propose a query-based cost-learning framework that estimates bounded costs for dynamically reachable ego trajectory queries, rather than dense BEV cells or a small regressed trajectory set. Compact joint scene tokens capture coherent multimodal agent futures, while contingency-aware cost aggregation and cost-guided intra-cluster MPPI mixing convert the learned cost topology into feasible ego plans. On nuScenes, our method improves over prior cost-estimation planners such as ST-P3 and NMP, outperforms most regression baselines in collision rate, while remaining competitive in L2, and retaining an interpretable cost interface. On real-world driving logs, the proposed planner reduces collision rates compared with SparseDrive and Alpamayo without fine-tuning, while maintaining a diverse set of candidate trajectories.
Figures & tables
Figure 1 : Multi-view images feed a sparse perception encoder with temporal instance memory. The proposed planner combines joint scene representations and reachable ego queries to estimate bounded costs and refine ego plans through intra-cluster mixing.
Figure 2 : Detailed view of the proposed planner. Joint scene tokens encode multimodal futures, while the ego branch samples and encodes reachable trajectory queries. A lightweight head predicts scene-conditioned costs with contingency aggregation, optional inference-time external-cost repair, and cost-guided intra-cluster mixing.
Method
L2 (m) ↓
Collision (%) ↓
1s
2s
3s
Avg.
1s
2s
3s
Avg.
Cost-based methods
NMP [ 31 ] †
–
–
2.31
–
–
–
1.92
–
SA-NMP [ 31 ] †
–
–
2.05
–
–
–
1.59
–
ST-P3 [ 9 ]
1.33
2.11
2.90
2.11
0.23
0.62
1.27
0.71
Ours
0.28
0.65
1.21
0.71
0.00
0.02
0.18
0.07
Table 1 : Open-loop planning performance on nuScenes under the UniAD metrics. † denotes LiDAR-based methods.
Method
L2 (m) ↓
Collision (%) ↓
Diversity (m) ↑
Params.
Latency (ms) ↓
1s
2s
3s
Avg.
1s
2s
3s
Avg.
Avg.
Endpoint
SparseDrive-S [ 24 ]
2.38
4.17
6.11
4.22
0.53
2.57
5.08
2.73
1.77
3.41
85.9M
115
Alpamayo 1.5 [ 27 ]
0.64
1.41
2.39
1.48
0.03
0.90
1.79
0.87
1.47
3.25
10.0B
1878
Ours
0.85
2.12
3.98
2.31
0.01
0.40
1.50
0.68
5.54
11.88
86.2M
104
Table 2 : Zero-shot real-world transfer. No method is fine-tuned on the real-world logs. Latency is measured on a workstation with a RTX 3090 GPU.
Collision supervision
L2 (m) ↓
Collision (%) ↓
1s
2s
3s
Avg.
1s
2s
3s
Avg.
Ground-truth collision
0.27
0.63
1.16
0.69
0.00
0.06
0.25
0.10
Top-1 predicted collision
0.25
0.59
1.10
0.64
0.02
0.07
0.26
0.12
Contingency collision risk
0.28
0.65
1.21
0.71
0.00
0.02
0.18
0.07
Table 3 : Ablation of collision supervision for the safety margin on nuScenes.
Objective Lrank
nuScenes
Real-world
L2 ↓
Coll. ↓
Cost AUC ↑
Cost Gap ↑
L2 ↓
Coll. ↓
Cost AUC ↑
Cost Gap ↑
NMP-style loss [ 31 ]
0.76
0.33
0.70
6.60
2.29
0.84
0.59
3.99
ST-P3-style loss [ 9 ]
0.61
0.12
0.79
7.80
3.26
0.83
0.61
4.24
Proposed safety margin
0.71
0.07
0.86
45.70
2.31
0.68
0.63
16.18
Table 4 : Comparison of cost-supervision objectives.
Margin source
nuScenes
Real-world
L2 ↓
Coll. ↓
Cost AUC ↑
Cost Gap ↑
L2 ↓
Coll. ↓
Cost AUC ↑
Cost Gap ↑
Distance
0.72
0.20
0.71
10.51
2.36
1.71
0.56
4.89
Collision + off-road
0.88
0.11
0.87
51.25
2.34
1.48
0.66
23.57
All
0.71
0.07
0.86
45.70
2.31
0.68
0.63
16.18
Table 5 : Ablation of safety-margin sources.
Figure 3 : Qualitative comparison using the conflict area of a red traffic light received from a V2X unit. SparseDrive’s low-diversity proposals already violate the rule, so post-hoc rescoring cannot recover a rule-compliant trajectory. Our planner applies the external penalty to the aggregated query-cost profile, producing safe low-cost candidates.
Figure 4 : Qualitative example under a temporally premature right-turn command. Panel (a) shows the final selected trajectory candidates projected into the camera views, with ours in green and SparseDrive in red. SparseDrive follows the premature command into the lane divider, while our planner selects the locally safe left bend.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5 : Command-conditioned reachable query sampling. The kinematic sampler generates clustered, dynamically executable ego trajectories that cover the diverse intentions implied by the navigation command while preserving local control variations within each cluster.
Tb (s)
L2 (m) ↓
Collision (%) ↓
1s
2s
3s
Avg.
1s
2s
3s
Avg.
0
0.33
0.80
1.50
0.88
0.02
0.14
0.35
0.17
1
0.28
0.65
1.21
0.71
0.00
0.02
0.18
0.07
2
0.30
0.69
1.28
0.76
0.02
0.07
0.29
0.13
3
0.32
0.78
1.47
0.86
0.03
0.16
0.37
0.19
Appendix
Table 6 : Ablation of Branching-time on nuScenes.
Variant
L2 (m) ↓
Collision (%) ↓
1s
2s
3s
Avg.
1s
2s
3s
Avg.
Global selection
0.29
0.70
1.30
0.77
0.02
0.07
0.31
0.13
Best per cluster
0.26
0.63
1.14
0.68
0.01
0.05
0.28
0.11
MPPI
0.28
0.65
1.21
0.71
0.00
0.02
0.18
0.07
MPPI + rescoring
0.28
0.67
1.22
0.73
0.01
0.12
0.31
0.15
Appendix
Table 7 : Ablation of Plan selection and rescoring on nuScenes.
Variant
L2 (m) ↓
Collision (%) ↓
Cost AUC ↑
Cost Gap ↑
1s
2s
3s
Avg.
1s
2s
3s
Avg.
Safety margin
0.35
0.89
1.72
0.99
0.00
0.151
0.41
0.19
0.81
36.17
+ candidate classification
0.34
0.82
1.54
0.90
0.04
0.19
0.44
0.22
0.81
38.64
+ candidate regression
0.38
1.00
1.91
1.10
0.04
0.11
0.31
0.15
0.82
35.76
+ classification + regression
0.28
0.65
1.21
0.71
0.00
0.02
0.18
0.07
0.86
45.70
Appendix
Table 8 : Ablation of planning loss components on nuScenes.
Collision signal
Negative agg.
L2 (m) ↓
Collision (%) ↓
Cost AUC ↑
Cost Gap ↑
1s
2s
3s
Avg.
1s
2s
3s
Avg.
Ground-truth collision
Max
0.29
0.69
1.30
0.76
0.01
0.06
0.27
0.11
0.85
41.67
Ground-truth collision
Top- k
0.27
0.63
1.16
0.69
0.00
0.06
0.25
0.10
0.86
46.66
Top-1 predicted collision
Max
0.27
0.62
1.16
0.68
0.00
0.04
0.19
0.08
0.88
42.56
Top-1 predicted collision
Top- k
0.27
0.61
1.12
0.67
0.04
0.12
0.23
0.13
0.87
45.71
Contingency collision risk
Max
0.28
0.66
1.23
0.72
0.01
0.10
0.28
0.13
0.84
41.22
Appendix
Table 9 : Ablation of hard-negative aggregation for safety-margin training on nuScenes.
Mixing score
L2 (m) ↓
Collision (%) ↓
Cost AUC ↑
Cost Gap ↑
1s
2s
3s
Avg.
1s
2s
3s
Avg.
Mean trajectory cost
0.28
0.65
1.21
0.71
0.00
0.02
0.18
0.07
0.86
45.70
Per-timestep cost
0.30
0.71
1.32
0.77
0.02
0.10
0.30
0.14
0.86
42.99
Appendix
Table 10 : Ablation of cost profiles for cost-guided cluster mixing on nuScenes.
Aggregation
L2 (m) ↓
Collision (%) ↓
Cost AUC ↑
Cost Gap ↑
1s
2s
3s
Avg.
1s
2s
3s
Avg.
Top-1
0.33
0.79
1.50
0.87
0.03
0.17
0.37
0.19
0.81
33.99
Worst Case
0.32
0.78
1.47
0.86
0.03
0.16
0.37
0.19
0.81
37.81
Mean
0.32
0.77
1.45
0.85
0.00
0.09
0.30
0.13
0.82
38.97
Probability-weighted
0.33
0.80
1.50
0.88
0.02
0.14
0.35
0.17
0.81
37.97
Contingency
0.28
0.65
1.21
0.71
0.00
0.02
0.18
0.07
0.86
45.70
Appendix
Table 11 : Ablation of scene-cost aggregation on nuScenes.
Figure 6 : Qualitative nuScenes examples. The camera views and BEV visualizations show representative interactive scenes in which the learned cost topology guides selection among reachable ego candidates.
Figure 7 : Zero-shot qualitative examples on real-world driving logs. Each row shows surround-view camera predictions, the BEV cost visualization with the selected trajectory, and ground-truth boxes with ego trajectory and semantic map support.
Institute of Measurement and Control Systems, Karlsruhe Institute of Technology (KIT), Karlsruhe, Germany · FZI Research Center for Information Technology, Karlsruhe, Germany