cs.RO · 2606.07170 Copy arXiv ID · Jun 5, 2026 Save Test-Time Trajectory Optimization for Autonomous Driving Authors: Yihong Xu , Eloi Zablocki , Yuan Yin , Elias Ramzi , Ellington Kirby , Alexandre Boulch , Matthieu Cord
Organizations: valeo.ai, Paris, France · Sorbonne Universit´e, CNRS, ISIR, F-75005 Paris, France
Abstract End-to-end planners for autonomous driving typically generate a set of candidate trajectories, score each one, and return the highest-scoring candidate. However, the scorer is applied only after the proposals are generated and cannot influence the set of trajectories: a weak set of candidates limits planning performance regardless of the scorer's quality. We instead treat the scorer as a learned trajectory-level reward function and search for trajectories that maximize it. Our method, TOAD, runs the Cross-Entropy Method at test time, warm-started from the planner's proposals. It requires no retraining and is plug-and-play for existing planners. Across six base planners, TOAD improves results on NAVSIM-v1 (94.7 PDMS), NAVSIM-v2 (56.3 EPDMS), and the closed-loop HUGSIM benchmark. The code will be made publicly available via the project page: https://valeoai.github.io/TOAD/ .
Explore similar work May 14, 2026 · Sining Ang, Yuguang Yang, Canyu Chen +1 Closed-Loop Reliability Closed-Loop Control
Jul 1, 2026 · Chong He, Yuechen Luo, Fang Li +2 Autonomous Driving Drives
Sep 1, 2026 · Yaguang Li, Jiaru Zhang, Chuheng Wei +2 Naturalistic Driving Data Diffusion Planning
May 14, 2026 · cs.RO J/K move · Enter open · S save
Sining Ang, Yuguang Yang, Canyu Chen, Yan Wang
Department of Automation, University of Science and Technology of China · Institute for AI Industry Research, Tsinghua University · School of Electronic Information Engineering, Beihang University · National College for Excellent Engineers, Beihang University
End-to-end autonomous driving planners are commonly trained by imitating a single logged trajectory, yet evaluated by rule-based planning metrics that measure safety, feasibility, progress, and comfort. This creates a training--evaluation mismatch: trajectories close to the logged path may violate planning rules, while alternatives farther from the demonstration can remain valid and high-scoring. The mismatch is especially limiting for proposal-selection planners, whose performance depends on candidate-set coverage and scorer ranking quality. We propose CLOVER, a Closed-LOop Value Estimation and Ranking framework for end-to-end autonomous driving planning. CLOVER follows a lightweight generator--scorer formulation: a generator produces diverse candidate trajectories, and a scorer predicts planning-metric sub-scores to rank them at inference time. To expand proposal support beyond single-trajectory imitation, CLOVER constructs evaluator-filtered pseudo-expert trajectories and trains the generator with set-level coverage supervision. It then performs conservative closed-loop self-distillation: the scorer is fitted to true evaluator sub-scores on generated proposals, while the generator is refined toward teacher-selected top-
k k k and vector-Pareto targets with stability regularization. We analyze when an imperfect scorer can improve the generator, showing that scorer-mediated refinement is reliable when scorer-selected targets are enriched under the true evaluator and updates remain conservative. On NAVSIM, CLOVER achieves 94.5 PDMS and 90.4 EPDMS, establishing a new state of the art. On the more challenging NavHard split, it obtains 48.3 EPDMS, matching the strongest reported result. On supplementary nuScenes open-loop evaluation, CLOVER achieves the lowest L2 error and collision rate among compared methods. Code data will be released at https://github.com/WilliamXuanYu/CLOVER.