Organizations: Chair of Robotics, Artificial Intelligence and Real-time Systems, Technical University of Munich, Garching, Germany · Computer Science and Technology School, Huaibei Normal University, Huaibei, China
We present S2Planner, a trajectory planner that combines three front-facing cameras with ego-motion history and the current driving command. A fine-tuned DINOv3 backbone and a Spatial Tuning Adapter produce multi-scale image features; a coarse-to-fine decoder then uses trajectory self-attention and camera-projected cross-attention to refine candidate waypoints. The contribution is the integration of ego-conditioned trajectory initialization with iterative, geometry-guided sampling of multi-scale image features, rather than a new visual backbone or attention operator. On the NAVSIM v1 non-reactive evaluation, the previously reported navtest run obtained 88.03 PDMS. Because that run was selected using navtest performance, this number is exploratory and cannot be interpreted as an unbiased test estimate. Validation-selected evaluation on unexposed data, repeated runs, and computational measurements are needed to establish generalization and efficiency.
Figures & tables
Fig. 1: S 2 Planner architecture. A fine-tuned DINOv3 backbone with a Spatial Tuning Adapter extracts multi-scale features from three concatenated camera images. Ego history and status condition coarse trajectory tokens and waypoints. Stacked trajectory self-attention and camera-projected cross-attention refine the candidates; the highest-scoring mode is returned.
Fig. 2: Concatenated front-facing images and a PCA visualization of DINOv3 patch features. This example illustrates the feature structure; a single visualization cannot establish a general loss of boundary or small-object information.
Fig. 3: DINOv3 and the Spatial Tuning Adapter architecture adopted from DEIMv2 [ 6 ] to form multi-scale features.
Fig. 4: Visualization of the anchor-based feature sampling. The interaction of the queries in the trajectory feature space is activated while the dense attention computation is prevented. For the sake of simplicity, the offset is not presented here.
Fig. 5: Illustration of the principle of Spatial Cross Attention. The reference anchors are directly derived from the actual trajectory coordinates, making sure that the metric consistency is preserved. For each trajectory query, only the feature from region where the projected 2D point locates inside the image plane would be seen as valid. The new trajectory feature is computed based on a weighted aggregation of the valid image features
Method
Sensors
Anchors
NC
DAC
TTC
C
EP
PDMS
Transfuser [ 41 ]
Camera + LiDAR
0
97.7
92.8
92.8
100
79.2
84.0
DRAMA [ 42 ]
Camera + LiDAR
0
98.0
93.1
94.8
100
80.1
85.5
VADv2-V 8192 [ 43 ]
Camera + LiDAR
8192
97.2
89.1
91.6
100
76.0
80.9
Hydra-MDP-V 8192 [ 3 ]
Camera + LiDAR
8192
97.9
91.7
92.9
100
77.6
83.0
Hydra-MDP-V 8192 -W-EP [ 3 ]
Camera + LiDAR
8192
98.3
96.0
94.6
100
78.7
86.5
DiffusionDrive [ 4 ]
Camera + LiDAR
20
98.2
96.2
94.7
100
82.2
88.1
TABLE I: Previously reported NAVSIM v1 navtest results (percent). Our row is test-selected and exploratory; the other rows are literature results reproduced from the original comparison, with their selection protocols and non-image inputs not audited here. Camera describes the sensor category only. NC, DAC, TTC, comfort (C), and EP are dataset summaries; dataset PDMS is aggregated from scenario scores. These rows do not establish a statistically significant ranking.
Fig. 6: Selected navtest examples. Green shows the displayed expert path and red the predicted path. The displayed paths may cover different time horizons; their apparent lengths should not be interpreted as a planning error or a progress score. These static overlays do not show how other agents would react to the plan.
Variant
DINOv3
STA
Deform.
PDMS
A0: ResNet
✗
✗
✗
72.4
A1: DINOv3
✓
✗
✗
84.9
A2: DINOv3 + STA
✓
✓
✗
86.3
A3: full model
✓
✓
✓
88.0
TABLE II: Previously reported component ablations on navtest. Values are exploratory and test-selected; repeated seeds and a validation-selected rerun are pending.
Variant
BEV head
PDMS
A1: direct regression
✗
84.9
A1.1: A1 + BEV auxiliary head
✓
83.5
TABLE III: Previously reported BEV auxiliary-head comparison on navtest. This tests one BEV-head design and loss setting only; values are exploratory.
End-to-end planners for autonomous driving typically generate a set of candidate trajectories, score each one, and return the highest-scoring candidate. However, the scorer is applied only after the proposals are generated and cannot influence the set of trajectories: a weak set of candidates limits planning performance regardless of the scorer's quality. We instead treat the scorer as a learned trajectory-level reward function and search for trajectories that maximize it. Our method, TOAD, runs the Cross-Entropy Method at test time, warm-started from the planner's proposals. It requires no retraining and is plug-and-play for existing planners. Across six base planners, TOAD improves results on NAVSIM-v1 (94.7 PDMS), NAVSIM-v2 (56.3 EPDMS), and the closed-loop HUGSIM benchmark. The code will be made publicly available via the project page: https://valeoai.github.io/TOAD/.
Yihong Xu, Eloi Zablocki, Yuan Yin +4
valeo.ai, Paris, France · Sorbonne Universit´e, CNRS, ISIR, F-75005 Paris, France
Vision-Language Models (VLMs) provide powerful semantic understanding and commonsense reasoning for End-to-End Autonomous Driving (E2E-AD) planning. However, trajectories directly generated by VLMs often encode only coarse driving intentions and remain insufficient for geometrically accurate, future-aware, and multi-view-grounded planning. To address these limitations, we develop the Layer-Wise World-Model-Guided Driving framework (LWDrive). LWDrive is a VLM planning framework that refines coarse trajectories through layer-wise world-model guidance. Instead of treating the VLM output as the final trajectory, LWDrive uses it as an intent-aware coarse plan, expands a diverse candidate space around it, and progressively refines the candidates through a Foresight Cascade Planner (FCP). Specifically, we introduce future-frame generation supervision to encourage the VLM to learn forward-looking scene representations, thereby injecting planning-relevant predictive dynamics into its internal hidden states. Built upon these world-model-supervised representations, FCP exploits VLM features across multiple layers and integrates historical temporal states, Action-Query representations, and current-frame multi-view Bird's-Eye-View (BEV) features to refine candidate trajectories in a coarse-to-fine manner. This design enables progressive correction of spatial positions and motion trends while grounding trajectory refinement with multi-view scene cues and preserving the high-level driving intention produced by the large model. Finally, a score head evaluates the refined candidates and selects the best trajectory as the final planning output. Experiments show that LWDrive achieves a score of 92.0 on the NAVSIM benchmark and 89.6 on NAVSIM-v2. Code and models will be made publicly available.
End-to-end autonomous driving planners typically generate trajectories from current observations alone. However, real-world driving is highly dynamic, and such reactive planning cannot anticipate future scene evolution, often leading to myopic decisions and safety-critical failures. We propose ProDrive, a world-model-based proactive planning framework that enables ego-environment co-evolution for autonomous driving. ProDrive jointly trains a query-centric trajectory planner and a bird's-eye-view (BEV) world model end-to-end: the planner generates diverse candidate trajectories and planning-aware ego tokens, while the world model predicts future scene evolution conditioned on them. By injecting planner features into the world model and evaluating all candidates in parallel, ProDrive preserves end-to-end gradient flow and allows future outcome assessment to directly shape planning. This bidirectional coupling enables proactive planning beyond current-observation-driven decision-making. Experiments on NAVSIM v1 show that ProDrive outperforms strong baselines in both safety and planning efficiency, while ablations validate the effectiveness of the proposed ego-environment coupling design.
Chuyao Fu, Shengzhe Gan, Zhuoli Ouyang +5
1Southern University of Science and Technology · 2Hong Kong University of Science and Technology