Organizations: Chair of Robotics, Artificial Intelligence and Real-time Systems, Technical University of Munich, Garching, Germany · Computer Science and Technology School, Huaibei Normal University, Huaibei, China
We present S2Planner, a trajectory planner that combines three front-facing cameras with ego-motion history and the current driving command. A fine-tuned DINOv3 backbone and a Spatial Tuning Adapter produce multi-scale image features; a coarse-to-fine decoder then uses trajectory self-attention and camera-projected cross-attention to refine candidate waypoints. The contribution is the integration of ego-conditioned trajectory initialization with iterative, geometry-guided sampling of multi-scale image features, rather than a new visual backbone or attention operator. On the NAVSIM v1 non-reactive evaluation, the previously reported navtest run obtained 88.03 PDMS. Because that run was selected using navtest performance, this number is exploratory and cannot be interpreted as an unbiased test estimate. Validation-selected evaluation on unexposed data, repeated runs, and computational measurements are needed to establish generalization and efficiency.
Figures & tables
Fig. 1: S 2 Planner architecture. A fine-tuned DINOv3 backbone with a Spatial Tuning Adapter extracts multi-scale features from three concatenated camera images. Ego history and status condition coarse trajectory tokens and waypoints. Stacked trajectory self-attention and camera-projected cross-attention refine the candidates; the highest-scoring mode is returned.
Fig. 2: Concatenated front-facing images and a PCA visualization of DINOv3 patch features. This example illustrates the feature structure; a single visualization cannot establish a general loss of boundary or small-object information.
Fig. 3: DINOv3 and the Spatial Tuning Adapter architecture adopted from DEIMv2 [ 6 ] to form multi-scale features.
Fig. 4: Visualization of the anchor-based feature sampling. The interaction of the queries in the trajectory feature space is activated while the dense attention computation is prevented. For the sake of simplicity, the offset is not presented here.
Fig. 5: Illustration of the principle of Spatial Cross Attention. The reference anchors are directly derived from the actual trajectory coordinates, making sure that the metric consistency is preserved. For each trajectory query, only the feature from region where the projected 2D point locates inside the image plane would be seen as valid. The new trajectory feature is computed based on a weighted aggregation of the valid image features
Method
Sensors
Anchors
NC
DAC
TTC
C
EP
PDMS
Transfuser [ 41 ]
Camera + LiDAR
0
97.7
92.8
92.8
100
79.2
84.0
DRAMA [ 42 ]
Camera + LiDAR
0
98.0
93.1
94.8
100
80.1
85.5
VADv2-V 8192 [ 43 ]
Camera + LiDAR
8192
97.2
89.1
91.6
100
76.0
80.9
Hydra-MDP-V 8192 [ 3 ]
Camera + LiDAR
8192
97.9
91.7
92.9
100
77.6
83.0
Hydra-MDP-V 8192 -W-EP [ 3 ]
Camera + LiDAR
8192
98.3
96.0
94.6
100
78.7
86.5
DiffusionDrive [ 4 ]
Camera + LiDAR
20
98.2
96.2
94.7
100
82.2
88.1
TABLE I: Previously reported NAVSIM v1 navtest results (percent). Our row is test-selected and exploratory; the other rows are literature results reproduced from the original comparison, with their selection protocols and non-image inputs not audited here. Camera describes the sensor category only. NC, DAC, TTC, comfort (C), and EP are dataset summaries; dataset PDMS is aggregated from scenario scores. These rows do not establish a statistically significant ranking.
Fig. 6: Selected navtest examples. Green shows the displayed expert path and red the predicted path. The displayed paths may cover different time horizons; their apparent lengths should not be interpreted as a planning error or a progress score. These static overlays do not show how other agents would react to the plan.
Variant
DINOv3
STA
Deform.
PDMS
A0: ResNet
✗
✗
✗
72.4
A1: DINOv3
✓
✗
✗
84.9
A2: DINOv3 + STA
✓
✓
✗
86.3
A3: full model
✓
✓
✓
88.0
TABLE II: Previously reported component ablations on navtest. Values are exploratory and test-selected; repeated seeds and a validation-selected rerun are pending.
Variant
BEV head
PDMS
A1: direct regression
✗
84.9
A1.1: A1 + BEV auxiliary head
✓
83.5
TABLE III: Previously reported BEV auxiliary-head comparison on navtest. This tests one BEV-head design and loss setting only; values are exploratory.