Visual foundation models (VFMs) are increasingly integrated into end-to-end autonomous driving for their powerful representations, yet it remains unclear when these representations improve driving performance. To investigate this question, we introduce ViRA, a planner-agnostic visual representation alignment framework that keeps the planner architecture and inference cost unchanged. Our study reveals three findings: (1) VFM-guided visual representations consistently improve driving performance across diverse end-to-end planners, with gains extending to zero-shot closed-loop evaluation. (2) The choice of VFM target matters for planning performance, and alignment to a different VFM can further benefit planners with pre-trained VFM encoders. (3) Auxiliary perception supervision reduces sensitivity to VFM target selection, narrowing the EPDMS spread across five targets from 2.7 to 0.5 points and potentially compensating for less effective VFM targets. Guided by these findings, we develop ViRA-Diffusion, a diffusion-based planner trained without auxiliary perception supervision, which achieves 92.3 EPDMS on NAVSIM v2 navtest, outperforming recent methods in our comparison by at least 1.9 points. The results motivate jointly considering target selection and planner supervision when integrating VFMs into end-to-end autonomous driving. The results and demo are available at https://github.com/OpenDriveLab/ViRA.
Figures & tables
Figure 1: Better visual representations improve planning, with gains depending on both VFMs and planners. (a) ViRA aligns planner representations with a frozen VFM during training, leaving inference unchanged. (b) Segmentation and visual attention suggest improved representations. (c) Better visual representations improve planning, but the gains depend on the VFM target and planner.
Figure 2: t-SNE visualization of visual representations. Planner representations without and with ViRA are shown alongside the reference VFM representations for (a) TransFuser and (b) Rap*. The visualization suggests that representation alignment with ViRA shifts planner representations toward those of the reference VFM in the t-SNE space.
Methods
Decoder
NC ↑
DAC ↑
DDC ↑
TLC ↑
EP ↑
TTC ↑
LK ↑
HC ↑
EC ↑
EPDMS ↑
Human Agent
-
100
100
99.8
100
97.8
100
100
98.1
90.1
97.9
Perception-based
TransFuser ( Chitta et al., 2023 )
97.1
89.6
99.0
99.8
97.9
95.9
95.9
98.3
87.3
83.6
TransFuser-ViRA
Regression
97.7
94.1
99.5
99.8
98.2
96.7
96.8
98.3
87.9
89.0 +5.4
DiffusionDrive ( Liao et al., 2025 )
98.2
95.9
99.4
99.8
87.5
97.3
96.8
98.3
87.7
84.5
DiffusionDrive-ViRA
Diffusion
98.2
96.5
99.6
99.8
98.2
97.4
97.4
98.3
88.0
91.9 +7.4
Table 1: Performance on the NAVSIM v2 navtest benchmark. Rap ∗ is our reimplementation with a different backbone from the official codebase and without the original data augmentation. For a fair comparison, all methods use camera input only, without LiDAR.
Methods
Decoder
Stage
NC ↑
DAC ↑
DDC ↑
TLC ↑
EP ↑
TTC ↑
LK ↑
HC ↑
EC ↑
EPDMS ↑
S1
94.4
98.8
100
99.5
100
93.5
99.3
87.7
36.0
PDM-Closed
-
S2
90.5
90.6
95.4
98.4
100
98.4
74.2
91.9
29.7
56.6
Perception-based
S1
96.4
73.8
95.6
99.3
94.8
95.6
88.0
97.6
82.2
TransFuser ( Chitta et al., 2023 )
S2
82.6
62.8
77.9
98.2
91.0
79.9
40.9
97.5
80.8
27.6
S1
97.4
80.0
98.7
99.3
98.0
95.6
95.8
97.8
81.3
Table 2: Performance on the NAVSIM v2 navhard benchmark. Rap ∗ is our reimplementation with a different backbone from the official codebase and without the original data augmentation. PDM-Closed plans with GT symbolic inputs, while all other methods use camera data only as input.
Method
Easy
Medium
Hard
Extreme
Overall
RC
HD-Score
RC
HD-Score
RC
HD-Score
RC
HD-Score
RC
HD-Score
Perception-based
TransFuser ( Chitta et al., 2023 )
66.8
56.7
28.1
11.6
24.5
9.2
26.2
10.9
36.4
22.1
TransFuser-ViRA
72.7
62.6
31.9
15.3
27.8
11.4
25.5
9.9
39.5
24.8 +2.7
DiffusionDrive ( Liao et al., 2025 )
68.9
56.8
29.8
12.4
25.9
9.7
25.7
9.8
37.6
22.2
DiffusionDrive-ViRA
77.5
69.3
37.7
20.7
30.0
12.0
27.9
12.3
43.3
28.6 +6.4
Table 3: Zero-shot Performance on the HUGSIM Benchmark. Rap ∗ is our reimplementation with a different backbone from the official codebase and without the original data augmentation.
Figure 3: Qualitative results showing improved perception and planning with ViRA.
VFM
Background ↑
Road ↑
Walkway ↑
Centerline ↑
Vehicle ↑
mIoU ↑
Baseline
85.4
62.3
56.3
23.6
27.1
36.4
Random-init
82.4
58.3
47.2
12.5
14.7
30.7 -5.7
DVGT ( Zuo et al., 2026a )
86.9
64.6
59.3
28.9
33.4
39.1 +2.7
Table 4: VFMs dependency on BEV segmentation of TransFuser. Static objects and pedestrians are omitted from the class-wise results; mIoU is computed over all seven classes.
Methods
VFMs
NC ↑
DAC ↑
DDC ↑
TLC ↑
LK ↑
EPDMS ↑
Perception-based
Baseline
97.1
89.6
99.0
99.8
95.9
83.6
Random-init
96.8
90.2
99.0
99.7
95.5
82.0 -1.6
TransFuser ( Chitta et al., 2023 )
DVGT ( Zuo et al., 2026a )
97.7
94.1
99.5
99.8
96.8
89.0 +5.4
Perception-free
Baseline
91.6
88.8
93.5
99.8
90.9
72.8
Table 5: Effect of different VFM teachers on TransFuser and Rap*. Performance on the NAVSIM v2 navtest benchmark.
Table 6: Rap ∗ ( Feng et al., 2026 ) with DINOv3 ( Siméoni et al., 2025 ) backbones. Performance on the NAVSIM v2 navtest benchmark under different VFM teachers.
Figure 5: Different VFM targets across planners.
Figure 6: Auxiliary perception supervision narrows the spread of planning gains across VFM targets. (a) BEV mIoU across VFM targets; the dashed line marks the baseline (36.4). (b) Planning gain Δ EPDMS with and without auxiliary perception supervision (baselines 83.6 and 82.6, respectively).
Method
NC ↑
DAC ↑
DDC ↑
TLC ↑
EP ↑
TTC ↑
LK ↑
HC ↑
EC ↑
EPDMS ↑
Epona ( Zhang et al., 2025 )
97.1
95.7
99.3
99.7
88.6
96.3
97.0
98.0
67.8
85.1
DiffusionDriveV2 ( Zou et al., 2025 )
97.7
96.6
99.2
99.8
88.9
97.2
96.0
97.8
91.0
87.5
Latent-WAM ( Wang et al., 2026b )
98.1
97.3
99.6
99.8
87.7
97.3
97.6
98.1
87.3
89.3
DriveFuture ( Hong et al., 2026 )
98.8
99.1
99.6
99.9
86.6
98.4
96.4
98.3
74.8
89.9
SparseDriveV2 ( Sun et al., 2026 )
98.1
98.1
99.6
99.8
91.1
97.3
96.9
98.2
78.4
90.1
Discrete-WAM ( Yao et al., 2026 )
98.5
98.2
99.7
99.8
90.5
97.9
97.2
98.3
78.1
90.4
Table 7: Comparison with state-of-the-art methods on the NAVSIM v2 navtest benchmark. ViRA-Diffusion is DiffusionDrive ( Liao et al., 2025 ) trained with ViRA and without auxiliary perception supervision.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Figure S.1: Pipeline of ViRA. Multi-view images are encoded by both the planner encoder and a frozen VFM target, producing student and VFM representations. The VFM representations supervise latent alignment only, without being injected into the policy head. The policy head operates solely on planner representations, and the VFM branch is removed after training, yielding no inference overhead. Q, K, and V denote query, key, and value; KL denotes Kullback–Leibler divergence; CE denotes cross-entropy.
Hyperpara.
TransFuser ( Chitta et al., 2023 )
DiffusionDrive ( Liao et al., 2025 )
Rap ∗ ( Feng et al., 2026 )
Model Configuration
Sensors
3×Cam.
3×Cam.
4×Cam.
Resolution
2048×512
2048×512
4×768×448
Horizon
4s
4s
5s
Frequency
2Hz
2Hz
2Hz
Backbone
R34 ( He et al., 2016 )
R34 ( He et al., 2016 )
R34 ( He et al., 2016 )
Appendix
Table S.1: Model and training hyperparameters.
MSE
Cosine
KL
Contrastive
DAC
DDC
TLC
LK
EPDMS ↑
✓
✗
✗
✗
90.9
98.6
99.8
92.3
84.4
✗
✓
✗
✗
90.8
99.0
99.8
95.6
85.1
✗
✗
✓
✗
90.3
98.9
99.7
94.6
84.3
✗
✗
✗
✓
90.4
98.9
99.8
95.8
84.6
✓
✓
✗
✗
89.6
99.0
99.7
95.9
83.6
✓
✗
✓
✗
92.9
99.4
99.8
95.8
87.8
Appendix
Table S.2: Effect of alignment objectives on navtest. MSE, Cosine, KL, and Contrastive denote mean squared error, cosine similarity, Kullback–Leibler divergence, and contrastive loss, respectively.
Figure S.2: Feature correlation analysis.
Teachers
NC ↑
DAC ↑
DDC ↑
TLC ↑
EP ↑
TTC ↑
LK ↑
HC ↑
EC ↑
EPDMS ↑
Baseline
97.1
89.6
99.0
99.8
97.9
95.9
95.9
98.3
87.3
83.6
Random init
96.8
90.2
99.0
99.7
98.0
96.1
95.5
98.3
86.7
82.0 -1.6
DVGT ( Zuo et al., 2026a )
97.7
94.1
99.5
99.8
98.2
96.7
96.8
98.3
87.9
89.0 +5.4
VGGT ( Wang et al., 2025 )
97.8
94.6
99.3
99.8
97.8
97.0
96.7
98.3
87.2
89.3 +5.7
DA3 ( Lin et al., 2025 )
97.9
94.2
99.5
99.8
98.0
97.0
95.9
98.3
87.5
89.1 +5.5
DINOv3-L ( Siméoni et al., 2025 )
97.6
94.4
99.3
99.7
98.0
97.0
96.7
98.3
86.6
88.8 +5.2
Appendix
Table S.3: Effect of different VFM teachers on TransFuser. Performance on the NAVSIM v2 navtest benchmark.
Teachers
NC ↑
DAC ↑
DDC ↑
TLC ↑
EP ↑
TTC ↑
LK ↑
HC ↑
EC ↑
EPDMS ↑
Baseline
97.0
89.9
98.5
99.8
97.9
96.3
94.3
98.4
86.8
82.6
Random init
95.4
90.9
97.7
99.5
98.1
94.2
93.6
98.3
84.8
80.1 -2.5
DVGT ( Zuo et al., 2026a )
97.9
90.9
99.1
99.7
98.0
96.8
96.5
98.3
87.2
88.2 +5.6
VGGT ( Wang et al., 2025 )
97.7
92.9
99.1
99.8
98.0
96.6
95.9
98.3
87.5
87.6 +5.0
DA3 ( Lin et al., 2025 )
97.9
92.4
99.1
99.7
97.9
96.7
95.9
98.3
85.8
87.2 +4.6
DINOv3-L ( Siméoni et al., 2025 )
97.9
90.9
99.3
99.8
98.2
96.9
96.7
98.5
88.6
89.1 +6.5
Appendix
Table S.4: Effect of different VFM teachers on TransFuser without auxiliary perception supervision. Performance on the NAVSIM v2 navtest benchmark.
Teachers
NC ↑
DAC ↑
DDC ↑
TLC ↑
EP ↑
TTC ↑
LK ↑
HC ↑
EC ↑
EPDMS ↑
Baseline
91.6
88.8
93.5
99.8
87.2
92.5
90.9
97.2
56.7
72.8
Random init
94.6
88.1
95.1
99.8
84.2
95.1
93.1
96.7
41.9
71.1 -1.7
DVGT ( Zuo et al., 2026a )
97.1
96.9
99.2
99.9
87.5
96.9
96.2
97.7
48.1
83.7 +10.9
VGGT ( Wang et al., 2025 )
96.4
95.6
98.6
99.9
84.0
96.9
95.1
97.4
36.5
79.4 +6.6
DA3 ( Lin et al., 2025 )
96.1
95.6
98.6
99.7
88.8
96.1
96.0
97.7
54.9
82.5 +9.7
DINOv3 ( Siméoni et al., 2025 )
97.6
96.4
98.9
100.0
87.8
97.7
95.9
98.0
58.4
84.9 +12.1
Appendix
Table S.5: Effect of different VFM teachers on Rap ∗ . Performance on the NAVSIM v2 navtest benchmark. Rap ∗ is our reimplementation with a different backbone from the official codebase and without the original data augmentation.
Teachers
BG ↑
Road ↑
Walkways ↑
Centerline ↑
SO ↑
Vehicles ↑
Pedestrians ↑
mIoU ↑
Baseline
85.4
62.3
56.3
23.6
0.1
27.1
0.0
36.4
DA3 ( Lin et al., 2025 )
86.6
63.8
59.1
28.2
0.3
32.3
0.0
38.6 +2.2
VGGT ( Wang et al., 2025 )
86.5
63.9
58.5
27.6
0.3
33.6
0.0
38.6 +2.2
SAM3 ( Carion et al., 2025 )
85.8
62.4
56.7
25.0
0.1
28.0
0.0
36.9 +0.5
DVGT ( Zuo et al., 2026a )
86.9
64.6
59.3
28.9
0.5
33.4
0.2
39.1 +2.7
DINOv3 ( Siméoni et al., 2025 )
87.1
64.7
60.1
29.8
0.3
33.5
0.2
39.4 +3.0
Appendix
Table S.6: VFMs dependency on BEV segmentation of TransFuser. Class-wise BEV segmentation performance of TransFuser under different VFM targets. BG indicates background, SO indicates static objects.
Figure S.3: Additional qualitative analysis for perception-based planners across different VFMs.
Figure S.4: Additional qualitative analysis for perception-based planners.
Figure S.5: Additional qualitative analysis for perception-free planners across different VFMs.
Figure S.6: Failure case analysis for a perception-free planner.
Figure S.7: Failure case analysis for perception-based planners.
Graduate School of Mobility, Korea Advanced Institute of Science & Technology (KAIST), Daejeon, Korea · Robotics Program, Korea Advanced Institute of Science & Technology (KAIST), Daejeon, Korea