Visual foundation models (VFMs) are increasingly integrated into end-to-end autonomous driving for their powerful representations, yet it remains unclear when these representations improve driving performance. To investigate this question, we introduce ViRA, a planner-agnostic visual representation alignment framework that keeps the planner architecture and inference cost unchanged. Our study reveals three findings: (1) VFM-guided visual representations consistently improve driving performance across diverse end-to-end planners, with gains extending to zero-shot closed-loop evaluation. (2) The choice of VFM target matters for planning performance, and alignment to a different VFM can further benefit planners with pre-trained VFM encoders. (3) Auxiliary perception supervision reduces sensitivity to VFM target selection, narrowing the EPDMS spread across five targets from 2.7 to 0.5 points and potentially compensating for less effective VFM targets. Guided by these findings, we develop ViRA-Diffusion, a diffusion-based planner trained without auxiliary perception supervision, which achieves 92.3 EPDMS on NAVSIM v2 navtest, outperforming recent methods in our comparison by at least 1.9 points. The results motivate jointly considering target selection and planner supervision when integrating VFMs into end-to-end autonomous driving. The results and demo are available at https://github.com/OpenDriveLab/ViRA.
Figures & tables
Figure 1: Better visual representations improve planning, with gains depending on both VFMs and planners. (a) ViRA aligns planner representations with a frozen VFM during training, leaving inference unchanged. (b) Segmentation and visual attention suggest improved representations. (c) Better visual representations improve planning, but the gains depend on the VFM target and planner.
Figure 2: t-SNE visualization of visual representations. Planner representations without and with ViRA are shown alongside the reference VFM representations for (a) TransFuser and (b) Rap*. The visualization suggests that representation alignment with ViRA shifts planner representations toward those of the reference VFM in the t-SNE space.
Methods
Decoder
NC ↑
DAC ↑
DDC ↑
TLC ↑
EP ↑
TTC ↑
LK ↑
HC ↑
EC ↑
EPDMS ↑
Human Agent
-
100
100
99.8
100
97.8
100
100
98.1
90.1
97.9
Perception-based
TransFuser ( Chitta et al., 2023 )
97.1
89.6
99.0
99.8
97.9
95.9
95.9
98.3
87.3
83.6
TransFuser-ViRA
Regression
97.7
94.1
99.5
99.8
98.2
96.7
96.8
98.3
87.9
89.0 +5.4
DiffusionDrive ( Liao et al., 2025 )
98.2
95.9
99.4
99.8
87.5
97.3
96.8
98.3
87.7
84.5
DiffusionDrive-ViRA
Diffusion
98.2
96.5
99.6
99.8
98.2
97.4
97.4
98.3
88.0
91.9 +7.4
Table 1: Performance on the NAVSIM v2 navtest benchmark. Rap ∗ is our reimplementation with a different backbone from the official codebase and without the original data augmentation. For a fair comparison, all methods use camera input only, without LiDAR.
Methods
Decoder
Stage
NC ↑
DAC ↑
DDC ↑
TLC ↑
EP ↑
TTC ↑
LK ↑
HC ↑
EC ↑
EPDMS ↑
S1
94.4
98.8
100
99.5
100
93.5
99.3
87.7
36.0
PDM-Closed
-
S2
90.5
90.6
95.4
98.4
100
98.4
74.2
91.9
29.7
56.6
Perception-based
S1
96.4
73.8
95.6
99.3
94.8
95.6
88.0
97.6
82.2
TransFuser ( Chitta et al., 2023 )
S2
82.6
62.8
77.9
98.2
91.0
79.9
40.9
97.5
80.8
27.6
S1
97.4
80.0
98.7
99.3
98.0
95.6
95.8
97.8
81.3
Table 2: Performance on the NAVSIM v2 navhard benchmark. Rap ∗ is our reimplementation with a different backbone from the official codebase and without the original data augmentation. PDM-Closed plans with GT symbolic inputs, while all other methods use camera data only as input.
Method
Easy
Medium
Hard
Extreme
Overall
RC
HD-Score
RC
HD-Score
RC
HD-Score
RC
HD-Score
RC
HD-Score
Perception-based
TransFuser ( Chitta et al., 2023 )
66.8
56.7
28.1
11.6
24.5
9.2
26.2
10.9
36.4
22.1
TransFuser-ViRA
72.7
62.6
31.9
15.3
27.8
11.4
25.5
9.9
39.5
24.8 +2.7
DiffusionDrive ( Liao et al., 2025 )
68.9
56.8
29.8
12.4
25.9
9.7
25.7
9.8
37.6
22.2
DiffusionDrive-ViRA
77.5
69.3
37.7
20.7
30.0
12.0
27.9
12.3
43.3
28.6 +6.4
Table 3: Zero-shot Performance on the HUGSIM Benchmark. Rap ∗ is our reimplementation with a different backbone from the official codebase and without the original data augmentation.
Figure 3: Qualitative results showing improved perception and planning with ViRA.
VFM
Background ↑
Road ↑
Walkway ↑
Centerline ↑
Vehicle ↑
mIoU ↑
Baseline
85.4
62.3
56.3
23.6
27.1
36.4
Random-init
82.4
58.3
47.2
12.5
14.7
30.7 -5.7
DVGT ( Zuo et al., 2026a )
86.9
64.6
59.3
28.9
33.4
39.1 +2.7
Table 4: VFMs dependency on BEV segmentation of TransFuser. Static objects and pedestrians are omitted from the class-wise results; mIoU is computed over all seven classes.
Methods
VFMs
NC ↑
DAC ↑
DDC ↑
TLC ↑
LK ↑
EPDMS ↑
Perception-based
Baseline
97.1
89.6
99.0
99.8
95.9
83.6
Random-init
96.8
90.2
99.0
99.7
95.5
82.0 -1.6
TransFuser ( Chitta et al., 2023 )
DVGT ( Zuo et al., 2026a )
97.7
94.1
99.5
99.8
96.8
89.0 +5.4
Perception-free
Baseline
91.6
88.8
93.5
99.8
90.9
72.8
Table 5: Effect of different VFM teachers on TransFuser and Rap*. Performance on the NAVSIM v2 navtest benchmark.
Table 6: Rap ∗ ( Feng et al., 2026 ) with DINOv3 ( Siméoni et al., 2025 ) backbones. Performance on the NAVSIM v2 navtest benchmark under different VFM teachers.
Figure 5: Different VFM targets across planners.
Figure 6: Auxiliary perception supervision narrows the spread of planning gains across VFM targets. (a) BEV mIoU across VFM targets; the dashed line marks the baseline (36.4). (b) Planning gain Δ EPDMS with and without auxiliary perception supervision (baselines 83.6 and 82.6, respectively).
Method
NC ↑
DAC ↑
DDC ↑
TLC ↑
EP ↑
TTC ↑
LK ↑
HC ↑
EC ↑
EPDMS ↑
Epona ( Zhang et al., 2025 )
97.1
95.7
99.3
99.7
88.6
96.3
97.0
98.0
67.8
85.1
DiffusionDriveV2 ( Zou et al., 2025 )
97.7
96.6
99.2
99.8
88.9
97.2
96.0
97.8
91.0
87.5
Latent-WAM ( Wang et al., 2026b )
98.1
97.3
99.6
99.8
87.7
97.3
97.6
98.1
87.3
89.3
DriveFuture ( Hong et al., 2026 )
98.8
99.1
99.6
99.9
86.6
98.4
96.4
98.3
74.8
89.9
SparseDriveV2 ( Sun et al., 2026 )
98.1
98.1
99.6
99.8
91.1
97.3
96.9
98.2
78.4
90.1
Discrete-WAM ( Yao et al., 2026 )
98.5
98.2
99.7
99.8
90.5
97.9
97.2
98.3
78.1
90.4
Table 7: Comparison with state-of-the-art methods on the NAVSIM v2 navtest benchmark. ViRA-Diffusion is DiffusionDrive ( Liao et al., 2025 ) trained with ViRA and without auxiliary perception supervision.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Figure S.1: Pipeline of ViRA. Multi-view images are encoded by both the planner encoder and a frozen VFM target, producing student and VFM representations. The VFM representations supervise latent alignment only, without being injected into the policy head. The policy head operates solely on planner representations, and the VFM branch is removed after training, yielding no inference overhead. Q, K, and V denote query, key, and value; KL denotes Kullback–Leibler divergence; CE denotes cross-entropy.
Hyperpara.
TransFuser ( Chitta et al., 2023 )
DiffusionDrive ( Liao et al., 2025 )
Rap ∗ ( Feng et al., 2026 )
Model Configuration
Sensors
3×Cam.
3×Cam.
4×Cam.
Resolution
2048×512
2048×512
4×768×448
Horizon
4s
4s
5s
Frequency
2Hz
2Hz
2Hz
Backbone
R34 ( He et al., 2016 )
R34 ( He et al., 2016 )
R34 ( He et al., 2016 )
Appendix
Table S.1: Model and training hyperparameters.
MSE
Cosine
KL
Contrastive
DAC
DDC
TLC
LK
EPDMS ↑
✓
✗
✗
✗
90.9
98.6
99.8
92.3
84.4
✗
✓
✗
✗
90.8
99.0
99.8
95.6
85.1
✗
✗
✓
✗
90.3
98.9
99.7
94.6
84.3
✗
✗
✗
✓
90.4
98.9
99.8
95.8
84.6
✓
✓
✗
✗
89.6
99.0
99.7
95.9
83.6
✓
✗
✓
✗
92.9
99.4
99.8
95.8
87.8
Appendix
Table S.2: Effect of alignment objectives on navtest. MSE, Cosine, KL, and Contrastive denote mean squared error, cosine similarity, Kullback–Leibler divergence, and contrastive loss, respectively.
Figure S.2: Feature correlation analysis.
Teachers
NC ↑
DAC ↑
DDC ↑
TLC ↑
EP ↑
TTC ↑
LK ↑
HC ↑
EC ↑
EPDMS ↑
Baseline
97.1
89.6
99.0
99.8
97.9
95.9
95.9
98.3
87.3
83.6
Random init
96.8
90.2
99.0
99.7
98.0
96.1
95.5
98.3
86.7
82.0 -1.6
DVGT ( Zuo et al., 2026a )
97.7
94.1
99.5
99.8
98.2
96.7
96.8
98.3
87.9
89.0 +5.4
VGGT ( Wang et al., 2025 )
97.8
94.6
99.3
99.8
97.8
97.0
96.7
98.3
87.2
89.3 +5.7
DA3 ( Lin et al., 2025 )
97.9
94.2
99.5
99.8
98.0
97.0
95.9
98.3
87.5
89.1 +5.5
DINOv3-L ( Siméoni et al., 2025 )
97.6
94.4
99.3
99.7
98.0
97.0
96.7
98.3
86.6
88.8 +5.2
Appendix
Table S.3: Effect of different VFM teachers on TransFuser. Performance on the NAVSIM v2 navtest benchmark.
Teachers
NC ↑
DAC ↑
DDC ↑
TLC ↑
EP ↑
TTC ↑
LK ↑
HC ↑
EC ↑
EPDMS ↑
Baseline
97.0
89.9
98.5
99.8
97.9
96.3
94.3
98.4
86.8
82.6
Random init
95.4
90.9
97.7
99.5
98.1
94.2
93.6
98.3
84.8
80.1 -2.5
DVGT ( Zuo et al., 2026a )
97.9
90.9
99.1
99.7
98.0
96.8
96.5
98.3
87.2
88.2 +5.6
VGGT ( Wang et al., 2025 )
97.7
92.9
99.1
99.8
98.0
96.6
95.9
98.3
87.5
87.6 +5.0
DA3 ( Lin et al., 2025 )
97.9
92.4
99.1
99.7
97.9
96.7
95.9
98.3
85.8
87.2 +4.6
DINOv3-L ( Siméoni et al., 2025 )
97.9
90.9
99.3
99.8
98.2
96.9
96.7
98.5
88.6
89.1 +6.5
Appendix
Table S.4: Effect of different VFM teachers on TransFuser without auxiliary perception supervision. Performance on the NAVSIM v2 navtest benchmark.
Teachers
NC ↑
DAC ↑
DDC ↑
TLC ↑
EP ↑
TTC ↑
LK ↑
HC ↑
EC ↑
EPDMS ↑
Baseline
91.6
88.8
93.5
99.8
87.2
92.5
90.9
97.2
56.7
72.8
Random init
94.6
88.1
95.1
99.8
84.2
95.1
93.1
96.7
41.9
71.1 -1.7
DVGT ( Zuo et al., 2026a )
97.1
96.9
99.2
99.9
87.5
96.9
96.2
97.7
48.1
83.7 +10.9
VGGT ( Wang et al., 2025 )
96.4
95.6
98.6
99.9
84.0
96.9
95.1
97.4
36.5
79.4 +6.6
DA3 ( Lin et al., 2025 )
96.1
95.6
98.6
99.7
88.8
96.1
96.0
97.7
54.9
82.5 +9.7
DINOv3 ( Siméoni et al., 2025 )
97.6
96.4
98.9
100.0
87.8
97.7
95.9
98.0
58.4
84.9 +12.1
Appendix
Table S.5: Effect of different VFM teachers on Rap ∗ . Performance on the NAVSIM v2 navtest benchmark. Rap ∗ is our reimplementation with a different backbone from the official codebase and without the original data augmentation.
Teachers
BG ↑
Road ↑
Walkways ↑
Centerline ↑
SO ↑
Vehicles ↑
Pedestrians ↑
mIoU ↑
Baseline
85.4
62.3
56.3
23.6
0.1
27.1
0.0
36.4
DA3 ( Lin et al., 2025 )
86.6
63.8
59.1
28.2
0.3
32.3
0.0
38.6 +2.2
VGGT ( Wang et al., 2025 )
86.5
63.9
58.5
27.6
0.3
33.6
0.0
38.6 +2.2
SAM3 ( Carion et al., 2025 )
85.8
62.4
56.7
25.0
0.1
28.0
0.0
36.9 +0.5
DVGT ( Zuo et al., 2026a )
86.9
64.6
59.3
28.9
0.5
33.4
0.2
39.1 +2.7
DINOv3 ( Siméoni et al., 2025 )
87.1
64.7
60.1
29.8
0.3
33.5
0.2
39.4 +3.0
Appendix
Table S.6: VFMs dependency on BEV segmentation of TransFuser. Class-wise BEV segmentation performance of TransFuser under different VFM targets. BG indicates background, SO indicates static objects.
Figure S.3: Additional qualitative analysis for perception-based planners across different VFMs.
Figure S.4: Additional qualitative analysis for perception-based planners.
Figure S.5: Additional qualitative analysis for perception-free planners across different VFMs.
Figure S.6: Failure case analysis for a perception-free planner.
Figure S.7: Failure case analysis for perception-based planners.
Recent Vision-Language-Action (VLA) models for autonomous driving have incorporated world modeling by predicting future driving scenes alongside driving actions, demonstrating strong planning performance. Future driving scenes are utilized as dense supervision, encouraging the policy to learn rich internal representations useful for planning. However, these World-Modeling VLAs rely on explicit future generation to learn such representations, thereby introducing two key limitations: additional training burden and inference latency. To address these limitations, we propose RAF-VLA (Representation Alignment with the Future), a VLA-based autonomous driving framework that shapes planning-relevant internal representations through direct guidance from future-frame representations. RAF-VLA employs Future-Aligned Supervised Fine-Tuning, in which a straightforward regularization aligns the policy's hidden states with future-frame representations obtained from a pretrained world encoder while learning driving actions. This simple alignment allows RAF-VLA to avoid the training burden and inference latency associated with future generation. Extensive experiments on the NAVSIM benchmark show that RAF-VLA achieves competitive planning performance against state-of-the-art VLA planners with substantially fewer training samples seen. Moreover, RAF-VLA incurs only 3.8% training overhead and a negligible 1 ms inference overhead.
Dogun Kim, Yongjae Lee, Joonhee Lim +4
Graduate School of Mobility, Korea Advanced Institute of Science & Technology (KAIST), Daejeon, Korea · Robotics Program, Korea Advanced Institute of Science & Technology (KAIST), Daejeon, Korea
Vision-Language Models (VLMs) provide powerful semantic understanding and commonsense reasoning for End-to-End Autonomous Driving (E2E-AD) planning. However, trajectories directly generated by VLMs often encode only coarse driving intentions and remain insufficient for geometrically accurate, future-aware, and multi-view-grounded planning. To address these limitations, we develop the Layer-Wise World-Model-Guided Driving framework (LWDrive). LWDrive is a VLM planning framework that refines coarse trajectories through layer-wise world-model guidance. Instead of treating the VLM output as the final trajectory, LWDrive uses it as an intent-aware coarse plan, expands a diverse candidate space around it, and progressively refines the candidates through a Foresight Cascade Planner (FCP). Specifically, we introduce future-frame generation supervision to encourage the VLM to learn forward-looking scene representations, thereby injecting planning-relevant predictive dynamics into its internal hidden states. Built upon these world-model-supervised representations, FCP exploits VLM features across multiple layers and integrates historical temporal states, Action-Query representations, and current-frame multi-view Bird's-Eye-View (BEV) features to refine candidate trajectories in a coarse-to-fine manner. This design enables progressive correction of spatial positions and motion trends while grounding trajectory refinement with multi-view scene cues and preserving the high-level driving intention produced by the large model. Finally, a score head evaluates the refined candidates and selects the best trajectory as the final planning output. Experiments show that LWDrive achieves a score of 92.0 on the NAVSIM benchmark and 89.6 on NAVSIM-v2. Code and models will be made publicly available.
Driving policies require capabilities of scene understanding and future evolution prediction. To achieve this goal, current end-to-end models typically construct complex perception-planning pipelines or introduce world models that explicitly predict future states, resulting in a complex system architecture. Inspired by the transferability of general-purpose visual representations, we argue that combining sufficiently strong visual representations with representation world modeling can support effective planning without relying on complex inference-time auxiliary modules. Based on this insight, we present ReDrive, an end-to-end driving framework that strengthens planning-oriented visual features via future representation prediction. To achieve this, ReDrive adopts a three-stage training pipeline consisting of driving video pretraining, joint world-modeling and planning training, and planner adaptation. This yields a strong planning-oriented representation and a high-performance planner, while requiring neither auxiliary perception modules nor future prediction at inference time. Experiments on NAVSIM demonstrate strong performance, achieving 91.0 PDMS on NAVSIM v1 and 90.8 EPDMS on NAVSIM v2. These results show that shaping representations with world modeling is sufficient to enable high-performance end-to-end planning while retaining a simple encoder-planner inference pipeline.
Yueting Zhu, Shaoyu Chen, Yuehao Song +4
Huazhong University of Science & Technology · Horizon Robotics