Driving policies require capabilities of scene understanding and future evolution prediction. To achieve this goal, current end-to-end models typically construct complex perception-planning pipelines or introduce world models that explicitly predict future states, resulting in a complex system architecture. Inspired by the transferability of general-purpose visual representations, we argue that combining sufficiently strong visual representations with representation world modeling can support effective planning without relying on complex inference-time auxiliary modules. Based on this insight, we present ReDrive, an end-to-end driving framework that strengthens planning-oriented visual features via future representation prediction. To achieve this, ReDrive adopts a three-stage training pipeline consisting of driving video pretraining, joint world-modeling and planning training, and planner adaptation. This yields a strong planning-oriented representation and a high-performance planner, while requiring neither auxiliary perception modules nor future prediction at inference time. Experiments on NAVSIM demonstrate strong performance, achieving 91.0 PDMS on NAVSIM v1 and 90.8 EPDMS on NAVSIM v2. These results show that shaping representations with world modeling is sufficient to enable high-performance end-to-end planning while retaining a simple encoder-planner inference pipeline.
Figures & tables
Figure 1 : Overview of existing visual representation learning paradigms and our proposed ReDrive for end-to-end autonomous driving.
Figure 2 : Overview of ReDrive and its three-stage training procedure. Self-Supervised Pretraining adapts the video encoder to the driving domain. Joint Training learns trajectory generation and future representation prediction from the shared history representation. Planner Adaptation freezes the encoder and future predictor and uses planner-generated trajectories as the condition for predictive supervision.
Type
Method
Inputs
NC ↑
DAC ↑
EP ↑
C ↑
TTC ↑
PDMS ↑
Perception-based
Transfuser [ 6 ]
C + L
97.7
92.8
79.2
100
92.8
84.0
VADv2 [ 15 ]
Camera
97.2
89.1
91.6
100
76.0
80.9
UniAD [ 13 ]
Camera
97.8
91.9
92.9
100
78.8
83.4
Hydra-MDP [ 26 ]
C + L
98.4
97.7
85.0
100
94.5
89.9
Hydra-MDP++ [ 22 ]
C + L
97.6
96.0
80.4
100
93.1
86.6
DiffusionDrive [ 29 ]
C + L
98.2
96.2
82.2
100
94.7
88.1
Table 1 : Comparison with state-of-the-art methods on NAVSIM v1. The best and second-best results are highlighted in bold and underlined, respectively.
Method
NC ↑
DAC ↑
DDC ↑
TLC ↑
EP ↑
TTC ↑
LK ↑
HC ↑
EC ↑
EPDMS ↑
DiffusionDrive [ 29 ]
98.2
95.9
99.4
99.8
87.5
97.3
96.8
98.3
87.7
84.5
DiffusionDriveV2 [ 54 ]
97.7
96.6
99.2
99.8
88.9
97.2
96.0
97.8
91.0
87.5
WAM-Diff [ 45 ]
99.0
98.4
99.3
99.9
87.0
98.6
96.2
98.1
78.5
89.7
DreamerAD [ 47 ]
98.0
97.2
99.5
99.8
87.8
97.4
97.5
98.3
72.4
87.7
Latent-WAM [ 38 ]
98.1
97.3
99.6
99.8
87.7
97.3
97.6
98.1
87.3
89.3
DriveFuture [ 11 ]
98.8
99.1
99.6
99.9
86.6
98.4
96.4
98.3
74.8
89.9
Table 2 : Comparison with state-of-the-art methods on NAVSIM v2. The best and second-best results are highlighted in bold and underlined, respectively. NC–EC uniformly report the corrected-evaluator metrics.
Method
Stage
NC ↑
DAC ↑
DDC ↑
TLC ↑
EP ↑
TTC ↑
LK ↑
HC ↑
EC ↑
S. ↑
EPDMS ↑
LTF [ 6 ]
S1
97.3
80.2
97.8
99.3
83.4
96.2
92.9
97.8
71.1
61.3
24.4
S2
79.4
69.0
85.6
98.5
83.8
76.7
47.9
97.0
70.6
39.2
DiffusionDrive [ 29 ]
S1
96.8
86.0
98.8
99.3
84.0
95.8
96.7
97.6
79.6
66.7
27.5
S2
80.1
72.8
84.4
98.4
85.9
76.6
46.4
96.3
72.8
40.5
GTRS-DP [ 27 ]
S1
94.7
78.8
96.1
99.5
83.0
94.4
92.0
97.5
72.8
–
23.8
S2
80.3
74.4
84.9
98.0
81.9
78.8
45.4
96.7
70.1
–
Table 3 : Comparison with state-of-the-art methods on NAVSIM v2 NavHard. S1 and S2 denote Stage-1 and Stage-2, respectively. S. denotes the per-stage score, and EPDMS reports the combined score. The best and second-best results are highlighted in bold and underlined, respectively.
Table 6Table 7
Figure 3 : Qualitative comparison of planning results using front-camera and bird’s-eye-view (BEV) visualizations. We compare the human trajectory, Drive-JEPA [ 39 ] , and ReDrive across representative driving scenarios.
Figure 4 : Sensitivity of the Future Predictor to trajectory conditions with the history representation held fixed. (a) Cosine similarities between future representations predicted under different lateral trajectory offsets. The horizontal axis and point color indicate the offsets of the two trajectory conditions being compared. (b) Pairwise cosine similarity matrix for all trajectory conditions.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5 : Illustration of camera alignment for PhysicalAI-Autonomous-Vehicles Dataset. The original wide-angle camera views are geometrically transformed using per-clip calibration to obtain aligned views with nuPlan-style camera geometry. The alignment reduces the discrepancy in camera projection and view distribution before self-supervised pretraining.
Figure 6 : Additional qualitative comparison of the human trajectory, Drive-JEPA, and ReDrive using front-camera and bird’s-eye-view (BEV) visualizations.
Module
Configuration
Setting
Encoder
Architecture
ViT-L
Depth
24
Hidden dimension
1024
Attention heads
16
Parameters
304M
Predictor
Depth
12
Appendix
Table 8 : Pretraining configuration.
Table 13
Figure 7 : Additional Future Predictor visualizations. Predicted future representations under different lateral trajectory conditions across diverse driving scenes. The representations exhibit consistent action-dependent variations across different scenarios.