Driving policies require capabilities of scene understanding and future evolution prediction. To achieve this goal, current end-to-end models typically construct complex perception-planning pipelines or introduce world models that explicitly predict future states, resulting in a complex system architecture. Inspired by the transferability of general-purpose visual representations, we argue that combining sufficiently strong visual representations with representation world modeling can support effective planning without relying on complex inference-time auxiliary modules. Based on this insight, we present ReDrive, an end-to-end driving framework that strengthens planning-oriented visual features via future representation prediction. To achieve this, ReDrive adopts a three-stage training pipeline consisting of driving video pretraining, joint world-modeling and planning training, and planner adaptation. This yields a strong planning-oriented representation and a high-performance planner, while requiring neither auxiliary perception modules nor future prediction at inference time. Experiments on NAVSIM demonstrate strong performance, achieving 91.0 PDMS on NAVSIM v1 and 90.8 EPDMS on NAVSIM v2. These results show that shaping representations with world modeling is sufficient to enable high-performance end-to-end planning while retaining a simple encoder-planner inference pipeline.
Figures & tables
Figure 1 : Overview of existing visual representation learning paradigms and our proposed ReDrive for end-to-end autonomous driving.
Figure 2 : Overview of ReDrive and its three-stage training procedure. Self-Supervised Pretraining adapts the video encoder to the driving domain. Joint Training learns trajectory generation and future representation prediction from the shared history representation. Planner Adaptation freezes the encoder and future predictor and uses planner-generated trajectories as the condition for predictive supervision.
Type
Method
Inputs
NC ↑
DAC ↑
EP ↑
C ↑
TTC ↑
PDMS ↑
Perception-based
Transfuser [ 6 ]
C + L
97.7
92.8
79.2
100
92.8
84.0
VADv2 [ 15 ]
Camera
97.2
89.1
91.6
100
76.0
80.9
UniAD [ 13 ]
Camera
97.8
91.9
92.9
100
78.8
83.4
Hydra-MDP [ 26 ]
C + L
98.4
97.7
85.0
100
94.5
89.9
Hydra-MDP++ [ 22 ]
C + L
97.6
96.0
80.4
100
93.1
86.6
DiffusionDrive [ 29 ]
C + L
98.2
96.2
82.2
100
94.7
88.1
Table 1 : Comparison with state-of-the-art methods on NAVSIM v1. The best and second-best results are highlighted in bold and underlined, respectively.
Method
NC ↑
DAC ↑
DDC ↑
TLC ↑
EP ↑
TTC ↑
LK ↑
HC ↑
EC ↑
EPDMS ↑
DiffusionDrive [ 29 ]
98.2
95.9
99.4
99.8
87.5
97.3
96.8
98.3
87.7
84.5
DiffusionDriveV2 [ 54 ]
97.7
96.6
99.2
99.8
88.9
97.2
96.0
97.8
91.0
87.5
WAM-Diff [ 45 ]
99.0
98.4
99.3
99.9
87.0
98.6
96.2
98.1
78.5
89.7
DreamerAD [ 47 ]
98.0
97.2
99.5
99.8
87.8
97.4
97.5
98.3
72.4
87.7
Latent-WAM [ 38 ]
98.1
97.3
99.6
99.8
87.7
97.3
97.6
98.1
87.3
89.3
DriveFuture [ 11 ]
98.8
99.1
99.6
99.9
86.6
98.4
96.4
98.3
74.8
89.9
Table 2 : Comparison with state-of-the-art methods on NAVSIM v2. The best and second-best results are highlighted in bold and underlined, respectively. NC–EC uniformly report the corrected-evaluator metrics.
Method
Stage
NC ↑
DAC ↑
DDC ↑
TLC ↑
EP ↑
TTC ↑
LK ↑
HC ↑
EC ↑
S. ↑
EPDMS ↑
LTF [ 6 ]
S1
97.3
80.2
97.8
99.3
83.4
96.2
92.9
97.8
71.1
61.3
24.4
S2
79.4
69.0
85.6
98.5
83.8
76.7
47.9
97.0
70.6
39.2
DiffusionDrive [ 29 ]
S1
96.8
86.0
98.8
99.3
84.0
95.8
96.7
97.6
79.6
66.7
27.5
S2
80.1
72.8
84.4
98.4
85.9
76.6
46.4
96.3
72.8
40.5
GTRS-DP [ 27 ]
S1
94.7
78.8
96.1
99.5
83.0
94.4
92.0
97.5
72.8
–
23.8
S2
80.3
74.4
84.9
98.0
81.9
78.8
45.4
96.7
70.1
–
Table 3 : Comparison with state-of-the-art methods on NAVSIM v2 NavHard. S1 and S2 denote Stage-1 and Stage-2, respectively. S. denotes the per-stage score, and EPDMS reports the combined score. The best and second-best results are highlighted in bold and underlined, respectively.
Table 6Table 7
Figure 3 : Qualitative comparison of planning results using front-camera and bird’s-eye-view (BEV) visualizations. We compare the human trajectory, Drive-JEPA [ 39 ] , and ReDrive across representative driving scenarios.
Figure 4 : Sensitivity of the Future Predictor to trajectory conditions with the history representation held fixed. (a) Cosine similarities between future representations predicted under different lateral trajectory offsets. The horizontal axis and point color indicate the offsets of the two trajectory conditions being compared. (b) Pairwise cosine similarity matrix for all trajectory conditions.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5 : Illustration of camera alignment for PhysicalAI-Autonomous-Vehicles Dataset. The original wide-angle camera views are geometrically transformed using per-clip calibration to obtain aligned views with nuPlan-style camera geometry. The alignment reduces the discrepancy in camera projection and view distribution before self-supervised pretraining.
Figure 6 : Additional qualitative comparison of the human trajectory, Drive-JEPA, and ReDrive using front-camera and bird’s-eye-view (BEV) visualizations.
Module
Configuration
Setting
Encoder
Architecture
ViT-L
Depth
24
Hidden dimension
1024
Attention heads
16
Parameters
304M
Predictor
Depth
12
Appendix
Table 8 : Pretraining configuration.
Table 13
Figure 7 : Additional Future Predictor visualizations. Predicted future representations under different lateral trajectory conditions across diverse driving scenes. The representations exhibit consistent action-dependent variations across different scenarios.
Generalization is a central challenge in autonomous driving, as real-world deployment requires robust performance under unseen scenarios, sensor domains, and environmental conditions. Recent world-model-based planning methods have shown strong capabilities in scene understanding and multi-modal future prediction, yet their generalization across datasets and sensor configurations remains limited. In addition, their loosely coupled planning paradigm often leads to poor video-trajectory consistency during visual imagination. To overcome these limitations, we propose DriveVA, a novel autonomous driving world model that jointly decodes future visual forecasts and action sequences in a shared latent generative process. DriveVA inherits rich priors on motion dynamics and physical plausibility from well-pretrained large-scale video generation models to capture continuous spatiotemporal evolution and causal interaction patterns. To this end, DriveVA employs a DiT-based decoder to jointly predict future action sequences (trajectories) and videos, enabling tighter alignment between planning and scene evolution. We also introduce a video continuation strategy to strengthen long-duration rollout consistency. DriveVA achieves an impressive PDM-based planning performance of 90.9 PDM score on the NAVSIM benchmark. Extensive experiments also demonstrate the zero-shot capability and cross-domain generalization of DriveVA, which reduces average L2 error and collision rate by 78.9% and 83.3% on nuScenes and 52.5% and 52.4% on the Bench2Drive built on CARLA v2 compared with the state-of-the-art world-model-based planner.
Mengmeng Liu, Diankun Zhang, Jiuming Liu +7
University of Twente, The Netherlands · Xiaomi EV, China · University of Cambridge, United Kingdom +1
Data scaling plays a pivotal role in the pursuit of general intelligence. However, the prevailing perception-planning paradigm in autonomous driving relies heavily on expensive manual annotations to supervise trajectory planning, which severely limits its scalability. Conversely, although existing perception-free driving world models achieve impressive driving performance, their real-world reasoning ability for planning is solely built on next frame image forecasting. Due to the lack of enough supervision, these models often struggle with comprehensive scene understanding, resulting in unsatisfactory trajectory planning. In this paper, we propose EponaV2, a novel paradigm of driving world models, which achieves high-quality planning with comprehensive future reasoning. Inspired by how human drivers anticipate 3D geometry and semantics, we train our model to forecast more comprehensive future representations, which can be additionally decoded to future geometry and semantic maps. Extracting the 3D and semantic modalities enables our model to deeply understand the surrounding environment, and the future prediction task significantly enhances the real-world reasoning capabilities of EponaV2, ultimately leading to improved trajectory planning. Moreover, inspired by the training recipe of Large Language Models (LLMs), we introduce a flow matching group relative policy optimization mechanism to further improve planning accuracy. The state-of-the-art (SOTA) performances of EponaV2 among perception-free models on three NAVSIM benchmarks (+1.3PDMS, +5.5EPDMS) demonstrate the effectiveness of our methods.
Jiawei Xu, Zhizhou Zhong, Zhijian Shu +8
1PCA Lab, VCIP, College of Computer Science, Nankai University · 2Horizon Robotics · 3HKUST +4
Existing latent world models are typically optimized for future predictability, yet the resulting representations are not necessarily useful for planning in autonomous driving. Predictions are commonly used for pretraining or auxiliary supervision rather than as direct conditioning signals for trajectory generation. We propose ForeDrive, which learns a planning-relevant latent representation and couples it asymmetrically to a Diffusion Transformer (DiT) planner. The planner consumes multi-horizon latent future representations learned with a JEPA-style world model; planning gradients update the shared online encoder, while stop-gradient routing trains the latent predictor with forecasting losses only. Because predicted futures have varying reliability across horizons and BEV trajectories are misaligned with image tokens, we use gated visual fusion, future-status injection, and Trajectory-Adaptive Bias (TAB) to inject future latents as guidance without overriding the current observation. Trained with pure imitation learning and using only the current front-view image as visual input at inference, ForeDrive attains 89.9 PDMS on NAVSIM v1 and 90.0 one-stage EPDMS on NAVSIM v2, without reinforcement learning or an external trajectory scorer.
Sinuo Wang, Zichong Gu, Yuhan Huang +10
Huazhong University of Science and Technology · Shanghai Zaofu Intelligent Technology Co., Ltd. · Tongji University