cs.ROSep 28, 2026

AD-E2E-JEPA: A Joint-Embedding Predictive Architecture For End-to-End Autonomous Driving

Authors: Haoran Zhu, Wancong Zhang, Yann LeCun, Anna Choromanska

Organizations: New York University · AMI Labs

Abstract

Autonomous driving requires \textit{world models} that can understand the physical world, reason and plan, and operate safely. In this paper, we first systematically evaluate existing action-conditioned joint-embedding predictive architecture (JEPA) world models, including LeWM, DINO-WM, and JEPA-WM for end-to-end autonomous driving (E2EAD). To isolate world-model quality from policy learning, we employ a goal-conditioned zero-shot planning setting that evaluates these models using ground-truth future observations as goals, without training any driving policy. We find that existing JEPA-based world models are either accurate for driving but computationally expensive, or computationally efficient but insufficient for planning. To address this trade-off, we propose \textbf{AD-E2E-JEPA}, which introduces a SIGReg-regularized learnable projector applied to projected patch embeddings. The projector reduces the number of planning patches by 16×16\times and the embedding dimension by 4×4\times, achieving a 100×100\times inference speedup while retaining planning performance, with a 0.8-second runtime for an 8-frame rollout over 256 candidate trajectories. \textit{Without} training any driving policy, the world model itself reaches the goals located 20 meters away on average within the displacement of respectively 4.0/2.8 meters, using world-model rollouts over trajectory vocabularies of respectively 256/8,192 candidates. On the NAVSIMv2 benchmark, it achieves 67.3/72.9 EPDMS with multiplicative safety metrics and 84.1/86.5 EPDMS†^{\dagger} without them in goal-conditioned zero-shot planning. Experiments further show that the self-supervised pretrained projector improves downstream imitation learning performance from 80.2 to 85.4 EPDMS. The source code is available at https://github.com/HaoranZhuExplorer/AD-E2E-JEPA

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Auto-JEPA: A Latent World Model of Continuous Intent for End-to-End Autonomous Driving

    Jul 31, 2026Jiwei Yang, Zhengxian Chen, Chaosheng Huang +1Autonomous DrivingVideo Joint Embedding Predictive Architecture

  2. Drive-JEPA: Video JEPA Meets Multimodal Trajectory Distillation for End-to-End Driving

    Jan 29, 2026Linhan Wang, Zichong Yang, Chen Bai +6Causality-Aware End-To-End Autonomous DrivingAutonomous Driving

  3. DriveVA: Video Action Models are Zero-Shot Drivers

    Apr 5, 2026Mengmeng Liu, Diankun Zhang, Jiuming Liu +7Autonomous DrivingDrives