cs.CVSep 29, 2026

V-JEPA Policy: Building Effective World-Action Models on Predictive Visual Latents

Authors: Yang Zhang, Jiangyuan Zhao, Chenyou Fan, Jiayu Hu, Xiu Yuan, Chenjia Bai, Xiu Li

Organizations: Tsinghua University · Shanghai Jiao Tong University · Fudan University · University of Science and Technology of China · Washington University in St. Louis · The Institute of Artificial Intelligence, China Telecom (TeleAI) · γ-Robotics

Abstract

World-action models (WAMs) couple future visual-state prediction with action generation. By adapting video generators or image-editing models pretrained at scale, a prominent line of recent WAMs inherits both predictive knowledge and the models in which it was learned. We ask whether a predictive visual latent space induced by large-scale predictive pretraining can instead provide a sufficient foundation for effective WAM learning without inheriting a complete pretrained visual generative model. To answer this question, we introduce V-JEPA Policy, a simple framework that builds a WAM on the latent space of a frozen V-JEPA 2.1 encoder. An instruction-conditioned future-latent predictor and a flow-matching action expert are jointly learned from scratch in a single downstream stage, with the predictor's future-informed context key--value states conditioning action generation. With 0.9B total parameters, of which 0.6B are trainable, V-JEPA Policy achieves competitive performance with representative WAM and vision-language-action baselines across LIBERO, LIBERO-Plus, and RoboCasa-GR1. Comparing visual foundations under the same downstream framework and training budget identifies V-JEPA latents as more effective than the discriminative, reconstructive, and video-understanding-oriented alternatives, particularly under distribution shifts. Beyond task-specific learning, pretraining the predictor on DROID video--instruction pairs without action labels and adapting it into a WAM yields substantial gains in downstream control and out-of-distribution generalization. Together, these findings establish predictive visual latents as a foundation for effective WAM learning from task-specific demonstrations and for transferring future-modeling knowledge acquired from broader in-the-wild videos. Our code is available at https://github.com/breez3young/VJEPA-Policy.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling

    Aug 10, 2026Yihan Lin, Jiawei He, Shifeng Bao +6Efficient World-Action ModelRobot Policies

  2. Latent evolving World Action Model

    Sep 23, 2026Xueji Fang, Boqiang Duan, Hua Wu +2Action GenerationWorld Models

  3. Foresight Without Seeing: Latent Futures for World Action Models

    Aug 12, 2026Jiakai Huang, Zhongbo Wu, Siyu Xu +5Action PredictionForesight