V-JEPA Policy: Building Effective World-Action Models on Predictive Visual Latents
Authors: Yang Zhang, Jiangyuan Zhao, Chenyou Fan, Jiayu Hu, Xiu Yuan, Chenjia Bai, Xiu Li
Organizations: Tsinghua University · Shanghai Jiao Tong University · Fudan University · University of Science and Technology of China · Washington University in St. Louis · The Institute of Artificial Intelligence, China Telecom (TeleAI) · γ-Robotics
World-action models (WAMs) couple future visual-state prediction with action generation. By adapting video generators or image-editing models pretrained at scale, a prominent line of recent WAMs inherits both predictive knowledge and the models in which it was learned. We ask whether a predictive visual latent space induced by large-scale predictive pretraining can instead provide a sufficient foundation for effective WAM learning without inheriting a complete pretrained visual generative model. To answer this question, we introduce V-JEPA Policy, a simple framework that builds a WAM on the latent space of a frozen V-JEPA 2.1 encoder. An instruction-conditioned future-latent predictor and a flow-matching action expert are jointly learned from scratch in a single downstream stage, with the predictor's future-informed context key--value states conditioning action generation. With 0.9B total parameters, of which 0.6B are trainable, V-JEPA Policy achieves competitive performance with representative WAM and vision-language-action baselines across LIBERO, LIBERO-Plus, and RoboCasa-GR1. Comparing visual foundations under the same downstream framework and training budget identifies V-JEPA latents as more effective than the discriminative, reconstructive, and video-understanding-oriented alternatives, particularly under distribution shifts. Beyond task-specific learning, pretraining the predictor on DROID video--instruction pairs without action labels and adapting it into a WAM yields substantial gains in downstream control and out-of-distribution generalization. Together, these findings establish predictive visual latents as a foundation for effective WAM learning from task-specific demonstrations and for transferring future-modeling knowledge acquired from broader in-the-wild videos. Our code is available at https://github.com/breez3young/VJEPA-Policy.
Figures & tables
Figure 1: Overview of V-JEPA Policy. A frozen V-JEPA 2.1 encoder defines the visual latent space. The future predictor jointly processes context tokens and learnable future queries in a single forward pass, providing layer-wise, future-informed context key–value states to a flow-matching action expert. Both modules receive language and proprioceptive conditioning and are jointly trained from scratch with future-latent regression and action flow matching.
LIBERO
LIBERO-Plus
Method
Params. (B)
Action P.T.
Avg.
Spatial
Object
Goal
Long
Avg.
Cam.
Init.
Lang.
Light
Texture
Noise
Layout
OpenVLA-OFT ( Kim et al., 2025 )
7.7
✓
97.1
97.6
98.4
97.9
94.5
69.6
56.4
31.9
79.5
88.7
93.3
75.8
74.2
π0 ( Black et al., 2024 )
3.3
✓
94.1
96.8
98.8
95.8
85.2
53.6
13.8
6.0
58.8
85.0
81.4
79.0
68.9
π0.5 ( Black et al., 2025 )
3.3
✓
96.9
98.8
98.2
98.0
92.4
80.7
64.0
58.0
88.5
96.6
81.4
87.5
85.9
VLA-JEPA ( Sun et al., 2026 )
2.3
✓
97.2
96.2
99.6
97.2
95.8
79.5
63.3
67.1
85.4
95.6
93.6
66.3
85.1
PRTS ( Zhang et al., 2026b )
5
✓
98.4
98.8
99.8
98.4
96.6
84.5
72.5
75.0
90.6
94.8
94.9
87.0
83.1
Table 1: In-distribution performance on LIBERO and generalization performance under distribution shifts on LIBERO-Plus. All scores are success rates (%). Bold and underline indicate the best and second-best reported scores in each column, including ties. Action P.T. denotes additional action-supervised embodied policy pretraining. Cam. , Init. , and Lang. denote camera viewpoint, robot initial state, and language instruction perturbation axes, respectively. † Micro-average estimated from the reported, rounded per-axis success rates of JEPA-WAM, weighted by the official LIBERO-Plus test-set sizes.
Method
Params.
Action P.T.
Success (%)
ABot-M0 ( Yang et al., 2026 )
4.6
✓
58.30
GR00T N1.6 ( NVIDIA et al., 2025 )
3
✓
47.60
π0.5 ( Black et al., 2025 )
3.3
✓
37.00
StarVLA- π ( Community, 2026 )
7.8
✗
43.90
V-JEPA Policy (From Scratch)
0.9
✗
50.92
V-JEPA Policy (Pretrained Predictor)
0.9
✗
55.58
Table 2: RoboCasa-GR1 Tabletop tasks.
Task
V-JEPA Policy (From scratch)
V-JEPA Policy (Predictor pretrained)
FastWAM ( Yuan et al., 2026 )
π0.5 ( Black et al., 2025 )
PRTS ( Zhang et al., 2026b )
Table Cleanup
11/20 (55%)
15/20 (75%)
12/20 (60%)
17/20 (85%)
20/20 (100%)
Saucer Racking
7/20 (35%)
17/20 (85%)
7/20 (35%)
10/20 (50%)
19/20 (95%)
Table 3: Real-world dual-arm success rates over 20 trials per task. Task descriptions, evaluation protocols and detailed fine-tuning schedules are provided in Appendix B .
LIBERO
LIBERO-Plus
Visual encoder
Param.
Avg.
Spatial
Object
Goal
Long
Avg.
Cam.
Init.
Lang.
Light
Tex.
Noise
Layout
Discriminative visual foundations
DINOv2 ViT-L/14 ( Oquab et al., 2024 )
304M
94.90
97.20
98.40
93.40
90.60
67.02
39.02
71.87
54.20
95.01
87.64
62.96
73.11
DINOv3 ViT-L/16 ( Siméoni et al., 2026 )
303M
93.35
95.00
98.00
92.20
88.20
64.28
28.52
63.35
49.77
94.83
82.81
73.39
71.80
Reconstructive visual foundations
WAN2.2 VAE ( Wan et al., 2025 )
150M
92.60
94.40
97.40
96.20
82.40
46.43
35.77
57.29
45.41
32.05
10.13
56.53
73.38
Table 4: Visual foundations, encoder generations, and model scales under a shared downstream recipe. All benchmark entries are success rates (%). Bold and underline indicate the best and second-best reported scores in each column, including ties. Param. gives the approximate frozen visual-encoder size.
Figure 6
Method
Mean Latency (ms)
P95 (ms)
Peak VRAM (GiB)
PRTS
115.89
118.99
9.96
π0.5
159.42
164.46
8.87
FastWAM
202.51
210.72
12.74
V-JEPA Policy
178.17
186.81
4.66
Table 6: Action-prediction core profile on an RTX 4090.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Property
Future predictor
Action expert
Transformer blocks
24
24
Hidden width
1024
512
Visual/action joint-attention heads
16
16
Joint-attention head dimension
64
64
Condition cross-attention heads
16
8
Condition-attention head dimension
64
64
Appendix
Table 7: Reference trainable network configuration. The action expert attends jointly to context-position keys and values from the predictor and its own action tokens. Condition cross-attention is separate from this joint attention.
Table 8: Mean task execution duration over successful rollouts.
Configuration
Future queries
Future loss
LIBERO
LIBERO-Plus
Context-only
✗
✗
92.55
65.86
Future queries, no future loss
✓
✗
91.65
68.81
Full model
✓
✓
97.25
79.25
Appendix
Table 9: Future-prediction controls under the shared downstream recipe. All entries are success rates (%).
Action Core Latency (ms)
PyTorch CUDA VRAM (GiB)
Process Memory
Method
Mean
P50
P95
Allocated
Reserved
Resident (MiB)
PRTS
115.89
115.85
118.99
9.96
10.30
11,014
π0.5
159.42
159.69
164.46
8.87
9.20
9,884
FastWAM
202.51
201.93
210.72
12.74
13.04
13,822
V-JEPA Policy
178.17
177.72
186.81
4.66
4.70
5,276
Appendix
Table 10: Matched three-view action-prediction core profile on an NVIDIA RTX 4090 (300 timed runs across 3 independent processes). Allocated and Reserved denote PyTorch framework allocations; Resident denotes total host OS process memory via NVML.
Figure 5: Success rate versus policy parameter count across three simulation benchmarks and two real-world tasks, with circled points marking nondominated policies. The base V-JEPA Policy lies on the empirical Pareto frontier in every panel.
Task
ABot-M0 ( Yang et al., 2026 )
GR00T N1.6 ( NVIDIA et al., 2025 )
StarVLA- π ( Community, 2026 )
V-JEPA Policy (From Scratch)
V-JEPA Policy (Pretrained Predictor)
PnPBottleToCabinetClose
86
52
26
78
64
PnPCanToDrawerClose
74
13
62
82
86
PnPCupToDrawerClose
48
9
42
48
48
PnPMilkToMicrowaveClose
46
14
50
52
68
PnPPotatoToMicrowaveClose
50
42
42
28
38
PnPWineToCabinetClose
66
17
32
62
64
Appendix
Table 11: Per-task success rates (%) on RoboCasa-GR1. Baseline results are taken from Community (2026) ; Yang et al. (2026) , with StarVLA- π following its official evaluation documentation. GR00T N1.6 task scores are rounded to integers for display; averages are computed over all 24 tasks before rounding. Bold indicates the best result in each row, including ties.
Robust robot control benefits from explicitly modeling state transitions, but video-generation world action models (WAMs) introduce substantial deployment cost. Existing latent WAMs avoid explicit future generation, but often compress predictive representations or separate predictive modeling from the representations used for action generation. We introduce JEPA-WAM, a latent WAM built in a pretrained V-JEPA space, which couples latent transition prediction with continuous action generation through a shared predictor. JEPA-WAM predicts a spatially structured joint current-future target that captures task-shared visual temporal structure between current and future observations, while preserving dense patch-level correspondence. Through the shared predictor, transition supervision directly shapes the backbone, from which dedicated representations are extracted for action prediction. The same design can also be instantiated in pretrained VLA policies while preserving their original perception and action pathways. On LIBERO-Plus, JEPA-WAM achieves 79.2%, the best result without large-scale robot-policy pretraining, while its pretrained π0.5 instantiation reaches 86.3%, achieving the best overall performance. Experiments on RoboTwin 2.0 and real-world bimanual manipulation further demonstrate strong generalization under visual and spatial shifts.
Yihan Lin, Jiawei He, Shifeng Bao +6
School of Information, Renmin University of China, Beijing, China · XYZ Embodied AI, Beijing, China · Key Laboratory of Data Engineering and Knowledge Engineering, Beijing, China +3
World Action Models (WAMs) jointly model action generation and environment dynamics and are mostly built on pretrained Video Diffusion Models (VDMs). In VDM-based WAMs, observations are first encoded by a VAE, and the resulting compressed latents are then processed by large video diffusion backbones to extract effective features for action generation. However, this paradigm ties WAM performance and training cost to large-scale video generation pretraining, limiting WAM efficiency and scalability. In this paper, we theoretically and empirically investigate how visual representations affect action generation in WAMs. Our results show that predictive embeddings from Joint-Embedding Predictive Architecture (JEPA) encoders better support action generation than compressed VAE latents, with I-JEPA performing best in our encoder comparison. Based on these findings, we propose LeWAM, which conditions action generation on JEPA embeddings and models environment evolution by predicting future embeddings in the same space, without relying on a video diffusion backbone. We further find that imitation learning matches demonstrated actions but does not distinguish better actions from worse ones, even though small action deviations can greatly affect task success. To address this limitation without additional environment interaction or the human oversight required for resets and safety, we introduce Demonstration-Guided DPO (DemoDPO), an offline preference refinement stage that derives preference supervision directly from demonstrations. With only 0.4B trainable parameters, LeWAM achieves an average success rate of 92.28% on RoboTwin 2.0, comparable to that of state-of-the-art VLAs and WAMs, and maintains practical effectiveness on real-world manipulation tasks.
Xueji Fang, Boqiang Duan, Hua Wu +2
Zhejiang University · Westlake University · Baidu Inc.
World Action Models (WAMs) connect visual prediction with robot control, but supplying predictive context often requires expensive future-video generation. Direct policies avoid this cost but lack an explicit interface for accessing future-indexed predictive information. We introduce ForeWAM, a World Action Model that separates forecasting from rendering to expose and shape latent predictive context for efficient control. Its core mechanism, Future-KV, performs a single Video DiT prefill over the current visual latent and noise-initialized future slots, then reuses the resulting key-value states throughout action denoising. To make this context relevant to control, we introduce dynamics registers supervised by latent actions from a frozen teacher during training, encouraging representations of interaction-induced transitions. This reusable context supports a lightweight, single-layer action decoder. We evaluate ForeWAM on LIBERO, LIBERO-Plus, RoboCasa, and real-world manipulation tasks. Without additional policy-level embodied pretraining, ForeWAM improves RoboCasa success by 9.7 percentage points over Fast-WAM at the same budget of 50 demonstrations per task, reaching 59.2%. With a single-layer decoder, it achieves 77.6% success on LIBERO-Plus and reduces policy-query latency to 88.7 ms on an NVIDIA A800, delivering a 6.27-fold speedup over Fast-WAM. These results show that latent predictive computation provides useful foresight for robust, efficient control without explicit future-video generation.
Jiakai Huang, Zhongbo Wu, Siyu Xu +5
Shanghai Jiao Tong University · ACE Robotics · Nanyang Technological University