V-JEPA Policy: Building Effective World-Action Models on Predictive Visual Latents
Authors: Yang Zhang, Jiangyuan Zhao, Chenyou Fan, Jiayu Hu, Xiu Yuan, Chenjia Bai, Xiu Li
Organizations: Tsinghua University · Shanghai Jiao Tong University · Fudan University · University of Science and Technology of China · Washington University in St. Louis · The Institute of Artificial Intelligence, China Telecom (TeleAI) · γ-Robotics
World-action models (WAMs) couple future visual-state prediction with action generation. By adapting video generators or image-editing models pretrained at scale, a prominent line of recent WAMs inherits both predictive knowledge and the models in which it was learned. We ask whether a predictive visual latent space induced by large-scale predictive pretraining can instead provide a sufficient foundation for effective WAM learning without inheriting a complete pretrained visual generative model. To answer this question, we introduce V-JEPA Policy, a simple framework that builds a WAM on the latent space of a frozen V-JEPA 2.1 encoder. An instruction-conditioned future-latent predictor and a flow-matching action expert are jointly learned from scratch in a single downstream stage, with the predictor's future-informed context key--value states conditioning action generation. With 0.9B total parameters, of which 0.6B are trainable, V-JEPA Policy achieves competitive performance with representative WAM and vision-language-action baselines across LIBERO, LIBERO-Plus, and RoboCasa-GR1. Comparing visual foundations under the same downstream framework and training budget identifies V-JEPA latents as more effective than the discriminative, reconstructive, and video-understanding-oriented alternatives, particularly under distribution shifts. Beyond task-specific learning, pretraining the predictor on DROID video--instruction pairs without action labels and adapting it into a WAM yields substantial gains in downstream control and out-of-distribution generalization. Together, these findings establish predictive visual latents as a foundation for effective WAM learning from task-specific demonstrations and for transferring future-modeling knowledge acquired from broader in-the-wild videos. Our code is available at https://github.com/breez3young/VJEPA-Policy.
Figures & tables
Figure 1: Overview of V-JEPA Policy. A frozen V-JEPA 2.1 encoder defines the visual latent space. The future predictor jointly processes context tokens and learnable future queries in a single forward pass, providing layer-wise, future-informed context key–value states to a flow-matching action expert. Both modules receive language and proprioceptive conditioning and are jointly trained from scratch with future-latent regression and action flow matching.
LIBERO
LIBERO-Plus
Method
Params. (B)
Action P.T.
Avg.
Spatial
Object
Goal
Long
Avg.
Cam.
Init.
Lang.
Light
Texture
Noise
Layout
OpenVLA-OFT ( Kim et al., 2025 )
7.7
✓
97.1
97.6
98.4
97.9
94.5
69.6
56.4
31.9
79.5
88.7
93.3
75.8
74.2
π0 ( Black et al., 2024 )
3.3
✓
94.1
96.8
98.8
95.8
85.2
53.6
13.8
6.0
58.8
85.0
81.4
79.0
68.9
π0.5 ( Black et al., 2025 )
3.3
✓
96.9
98.8
98.2
98.0
92.4
80.7
64.0
58.0
88.5
96.6
81.4
87.5
85.9
VLA-JEPA ( Sun et al., 2026 )
2.3
✓
97.2
96.2
99.6
97.2
95.8
79.5
63.3
67.1
85.4
95.6
93.6
66.3
85.1
PRTS ( Zhang et al., 2026b )
5
✓
98.4
98.8
99.8
98.4
96.6
84.5
72.5
75.0
90.6
94.8
94.9
87.0
83.1
Table 1: In-distribution performance on LIBERO and generalization performance under distribution shifts on LIBERO-Plus. All scores are success rates (%). Bold and underline indicate the best and second-best reported scores in each column, including ties. Action P.T. denotes additional action-supervised embodied policy pretraining. Cam. , Init. , and Lang. denote camera viewpoint, robot initial state, and language instruction perturbation axes, respectively. † Micro-average estimated from the reported, rounded per-axis success rates of JEPA-WAM, weighted by the official LIBERO-Plus test-set sizes.
Method
Params.
Action P.T.
Success (%)
ABot-M0 ( Yang et al., 2026 )
4.6
✓
58.30
GR00T N1.6 ( NVIDIA et al., 2025 )
3
✓
47.60
π0.5 ( Black et al., 2025 )
3.3
✓
37.00
StarVLA- π ( Community, 2026 )
7.8
✗
43.90
V-JEPA Policy (From Scratch)
0.9
✗
50.92
V-JEPA Policy (Pretrained Predictor)
0.9
✗
55.58
Table 2: RoboCasa-GR1 Tabletop tasks.
Task
V-JEPA Policy (From scratch)
V-JEPA Policy (Predictor pretrained)
FastWAM ( Yuan et al., 2026 )
π0.5 ( Black et al., 2025 )
PRTS ( Zhang et al., 2026b )
Table Cleanup
11/20 (55%)
15/20 (75%)
12/20 (60%)
17/20 (85%)
20/20 (100%)
Saucer Racking
7/20 (35%)
17/20 (85%)
7/20 (35%)
10/20 (50%)
19/20 (95%)
Table 3: Real-world dual-arm success rates over 20 trials per task. Task descriptions, evaluation protocols and detailed fine-tuning schedules are provided in Appendix B .
LIBERO
LIBERO-Plus
Visual encoder
Param.
Avg.
Spatial
Object
Goal
Long
Avg.
Cam.
Init.
Lang.
Light
Tex.
Noise
Layout
Discriminative visual foundations
DINOv2 ViT-L/14 ( Oquab et al., 2024 )
304M
94.90
97.20
98.40
93.40
90.60
67.02
39.02
71.87
54.20
95.01
87.64
62.96
73.11
DINOv3 ViT-L/16 ( Siméoni et al., 2026 )
303M
93.35
95.00
98.00
92.20
88.20
64.28
28.52
63.35
49.77
94.83
82.81
73.39
71.80
Reconstructive visual foundations
WAN2.2 VAE ( Wan et al., 2025 )
150M
92.60
94.40
97.40
96.20
82.40
46.43
35.77
57.29
45.41
32.05
10.13
56.53
73.38
Table 4: Visual foundations, encoder generations, and model scales under a shared downstream recipe. All benchmark entries are success rates (%). Bold and underline indicate the best and second-best reported scores in each column, including ties. Param. gives the approximate frozen visual-encoder size.
Figure 6
Method
Mean Latency (ms)
P95 (ms)
Peak VRAM (GiB)
PRTS
115.89
118.99
9.96
π0.5
159.42
164.46
8.87
FastWAM
202.51
210.72
12.74
V-JEPA Policy
178.17
186.81
4.66
Table 6: Action-prediction core profile on an RTX 4090.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Property
Future predictor
Action expert
Transformer blocks
24
24
Hidden width
1024
512
Visual/action joint-attention heads
16
16
Joint-attention head dimension
64
64
Condition cross-attention heads
16
8
Condition-attention head dimension
64
64
Appendix
Table 7: Reference trainable network configuration. The action expert attends jointly to context-position keys and values from the predictor and its own action tokens. Condition cross-attention is separate from this joint attention.
Table 8: Mean task execution duration over successful rollouts.
Configuration
Future queries
Future loss
LIBERO
LIBERO-Plus
Context-only
✗
✗
92.55
65.86
Future queries, no future loss
✓
✗
91.65
68.81
Full model
✓
✓
97.25
79.25
Appendix
Table 9: Future-prediction controls under the shared downstream recipe. All entries are success rates (%).
Action Core Latency (ms)
PyTorch CUDA VRAM (GiB)
Process Memory
Method
Mean
P50
P95
Allocated
Reserved
Resident (MiB)
PRTS
115.89
115.85
118.99
9.96
10.30
11,014
π0.5
159.42
159.69
164.46
8.87
9.20
9,884
FastWAM
202.51
201.93
210.72
12.74
13.04
13,822
V-JEPA Policy
178.17
177.72
186.81
4.66
4.70
5,276
Appendix
Table 10: Matched three-view action-prediction core profile on an NVIDIA RTX 4090 (300 timed runs across 3 independent processes). Allocated and Reserved denote PyTorch framework allocations; Resident denotes total host OS process memory via NVML.
Figure 5: Success rate versus policy parameter count across three simulation benchmarks and two real-world tasks, with circled points marking nondominated policies. The base V-JEPA Policy lies on the empirical Pareto frontier in every panel.
Task
ABot-M0 ( Yang et al., 2026 )
GR00T N1.6 ( NVIDIA et al., 2025 )
StarVLA- π ( Community, 2026 )
V-JEPA Policy (From Scratch)
V-JEPA Policy (Pretrained Predictor)
PnPBottleToCabinetClose
86
52
26
78
64
PnPCanToDrawerClose
74
13
62
82
86
PnPCupToDrawerClose
48
9
42
48
48
PnPMilkToMicrowaveClose
46
14
50
52
68
PnPPotatoToMicrowaveClose
50
42
42
28
38
PnPWineToCabinetClose
66
17
32
62
64
Appendix
Table 11: Per-task success rates (%) on RoboCasa-GR1. Baseline results are taken from Community (2026) ; Yang et al. (2026) , with StarVLA- π following its official evaluation documentation. GR00T N1.6 task scores are rounded to integers for display; averages are computed over all 24 tasks before rounding. Bold indicates the best result in each row, including ties.
School of Information, Renmin University of China, Beijing, China · XYZ Embodied AI, Beijing, China · Key Laboratory of Data Engineering and Knowledge Engineering, Beijing, China +3