World action models (WAMs) have emerged as an important class of generalist robot policies, aiming to transfer video prediction priors to action learning. However, we find that existing WAMs frequently produce incorrect motion predictions in open-ended environment, leading to erroneous actions. We attribute this limitation to two factors: (1) base video models are not optimized for manipulation, and (2) naively incorporating action components into video models can substantially degrade their generalization capabilities. We introduce Video Prediction Policy 2 (VPP2), a WAM that enables strong zero-shot generalization in both video prediction and action generation. First, we curate a large-scale, diverse dataset of manipulation videos to continue pretraining the base video foundation model. We annotate video clips with detailed captions and perform \textit{event-level} video pretraining to promote generalization across open-ended manipulation tasks. Second, we post-train and distill the video model into a single-step visual planner with fixed prediction horizon. Finally, we introduce action module via a mixture-of-transformers (MoT) architecture to learn implicit inverse dynamics model. Experiments demonstrate three key results: (1) VPP2-14B outperforms Cosmos3-64B by 11.0% points in video prediction instruction-following success rate on open-ended tasks; (2) VPP2 surpasses the strongest baseline by 18.5% points in success rate on real-world zero-shot ALOHA manipulation tasks; and (3) following benchmark-specific post-training, VPP2 achieves the highest success rates among evaluated methods on the challenging LIBERO-Pro, LIBERO-OOD, and RoboDojo benchmarks.
Figures & tables
Figure 1: A typical robot policy maps an simple instruction and an observation into an short action chunk. As datasets scale, this mapping becomes increasingly multi-modes and uncertain, leading models to learn spurious short-horizon correlations. VPP2 first establishes consistent semantic-to-trajectory mappings via next event video prediction with detailed caption, then post-train and distill model to generate short action chunk.
Table 1: We use various types of manipulation datasets and perform different filtering strategy for different datasets to ensure diversity and high-quality.
Figure 2: An example of our data process pipeline. The central objective is to reduce uncertainty in future prediction, thereby encouraging the model to learn a consistent mapping from conditioning information to sub-task trajectories. To this end, each caption includes detailed task descriptions, explicit End-Effector Identification, visible target object, and camera-view changes.
Figure 3: (a) Bounding boxes show the aligned workspaces of different datasets. (b) Aligned end-effector coordinate frames.
Figure 4: The VPP2 training pipeline is designed to maximize generalization. Stage 1 uses large-scale, event-level video pretraining to learn a generalizable video model for manipulation. Stage 2 post-trains and distills this model to predict long-horizon video chunks spanning 8 seconds in a single forward pass, taking approximately 0.1 seconds. Finally, Stage 3 trains an action expert to generate 2-second action chunks conditioned on the one-step video latents.
Figure 5: VPP2 video predictions for human-hand and robot manipulation. Trained on large-scale, diverse manipulation datasets, VPP2 generalizes across human hands and a wide range of robot embodiments. For robot manipulation, VPP2 jointly predicts one to three camera views arranged in a T-shaped composite. Due to space constraints, single-step video predictions after distillation are shown in Figure 10 from Appendix.
Figure 7
Figure 7: Comparisons on instruction following capability between VPP2, Wan-14B, Cosmos3-64B. VPP2 demonstrates better instruction following on complex tasks requiring spatial understanding.
Figure 8: Zero-shot success rates on the real-world ALOHA across 10 randomly selected task categories. For a fair comparison, we finetune π0.5 and FastWAM on all ALOHA datasets exposed in VPP2’s training data and use the 14B variant of FastWAM.
LIBERO-ID
LIBERO-Pro
LIBERO-OOD
Method
Overall
Position
Task
Overall
Spatial
Object
Goal
Overall
π0 ( Black et al., 2024 )
94.2
0.5
0.0
0.3
0.7
0.3
4.3
1.7
π0.5 ( Intelligence et al., 2025 )
96.9
20.8
1.3
11.0
36.7
2.3
41.7
26.8
MolmoAct ( Lee et al., 2025 )
86.6
1.5
1.5
1.5
–
–
–
–
X-VLA ( Zheng et al., 2026 )
98.1
0.8
6.8
3.8
–
–
–
–
AtomVLA ( Sun et al., 2026 )
97.0
7.3
5.3
6.3
–
–
–
–
Table 3: Post-training on LIBERO-ID, LIBERO-Pro, and LIBERO-OOD benchmarks. All models are trained exclusively on the four standard LIBERO suites ( Liu et al., 2023 ) and are evaluated under three complementary settings.
π0
π0.5
X-VLA
Fast-WAM
OpenWAM- α
Xiaomi-Robotics-1
GPT-6-Astra
VPP2 (Ours)
Avg. Score
3.48
11.44
10.13
3.48
17.18
20.07
28.97
35.51
Success Rate (%)
1.53
6.93
6.52
2.03
11.92
13.93
22.48
29.47
Table 4: Post-training results on the RoboDojo simulation benchmark. Average score captures partial task progress, and success rate measures binary task completion; both are averaged over the five capability dimensions. Baseline results are taken from the official leaderboard.
Figure 9: Success rates with and without VLM-based subtask planning (50 trials per task). Average is the five-task mean; the baseline checkpoint matches Table 4 .
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 10: Additional video prediction results on open-ended tasks. We compare single-step video predictions before and after distillation. Distillation enables high-quality video prediction with just one sampling step.
Method
Spatial
Object
Goal
Long
Avg.
π0 ( Black et al., 2024 )
96.8
98.8
95.8
85.2
94.2
π0.5 ( Intelligence et al., 2025 )
98.8
98.2
98.0
92.4
96.9
MolmoAct ( Lee et al., 2025 )
87.0
95.4
87.6
77.2
86.6
X-VLA ( Zheng et al., 2026 )
98.2
98.6
97.8
97.6
98.1
AtomVLA ( Sun et al., 2026 )
96.4
99.6
97.6
94.4
97.0
Cosmos-Policy ( Kim et al., 2026 )
98.1
100.0
98.2
97.6
98.5
Appendix
Table 5: Per-suite success rates (%) on the standard LIBERO benchmark. Best results in each column are in bold.
Generalization
Method
Std.
Rand.
Precision
Long-Horizon
Memory
Open
Avg.
π0 ( Black et al., 2024 )
7.18
0.71
3.56
6.19
3.47
0.25
3.48
π0.5 ( Intelligence et al., 2025 )
20.93
5.82
12.40
23.54
5.89
1.98
11.44
X-VLA ( Zheng et al., 2026 )
17.90
3.04
18.32
16.53
4.76
0.55
10.13
Fast-WAM ( Yuan et al., 2026 )
4.33
0.34
1.96
9.14
3.55
0.42
3.48
OpenWAM- α ( Wang et al., 2026 )
33.16
8.26
18.45
34.93
10.41
1.41
17.18
Appendix
Table 6: Per-dimension average score on the RoboDojo simulation benchmark. Generalization is evaluated under standard (Std.) and randomized (Rand.) settings. Avg. is the mean over the five capability dimensions, where the generalization score is the mean of the Std. and Rand. settings. For VPP2, we report the generalization result pooled over both settings (25 episodes each, following the official protocol), which equals their mean. Baseline results are taken from the official leaderboard.
Generalization
Method
Std.
Rand.
Precision
Long-Horizon
Memory
Open
Avg.
π0 ( Black et al., 2024 )
4.89
0.22
0.75
2.00
2.11
0.25
1.53
π0.5 ( Intelligence et al., 2025 )
14.89
1.44
5.50
14.67
4.67
1.67
6.93
X-VLA ( Zheng et al., 2026 )
12.22
1.33
12.00
9.75
3.56
0.50
6.52
Fast-WAM ( Yuan et al., 2026 )
2.11
0.11
0.00
5.17
3.44
0.42
2.03
OpenWAM- α ( Wang et al., 2026 )
25.56
4.11
9.25
25.33
9.11
1.08
11.92
Appendix
Table 7: Per-dimension success rate (%) on the RoboDojo simulation benchmark, computed in the same way as Table 6 .