Learning to predict how the world evolves can provide vision-language-action (VLA) policies with predictive context for long-horizon control, but its effectiveness depends on what future representation is modeled and how it conditions action generation. We introduce PLaW-VLA, which models task-relevant future states in a pretrained prediction-oriented representation space, reducing the need to predict control-irrelevant visual details. Built on a Mixture-of-Transformers architecture, PLaW-VLA conditions action generation on observation history, current task semantics, and predicted future states through structured causal attention. Experiments show a +11.8 percentage-point (pp) gain over reactive policies on RoboTwin Hard Horizon III and a +1.77 pp gain over reconstruction-oriented latent prediction on zero-shot LIBERO-Plus, supporting improved long-horizon control and generalization under distribution shift, respectively. By avoiding low-level visual reconstruction, PLaW-VLA lowers the burden of future prediction, enabling a lightweight latent world model with parallel future prediction and about 1/19 the inference latency of generative world-action modeling at comparable policy performance.
Figures & tables
Figure 1: PLaW-VLA. Videos and robot trajectories train a latent world model whose predicted future representations condition robot actions. We evaluate this approach on simulated and real-world manipulation, including multi-step and deformable-object tasks.
Figure 2: PLaW-VLA overview. Three experts share structured attention for semantic grounding, latent future prediction, and continuous action generation. The world model predicts future representations in a frozen V-JEPA 2 latent space using learnable queries. The action expert attends to the resulting future prediction tokens together with the observed context.
Figure 3: Cross-expert attention.
Method
LIBERO-Spatial
LIBERO-Object
LIBERO-Goal
LIBERO-Long
Average
π0 [ 4 ]
96.8
98.8
95.8
85.2
94.2
π0.5 [ 33 ]
98.8
98.2
98.0
92.4
96.9
FLOWER [ 35 ]
97.5
99.1
96.1
94.9
96.9
JALA-dino [ 25 ]
96.0
98.2
97.4
96.0
96.9
OpenVLA-OFT [ 16 ]
97.6
98.4
97.9
94.5
97.1
UniVLA [ 43 ]
95.4
98.8
93.6
94.0
95.5
Table 1: LIBERO results. Task success rates (%). The best and second-best suite scores are indicated by boldface and underlining , respectively. Shading marks methods with explicit future-observation prediction.
Method
Robot
Layout
Light
Background
Language
Noise
Camera
Average
π0 [ 4 ]
6.0
68.9
85.0
81.4
58.8
79.0
13.8
53.6
π0 -Fast [ 4 , 32 ]
21.6
68.8
73.2
73.2
61.0
74.4
65.1
61.6
RIPT-VLA [ 40 ]
31.2
74.2
88.4
91.6
77.6
73.5
55.2
68.4
OpenVLA-OFT [ 16 ]
31.9
74.2
88.7
93.3
79.5
75.8
56.4
69.6
WorldVLA [ 9 ]
27.9
38.0
43.7
17.1
41.6
10.9
0.1
25.0
UniVLA [ 6 ]
46.2
31.9
69.0
81.0
69.6
21.2
1.8
42.9
Table 2: LIBERO-Plus results. Zero-shot success rates (%); averages use official task weights (UniVLA recomputed). Boldface and underlining indicate best and second-best scores. Shading marks explicit future-observation prediction.
Metric
X-VLA [ 47 ]
π0 [ 4 ]
π0.5 [ 33 ]
PLaW-VLA (Ours)
Easy
Hard
Easy
Hard
Easy
Hard
Easy
Hard
Horizon I
81.6
82.5
66.5
61.6
85.1
80.2
87.9 (+2.8)
87.8 (+5.3)
Horizon II
59.3
55.9
66.1
54.7
79.3
73.0
82.1 (+2.8)
75.5 (+2.5)
Horizon III
61.2
66.0
61.6
50.2
78.6
67.4
82.4 (+3.8)
79.2 (+11.8)
Average
72.9
72.8
65.9
58.4
82.7
76.8
85.6 (+2.9)
83.2 (+6.4)
Table 3: Results on RoboTwin. Success rates (%) are reported. The best results are shown in bold , the second-best results are underlined , and values in parentheses indicate the absolute gains of PLaW-VLA over the second-best method.
Figure 4: Real-world robot evaluation. Task progress across four manipulation tasks on a dual-arm wheeled humanoid robot. * denotes zero-shot variants with unseen object appearances or receptacles.
Table 4: Predictive world-model contribution and use for control. Success rates averaged over three evaluation seeds (%). Parentheses give percentage-point changes relative to PLaW-VLA (full). Group (a) omits Stage I and use matched Stage II data. Masking is inference-only; a dash means not evaluated.
Figure 5: Inference efficiency. LIBERO success versus single-H100 inference latency under matched settings. Marker size indicates model parameter count. The mimic-video score averages three suites.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Method
LIBERO-Spatial
LIBERO-Object
LIBERO-Goal
LIBERO-Long
Average
TraceVLA [ 48 ]
84.6
85.2
75.1
54.1
74.8
Octo [ 30 ]
78.9
85.7
84.6
51.1
75.1
OpenVLA [ 17 ]
84.7
88.4
79.2
53.7
76.5
SpatialVLA [ 34 ]
88.2
89.9
78.6
55.5
78.1
LAPA [ 45 ]
83.4
87.6
78.2
68.8
79.5
WorldVLA [ 9 ]
87.6
96.2
83.4
60.0
81.8
Appendix
Table 5: Performance comparison on LIBERO. CoT-VLA and MolmoAct averages are recomputed from the displayed suite scores. The mimic-video average covers only the three reported suites.
Method
Robot
Layout
Light
Background
Language
Noise
Camera
Average
OpenVLA [ 17 ]
3.5
28.5
8.1
34.8
23.0
15.2
0.8
15.6
NORA [ 15 ]
37.0
62.1
45.7
58.6
65.1
12.8
2.2
39.0
π0 [ 4 ]
6.0
68.9
85.0
81.4
58.8
79.0
13.8
53.6
π0 -Fast [ 4 , 32 ]
21.6
68.8
73.2
73.2
61.0
74.4
65.1
61.6
RIPT-VLA [ 40 ]
31.2
74.2
88.4
91.6
77.6
73.5
55.2
68.4
OpenVLA-OFT [ 16 ]
31.9
74.2
88.7
93.3
79.5
75.8
56.4
69.6
Appendix
Table 6: Full LIBERO-Plus comparison. Zero-shot task success rates (%); the average is weighted by official task counts. UniVLA’s average is recomputed from its perturbation scores. The best and second-best perturbation scores are indicated by boldface and underlining , respectively.
Task
Horizon
X-VLA [ 47 ]
π0 [ 4 ]
π0.5 [ 33 ]
Ours
Easy
Hard
Easy
Hard
Easy
Hard
Easy
Hard
Adjust Bottle
I
100%
99%
99%
95%
100%
99%
96%
99%
Beat Block Hammer
I
92%
88%
79%
84%
96%
93%
96%
98%
Blocks Ranking RGB
III
83%
83%
80%
63%
92%
85%
95%
89%
Blocks Ranking Size
III
67%
74%
14%
5%
49%
26%
69%
53%
Click Alarmclock
I
99%
99%
77%
68%
98%
89%
95%
91%
Appendix
Table 7: Evaluation on RoboTwin 2.0 Simulation: Easy vs Hard (50 tasks). RoboTwin 2.0 is a challenging bimanual manipulation benchmark requiring coordinated dual-arm control. ”Easy” uses fixed initial configurations, while ”Hard” involves randomized object poses and scene layouts.
Method
ckpt
LIBERO-Spatial
LIBERO-Object
LIBERO-Goal
LIBERO-Long
Average
Without Stages I–II
30k
97.00
98.40
91.20
89.60
94.05
PLaW-VLA
30k
98.90
99.00
97.10
94.60
97.40
Δ
30k
+1.90
+0.60
+5.90
+5.00
+3.35
Without Stages I–II
50k
97.40
98.80
94.40
93.20
95.95
PLaW-VLA
50k
98.90
98.90
98.30
91.70
96.95
Δ
50k
+1.50
+0.10
+3.90
-1.50
+1.00
Appendix
Table 8: Effect of Stages I–II pretraining on LIBERO. Matched Stage III checkpoints; all values are success rates (%).
Method
ckpt
Robot
Layout
Light
Background
Language
Noise
Camera
Average
Without Stages I–II
50k
63.23
75.08
86.78
90.52
65.91
61.40
44.59
67.79
PLaW-VLA
50k
68.40
84.00
90.20
91.40
75.10
53.30
45.10
70.62
Δ
50k
+5.17
+8.92
+3.42
+0.88
+9.19
-8.10
+0.51
+2.83
Appendix
Table 9: Effect of Stages I–II pretraining on LIBERO-Plus at 50k steps. Success rates (%); the average is weighted by perturbation task counts.
Figure 6: Stage III training dynamics. Training loss and LIBERO success rate across checkpoints, showing stable downstream performance after 30k updates.
Method
LIBERO-Spatial
LIBERO-Object
LIBERO-Goal
LIBERO-Long
Average
Without all future states
89.60
95.40
97.40
89.80
93.05
Without far-future states
89.60
97.40
99.80
95.00
95.45
Without near-future states
92.40
97.00
99.00
94.80
95.80
PLaW-VLA
98.90
99.00
97.10
94.60
97.40
Appendix
Table 10: Inference-time masking of future predictions in LIBERO. All values are success rates (%).
Figure 7: Visualization of real-robot fine-tuning tasks across different objects.
Hyperparameters
Stage I
Stage II
Stage III
LIBERO
RoboTwin
Real-world
Dataset Regime
Dwarm
Dpre
Dpost
Dpost
Dpost
Trainable Modules
LWM only
Full MoT
Full MoT
Full MoT
Full MoT
Batch Size
512
256
256
256
256
Learning Rate
5.0×10−5
5.0×10−5
5.0×10−5
5.0×10−5
5.0×10−5
LR Scheduler
Constant
Constant
Warmup + constant
Warmup + constant
Warmup + constant
Appendix
Table 11: Training configurations across the three stages of PLaW-VLA. Stage I corresponds to Dwarm , Stage II to Dpre , and Stage III to task-specific Dpost . All stages use the same temporal sampling frequency, and all action-annotated robot data are represented in a unified end-effector (EEF) delta-action space. Stage III uses a linear warmup followed by a constant learning rate. LIBERO warms up for 10k steps, trains for 50k steps, and reports the 30k checkpoint.
Table 12: Heterogeneous data corpus used for Stage I and Stage II. Stage I uses visual sequences from all five sources without action annotations. Stage II uses the four robot datasets with actions aligned to a shared end-effector representation. EgoDex is used only in Stage I. Stage III settings appear in Table 11 .
Vision-Language-Action (VLA) models have emerged as a promising paradigm for building embodied agents that ground perception and language into action. However, most existing approaches rely on direct action prediction, lacking the ability to reason over long-horizon trajectories and evaluate their consequences, which limits performance in complex decision-making tasks. In this work, we introduce World-Value-Action (WAV) model, a unified framework that enables implicit planning in VLA systems. Rather than performing explicit trajectory optimization, WAV model learn a structured latent representation of future trajectories conditioned on visual observations and language instructions. A learned world model predicts future states, while a trajectory value function evaluates their long-horizon utility. Action generation is then formulated as inference in this latent space, where the model progressively concentrates probability mass on high-value and dynamically feasible trajectories. We provide a theoretical perspective showing that planning directly in action space suffers from an exponential decay in the probability of feasible trajectories as the horizon increases. In contrast, latent-space inference reshapes the search distribution toward feasible regions, enabling efficient long-horizon decision making. Extensive simulations and real-world experiments demonstrate that the WAV model consistently outperforms state-of-the-art methods, achieving significant improvements in task success rate, generalization ability, and robustness, especially in long-horizon and compositional scenarios. Code is available at https://github.com/Win-commit/WAV.
Runze Li, Hongyin Zhang, Junxi Jin +5
Westlake University Hangzhou, China · Nanjing University Suzhou Campus, 1520 Taihu Avenue, Suzhou, China
Vision-Language-Action models (VLAs) leverage large-scale vision-language pretraining for semantic robot control, but often lack explicit foresight into how robot actions change the scene. World-Action Models (WAMs) address this limitation by conditioning policies on predicted futures, yet existing approaches typically rely on computationally expensive video generation with substantial pixel-level redundancy. We present LaWAM, a Latent World Action Model that exposes predictive dynamics to robot policies through compact latent visual subgoals instead of reconstructed future video. At the core of LaWAM is a latent-action-conditioned Latent World Model (LaWM). We obtain LaWM by training a latent action model in the latent space of a pretrained vision foundation model and repurposing its forward decoder to predict future observation features for scene evolution. LaWAM then conditions action generation on these predicted latent visual subgoals to enable dynamics-aware robot control. LaWAM achieves state-of-the-art or competitive success rates (SRs) across LIBERO (98.6% SR), RoboTwin (91.22% SR), and real-world manipulation tasks while retaining low-latency inference. LaWAM runs in 187 ms per action-chunk prediction and achieves up to 24x lower wall-clock latency than pixel-space WAMs.
Jialei Chen, Kai Wang, Kang Chen +9
Jilin University · Zhongguancun Academy · Nankai University +4
Future prediction is increasingly used to improve vision-language-action (VLA) policies, based on the premise that anticipating scene evolution encourages representations useful for control. However, forecast quality alone does not establish that a policy has learned a better representation for action. This distinction matters under distribution shift, where successful control depends on preserving spatial state and likely scene change beyond familiar configurations. We study what determines whether predictive supervision improves the visual representation used by a VLA policy. Through controlled comparisons with matched target constructions, prediction horizons, and training conditions, we find that different prediction interfaces produce markedly different forecasts and visual representations, including in the spatial, dynamics, and action information that transfers beyond familiar scenes. We trace these differences to how predictive errors shape the policy's visual stream. Consistent with this controlled finding, VLA policies trained with more direct, scene-matched future supervision show stronger robustness under simulated and physical distribution shifts. Together, our results frame future prediction as a representation-learning design problem whose value for control depends on whether its supervision reaches the representations through which the policy acts.
Hanseul Kim, Jewon Yeom, Youngjoon Jeong +2
Graduate School of Data Science Seoul National University