World models offer a promising way to help robots understand how the physical world evolves and plan complex behaviours through imagination. Yet existing studies mainly demonstrate what these models can accomplish, leaving unclear when their predictions remain useful for planning and where they fail. We study this question using action-conditioned predictors built on five frozen self-supervised visual backbones: V-JEPA 2, V-JEPA 2.1, VideoMAEv2, VideoPrism, and DINOv2. We use frozen backbones to test representations intended to transfer across environments. We evaluate these models on diverse Meta-World manipulation tasks and real-robot interactions from BridgeData V2. We find that a world model guides action selection reliably only when the goal lies within, or slightly beyond, the trajectory it imagines during planning. With five-step rollouts, the length the predictor was trained on, the world model ranks actions reliably only for targets five to ten control steps ahead, whereas task goals lie 16 to 53 steps away. Neither an 81-fold larger predictor nor longer-rollout training extends this range; the encoder affects both range and closed-loop success, with V-JEPA 2.1 performing most consistently. More fundamentally, the limit persists under perfect prediction: using the real simulator, success falls from 92% to 41% as the target moves from five to twenty steps ahead of a five-step rollout. Planning therefore requires either longer imagined trajectories or closer subgoals. For distant goals, pure imagination succeeds in 23% of episodes, planning with feedback (MPC) raises success to 30%, imagining as far as the goal to 47%, and nearby expert subgoals to 76%. Used within its plannable range, a world model can also improve a vision-language-action (VLA) policy: choosing among eight actions the VLA proposes raises its success from 65% to 77% across 16 different tasks.
Figures & tables
Figure 1: Planning fails when the target lies beyond the imagined rollout, even with a perfect predictor. (a) A planner using model predictive control (MPC) imagines five control steps ahead (filled squares) and picks the actions whose imagined states come closest to the target (star), five or twenty steps ahead. (b) Success on sixteen Meta-World tasks with the learned world model (V-JEPA 2.1) and with the real simulator as a perfect predictor.
Figure 2: With a five-step rollout, the world model works best for targets five steps ahead. (a) V-JEPA 2.1 as the target moves further ahead. Left axis (higher is better): how often the model identifies the executed action sequence among fifteen random ones, imagining five steps (blue) or up to the target (grey). Right axis (lower is better): expert percentile with a five-step rollout, the share of 100 random sequences scored closer to the target than the expert’s (red); the dashed line marks the threshold of 10. (b) For each task, the plannable range with a five-step rollout (red) and the control steps the expert needs to finish the task (black). Triangles mark ranges that reach the largest tested lookahead.
Figure 3: Only a longer rollout extends the plannable range beyond ten steps. Expert percentile on the four diagnostic tasks; lower is better, chance is 50 and the threshold is 10. (a) Number of recursive training steps. (b) Rollout length; solid and dashed lines are V-JEPA 2.1 predictors trained with five and twenty recursive steps. (c) Frozen encoder; circles mark each encoder’s plannable range.
Figure 4: Keeping the target within the imagined rollout improves planning success. (a) V-JEPA 2.1 on the sixteen main tasks, planning towards a goal image from the same episode with a five-step (red) or thirty-step (blue) rollout, in pure imagination or with MPC. The open square uses the same computation as the five-step planner; the filled thirty-step points use six times more. The dashed line and grey band show MPC with a five-step rollout and subgoals five steps ahead. (b) Three encoders with MPC and a five-step rollout, planning towards a goal image from another reset or towards subgoals twenty or five steps ahead. Subgoals come from an expert rollout of the same episode.
Figure 5: VLA–WM control. The VLA proposes N candidate action sequences from the current observation and the task instruction. The world model imagines and scores each candidate against the target, and the robot executes the best one and then observes again.
Figure 6: World-model selection improves the VLA on trained tasks and near variants. (a) Success on the main and held-out task groups. (b) Learned, untrained and simulator-based critics on a seven-task subset. (c) Candidate-count sweep on the same subset.
Appendix figures & tables25 assets
Supplementary material from the paper’s appendix.
Appendix
Experiment
Reported in
K
Cand. / elites
Feedback
Target
Goal image, 16 tasks
Fig. 4 (a), Tab. 6
5
100 / 10
MPC or none
final frame
30
100 / 10
MPC or none
final frame
30
17 / 2
MPC
final frame
Encoders, subgoals
Figs. 4 , 17
5
100 / 10
MPC
L=5 , 20 or final frame
Far subgoals, 8 tasks
Tab. 5
5
100 / 10
MPC
L=20
20
100 / 10
MPC
L=20
Appendix
Table 1: Planner settings in the closed-loop experiments. K is the rollout horizon in control steps. Every CEM planner uses three iterations and a warm start. “MPC” encodes a new observation at every control step, and “none” encodes only the first. A subgoal lies L control steps ahead on an expert rollout of the same episode. A final frame is the last frame of an expert rollout, from the same episode or from another reset, as stated with each result.
Figure 7: Expert-action percentile in Meta-World simulation and WidowX recordings. Lower is better; dashed and dotted lines mark the threshold and chance. Candidate sampling and control rates differ across datasets, and V-JEPA 2-AC is evaluated outside its post-training distribution.
Encoder
Threshold 5
Threshold 10
Threshold 20
DINOv2
5
10
20
VideoMAEv2
5
5
5
VideoPrism
0
5
5
V-JEPA 2.1
5
10
10
V-JEPA 2
5
5
10
Appendix
Table 2: Plannable range in control steps at three percentile thresholds, computed from the same four-task aggregate curves with a five-step rollout. A value of twenty reaches the largest tested lookahead. At every threshold, the ranges remain shorter than typical task lengths.
Figure 8: Decoded object-position error versus prediction lookahead for V-JEPA 2.1 and two baselines, constant velocity and predict-the-mean. Lower is better.
Encoder
L=1
L=2
L=3
L=5
L=10
L=20
DINOv2
0.043
0.044
0.061
0.173
0.102
0.047
VideoMAEv2
0.130
0.149
0.179
0.259
0.164
0.071
VideoPrism
0.016
0.021
0.027
0.059
0.052
0.039
V-JEPA 2.1
0.071
0.142
0.243
0.401
0.168
0.066
V-JEPA 2
0.043
0.082
0.133
0.230
0.151
0.078
V-JEPA 2.1, fixed window
0.085
0.106
0.201
0.403
0.104
0.034
Appendix
Table 3: Relative spread of the goal-distance score across random candidates, by goal lookahead, at a fixed five-step rollout. The largest value in each row is bold and falls at L=5 for every encoder. The three DINOv2 seeds agree to within 0.02 at every lookahead; seed 0 is shown. The final row is the fixed-window condition on one task, which holds the window population at 36 across the whole curve.
Predictor
K
L=1
L=2
L=3
L=5
L=10
L=20
P∗
Five-step
5
2.27
1.13
0.77
0.51
5.16
32.90
10
10
2.26
1.06
0.92
0.45
0.00
15.71
10
20
2.33
1.00
0.84
0.47
0.00
0.00
20
Twenty-step
5
0.77
0.44
0.05
0.07
3.16
37.07
10
10
0.80
0.42
0.12
0.04
0.00
15.86
10
20
0.75
0.33
0.06
0.04
0.00
0.00
20
Appendix
Table 4: Mean expert percentile by rollout horizon K and goal lookahead L , on four articulated tasks with matched window populations. Lower is better; chance is 50 and the reporting threshold is 10. P∗ is the largest tested L at or below the threshold. The L=20 column at K=20 and the resulting P∗ are bold. Twenty is the largest lookahead in this sweep, so P∗=20 is a lower bound.
Task
K=5
K=20
K=20
100×5×3
25×20×3
100×20×3
1500
1500
6000
assembly
0.00
0.00
0.00
drawer-open
1.00
0.75
1.00
door-open
0.00
0.00
0.05
pick-place
0.00
0.05
0.05
Appendix
Table 5: Closed-loop success with the learned predictor at a five-step and a twenty-step rollout, on eight tasks with twenty episodes each. Budgets are candidates × rollout steps × iterations. The exact-dynamics rows are measured with true simulator transitions and a state-based distance, and are shown for reference only.
Figure 9: Matching the rollout to the goal in closed loop with the learned predictor. (a) Every task at a five-step and a twenty-step rollout with the search held fixed. (b) Both budget controls, with 95% Wilson intervals. The dashed lines are the exact-dynamics reference, shown for comparison only.
Rollout K
Feedback
Budget
Success
Gains/losses
Exact McNemar p
Image of the completed task, same episode
5
MPC
100×5×3
0.3031
—
—
30
MPC
17×30×3
0.4656
57/5
3.1×10−12
30
MPC
100×30×3
0.5156
70/2
1.1×10−18
5
none
100×5×3
0.2250
—
—
30
none
100×30×3
0.3594
58/15
4.1×10−7
Appendix
Table 6: Closed-loop success on the sixteen main tasks with the V-JEPA 2.1 predictor trained with five recursive steps, twenty paired episodes per task, and no subgoals. MPC observes the environment after every control step; “none” plans from imagined states only. Budgets are candidates × rollout steps × iterations. Gains and losses are paired episodes against the five-step row of the same setting.
(a) Expert percentile against perturbed-expert alternatives
Matched rollout
σ=0.1
σ=0.25
σ=0.5
K=20 (4 tasks)
25.00
12.25
2.46
K=30 (3 tasks)
21.64
6.54
0.64
(b) The five encoders at K=20 , σ=0.25 , on three tasks
Encoder
Perturbed-expert
Random Gaussian
V-JEPA 2.1
13.02
0.00
Appendix
Table 7: Expert percentile against perturbed-expert alternatives at a matched rollout. Lower is better. Panel (a) varies the perturbation σ , where smaller is harder, and the two rows are computed on different task sets. Panel (b) compares the five encoders at K=20 and σ=0.25 beside the same encoders measured against random Gaussian alternatives, which is a different quantity reported for context only.
Figure 10: Encoder differences return when the alternatives are hard. Lower is better in every panel. (a) Expert percentile against perturbed-expert alternatives by perturbation size, at two matched rollouts measured on different task sets. (b) The five encoders against hard alternatives. (c) The same five encoders where the random-alternative measurement saturates. Panels (b) and (c) use separate axes.
Protocol
L=1
L=2
L=3
L=5
L=10
L=15
L=20
Unrolled to the goal
0.506
0.600
0.682
0.812
0.965
0.978
0.982
Five-step rollout
0.140
0.211
0.365
0.812
0.650
0.494
0.342
Appendix
Table 8: Recorded-action identification for V-JEPA 2.1 under both protocols. The first protocol unrolls the model to the goal; the second holds the rollout at five steps and moves the goal. The two agree exactly at a goal five steps ahead, where they are the same computation.
Figure 11: Three measurements under one protocol, each best where the rollout ends. (a) Candidate score separation, where higher is better. (b) Expert action ranking, where lower is better. (c) Recorded-action identification for V-JEPA 2.1 under both protocols, where higher is better. The vertical line marks the five-step rollout.
Figure 12: Plannable range with a five-step rollout under changes to (a) predictor capacity, including the 3.75 M-parameter baseline, and (b) number of recursive training steps. Higher is better; both panels use four-task aggregate curves.
Figure 13: Target distance along observed demonstrations and expert-action ranking through learned rollouts. The two measurements use observed and predicted states, respectively.
Figure 14: Closed-loop success grouped by offline plannable range. Circles denote encoder–task pairs, diamonds denote eight additional tasks, and horizontal marks show group means.
Comparison
Difference
95% interval
VLA to VLA–WM (16 tasks)
0.122
[0.038, 0.219]
WM to simulator critic (7 tasks)
0.057
[-0.007, 0.121]
VLA to VLA–WM (4 near tasks)
0.237
[0.075, 0.412]
Appendix
Table 9: Paired success differences and 95% task-bootstrap intervals. Positive values favour the second named system.
Target lookahead
Training
Held out
All sixteen
5 control steps
0.973
0.700
0.922
20 control steps
0.477
0.100
0.406
Appendix
Table 10: Closed-loop success with simulator dynamics and a fixed five-step rollout. Only the target lookahead changes. The comparison removes learned prediction error but retains the mismatch between rollout length and target distance in the second row.
Figure 15: Success versus control steps without new observations on four near variants and four new tasks. The world model replans from predicted states, whereas the VLA executes its predicted chunk. The vertical line marks the aggregate offline range.
Figure 16: Pure imagination, MPC, and MPC with subgoals five steps ahead over twelve encoder–task conditions (three encoders, four diagnostic tasks, five-step rollout). Error bars show 95% Wilson intervals.
Figure 17: Standalone VLA and world-model success on thirteen training tasks, four near variants, and four tasks with new objects or motion patterns. The world-model planner receives demonstration subgoals taken from an expert rollout of the same episode; error bars show 95% Wilson intervals.
Method
Near variants
New tasks
VLA
25.0%
2.5%
VLA + world-model selection
48.7%
6.2%
Standalone world-model planner
98.8%
10.0%
Appendix
Table 11: Success on held-out tasks. The VLA–WM and standalone world-model methods use demonstration subgoals from expert rollouts of the same episodes. Their candidate-generation procedures differ, so the table compares observed capabilities rather than providing a controlled ranking.
Figure 18: Additional VLA–WM evaluations. (a) Sampling-temperature sweep with a matched VLA baseline on sixteen tasks. (b) Learned and simulator-based critics on near variants and new tasks. Error bars show 95% Wilson intervals; annotations use paired episode outcomes.
Task
Split
VLA
VLA–WM
Simulator
assembly
train
5/20
15/20
14/20
button press topdown
train
20/20
20/20
20/20
coffee button
train
20/20
20/20
20/20
dial turn
train
10/20
19/20
20/20
door close
train
20/20
20/20
20/20
door open
train
19/20
20/20
20/20
Appendix
Table 12: Main evaluation counts at eight candidates for selection conditions. The VLA baseline executes one policy sample per environment step (Appendix N ).
Figure 19: Per-task success of VLA and VLA–WM across the sixteen main tasks. Each task uses twenty episodes; horizontal error bars show 95% Wilson intervals.