World models such as DINO-WM and LeWM specify the goal with an image, which is difficult to obtain in advance for novel tasks. We present the Grounded World Model (GWM), a latent world model that enables zero-shot planning in the real world from language goals alone. Given a candidate action sequence and the current observation, GWM predicts the future in the visual space of a pretrained video-language embedding model. The frozen readout of this embedding model maps this imagined future and the task description into the same embedding space, where their negative cosine similarity serves as the planning cost. Training GWM requires only offline and task-agnostic video-action pairs and no language labels. In simulated experiments on WISER, planning with GWM, which executes the candidate action of lowest cost, solves 87% of 288 tasks with unseen instructions and visual signals, while ten fine-tuned VLAs average 22%. We then scale GWM up with real robot data, and use it for zero-shot planning in realistic simulation and real scenes. In the IsaacSim evaluation, planning with GWM completes all 70 trials across 14 tasks that require reasoning over referring expressions, matching a modular planner grounded by a frontier VLM, while pi0.5 reaches 37/70. Deployed on a real Franka, the same stack completes 55/60 separately evaluated pick-and-place sub-tasks, comparable to the modular planner's 52/60, with the full system running locally on a single consumer GPU. Project website: https://quanyili.github.io/gwm-wiser/.
Figures & tables
Figure 1: World Model planning with image goals and with language goals. Left : the image interface, where candidate futures are imagined and compared to the goal in a pure visual latent space. Right : planning with GWM, where the same imagination step runs inside the visual latent space of a video-language embedding model and a fixed readout compares the imagined future to the task description.
Figure 2: Training GWM and reading out its predictions for planning, shown for one candidate of a bowl pick-up task. Blue flow : the candidate trajectory at:t+c is rendered by R into robot-only frames and encoded together with the current observation into et by the frozen vision encoder E ; the learned predictor Pθ outputs the imagined future latent pt , which the frozen readout B projects to zt for comparison with the instruction embedding zg , where 1−cos ranks candidates as Eq. 2 does. Pink flow : during training the ground-truth future clip is encoded by the same E into the target eˉt of Eq. 1 .
Figure 3: The WISER testbed. Each observation combines the instruction ℓ , joint positions, gripper state, and camera input ot . Each of the 24 world-knowledge categories has a training scene and a paired test scene. The two splits share the same 12 motions but have disjoint images, referring expressions, and cube colors. Cube order is randomized across categories, so even a fixed position, such as the second cube from the left, carries different visual and semantic cues across the splits.
Method
Train Success
Test Success
GWM (Ours)
0.92
0.87
GWM w/ action encoder
0.74
0.24
GWM w/ pooled target
0.92
0.61
GWM w/ WEMM
0.92
0.76
GWM w/ xArm6
0.87
0.83
Oracle
0.90
0.93
Table 1: WISER success of GWM and its variants.
Figure 4: WISER success of planning with GWM and of ten fine-tuned VLAs on training tasks (yellow) and semantically disjoint test tasks (green). Dashed lines show the VLA averages.
Figure 5: GWM-based scoring landscapes after scaling GWM to real-world data. Each row begins with an external scoring view. Top : scores over end-effector hover positions for four pointing instructions. The orange path follows the CEM mean, and the star marks the selected final gripper pose. Darker blue denotes a higher value of the probe’s scoring objective. Bottom : scores over slide endpoints for three pushing instructions, followed by the named cube’s final positions under each instruction. The outcome legend distinguishes “moved” (at least 1 cm) from “pushed” (at least 3 cm).
Figure 6: Zero-shot planning in DROID-sim. Left : The scoring view, containing a banana, a cube, a bowl, and two colored bins. Middle : 16 language-blind grasp candidates, visualized as gripper outlines from the wrist camera at the home pose. Right : 16 placement candidates, shown as stars after the block has been grasped. Bottom : The 14 instructions and the target.
Method
Success
Grounding failures
Planning failures
GWM (Ours)
70/70
0
0
Opposite View
65/70
5
0
TiPToP
66/70
1
3
π0.5
37/70
-
V-JEPA 2-AC
25/70
45
0
Table 2: Success over 14 tasks and five trials per task, with failures attributed where the method exposes the distinction. π0.5 is scored under a relaxed pick criterion (Appendix D ).
Figure 9
Method
Scene 1
Scene 2
Total
Ground. fail.
Plan. fail.
GWM (Ours)
36/40
19/20
55/60
2
3
TiPToP
37/40
15/20
52/60
0
8
Table 3: Real-robot success over 60 separately scored pick-and-place sub-task trials per method.
Figure 7: The two real-robot scenes from the external scoring camera, each captured while executing the pick instruction in its title. The figure also lists the twelve instructions and their object or container referents. Scene 1 was recorded under artificial light in the evening, and scene 2 under daylight.
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: GWM (left) and its raw-action-conditioned variant (right). With RAT, actions are rendered into robot-only frames and read by the same frozen vision encoder as the observation. The variant instead encodes numeric actions with a learnable module. Captured images are exemplary.
Hyperparameter
Value
Architecture
Hidden dimension ( dmodel )
4096
FFN intermediate dimension ( dffn )
8192
Attention head dimension ( dhead )
128
Number of layers
5
Number of attention heads
32
Appendix
Table 4: Transformer configuration and training hyperparameters of GWM.
Figure 9: Held-out MSE and cosine similarity between predicted and ground-truth future latents over the real-world training run, evaluated every 2k steps on 256 held-out windows at time scales 1.0, 0.5, and 1.5. The gray curve is the training-batch value smoothed over 200 steps.
Figure 10: Oracle sweeps of the replanning interval, prediction horizon, and keyframe count on the WISER training and test splits, and the Oracle in the 8B and 2B embedding spaces. The training split selects a replanning interval of 20, c=60 , and K=6 , and the test split selects the same optimum in each sweep.
Prompt
Training Set
Test Set
Grasp
Reach
Success
Grasp
Reach
Success
GWM ℓpick + ℓplace
0.97
0.95
0.92
0.99
0.88
0.87
GWM ℓ
0.93
0.88
0.82
0.92
0.76
0.73
InstructVLA ℓpick + ℓplace
0.98
0.53
0.52
0.80
0.30
0.30
InstructVLA ℓ
0.98
0.92
0.89
0.79
0.51
0.47
Appendix
Table 5: Prompt decomposition for GWM (Ours) and InstructVLA on WISER.
Embedding space
Grasp
Reach
Success
Qwen3-VL-Embedding-8B
1.00
0.93
0.93
WEMM-Embedding-9B [ Zhou et al., 2026 ]
0.94
0.90
0.88
Qwen3-VL-Embedding-2B
0.67
0.40
0.33
Perception Encoder [ Bolya et al., 2025 ]
-
-
0.00
Appendix
Table 6: The same Oracle planner built in different embedding spaces (WISER test split).
Figure 18
Model
c
Replan Interval
Dataloader
Implementation
Action
Motus
48
40
LeRobot
Official
Absolute
XVLA
40
20
LeRobot
LeRobot
Absolute
GR00T-N1.6
40
40
LeRobot
Official
Relative
InstructVLA
16
16
RLDS
Official
Absolute
SmolVLA
20
20
LeRobot
LeRobot
Absolute
Wall-OSS
20
20
LeRobot
LeRobot
Absolute
Appendix
Table 7: Configuration of the evaluated VLAs.
Method
H100 GPU Hours
Training Set
Test Set
Grasp
Reach
Success
Grasp
Reach
Success
GWM (Ours)
20
0.97
0.95
0.92
0.99
0.88
0.87
Fine-tuned VLAs
InstructVLA [ Yang et al., 2025 ]
70
0.98
0.92
0.89
0.79
0.51
0.47
SmolVLA [ Shukor et al., 2025 ]
75
0.99
1.00
0.99
0.29
0.31
0.08
Wall-OSS [ Zhai et al., 2025 ]
80
1.00
1.00
1.00
0.68
0.50
0.40
Appendix
Table 8: Full WISER results for the ten fine-tuned VLAs of Fig. 4 and for the GWM variants and diagnostics, with Grasp and Reach alongside Success. Trainable models use the full demonstration corpus except the reduced-data variant, whose category coverage and demonstration count are specified in Appendix B.7 . The GWM rows use N=12 retrieved candidates and the same cosine comparison. They use the same frozen embedding model except for the WEMM variant. GWM (Ours) is repeated from Table 1 for reference.
Task (referring expression)
Target
GWM cam2
GWM cam1
GWM fusion
TiPToP
π0.5
V-JEPA 2-AC
“pick up the fruit”
banana
5/5
5/5
5/5
4/5
0/5
0/5
“…the yellow object”
banana
5/5
0/5
5/5
5/5
5/5
0/5
“…the thing you could eat”
banana
5/5
5/5
5/5
5/5
1/5
0/5
“…neither a toy nor a container”
banana
5/5
5/5
5/5
5/5
2/5
0/5
“…the puzzle toy”
cube
5/5
5/5
5/5
4/5
5/5
0/5
“…the most colorful object”
cube
5/5
5/5
5/5
4/5
5/5
0/5
Appendix
Table 9: Per-task results in the IsaacSim evaluation (successes out of 5 trials). Cam2 is the shipped scoring view used in § 6 . Cam1 is the symmetric opposite view. Fusion averages the two views’ scores per candidate (Appendix F ). The π0.5 column uses the relaxed pick criterion of this appendix.
Scene
Task
GWM best view
GWM shoulder view
TiPToP
1
pick the cylinder
5/5
5/5
5/5
1
place it inside the box
4/5
5/5
5/5
1
pick the object that can eat
5/5
5/5
5/5
1
place it inside the cup
4/5
3/5
5/5
1
pick the yellow object
5/5
4/5
5/5
1
place it in the container in the sky color
5/5
3/5
3/5
Appendix
Table 10: Per-task real-robot results. Each row reports successes out of 5 trials of one pick or place sub-task; totals count the two stages separately.
Figure 11: Two scoring views of hardware scene 1, shown at the start of separate “pick up the cylinder” trials. Success counts cover all eight scene-1 pick and place instructions, with five trials each. Each evaluation uses one camera.
Task
cam2
cam1
fusion
fruit
+0.104
+0.014
+0.057
yellow
+0.105
+0.002 ✗
+0.053
eat
+0.091
+0.018
+0.056
negation
+0.094
+0.019
+0.065
puzzle
+0.052
+0.076
+0.069
colorful
+0.054
+0.057
+0.062
Appendix
Table 11: Object-selection margins per scoring view. The margin is the aggregate cosine of the best object minus the second best, and ✗ marks a wrong selection.
Building generalizable agents for diverse applications remains a fundamental challenge. While imitation learning-based policies succeed in specific training environments, they often fail to generalize to novel scenes and tasks. In this work, we propose World Action Planner, a robot planning system that leverages the reasoning capabilities of Vision-Language Models (VLMs) and the physical grounding of a multi-task pose-image conditioned world model. Our system enables an agent to propose initial action plans and iteratively refine them via optimization and search, reasoning over imagined world model rollouts. We demonstrate that our approach achieves superior performance across compositional tasks, new layouts, and zero-shot generalization scenarios, significantly outperforming state-of-the-art end-to-end policy models such as VLAs and WAMs. Project website at worldactionplanner.github.io
World models offer a promising way to help robots understand how the physical world evolves and plan complex behaviours through imagination. Yet existing studies mainly demonstrate what these models can accomplish, leaving unclear when their predictions remain useful for planning and where they fail. We study this question using action-conditioned predictors built on five frozen self-supervised visual backbones: V-JEPA 2, V-JEPA 2.1, VideoMAEv2, VideoPrism, and DINOv2. We use frozen backbones to test representations intended to transfer across environments. We evaluate these models on diverse Meta-World manipulation tasks and real-robot interactions from BridgeData V2. We find that a world model guides action selection reliably only when the goal lies within, or slightly beyond, the trajectory it imagines during planning. With five-step rollouts, the length the predictor was trained on, the world model ranks actions reliably only for targets five to ten control steps ahead, whereas task goals lie 16 to 53 steps away. Neither an 81-fold larger predictor nor longer-rollout training extends this range; the encoder affects both range and closed-loop success, with V-JEPA 2.1 performing most consistently. More fundamentally, the limit persists under perfect prediction: using the real simulator, success falls from 92% to 41% as the target moves from five to twenty steps ahead of a five-step rollout. Planning therefore requires either longer imagined trajectories or closer subgoals. For distant goals, pure imagination succeeds in 23% of episodes, planning with feedback (MPC) raises success to 30%, imagining as far as the goal to 47%, and nearby expert subgoals to 76%. Used within its plannable range, a world model can also improve a vision-language-action (VLA) policy: choosing among eight actions the VLA proposes raises its success from 65% to 77% across 16 different tasks.
Planning with a learned latent world model is a promising route to control from raw pixels, but a strong world model alone is not enough. We show this experimentally: even with a perfect world model (operationalized by replacing the learned forward predictor with an idealized rollout of the true environment dynamics), a finite-budget sample-based planner still fails on some tasks, indicating that the bottleneck can lie in search rather than in world-model accuracy. Motivated by this gap, we propose IMWM (Intuition Model + World Model), which pairs the world model with an intuition model trained from demonstrations to recognize promising actions. The two models collaborate through three lightweight components: (i) Retrieval Initialization, which initializes the planner's action proposal from a retrieved demonstration; (ii) Hybrid Cost, which combines the intuition score with the world-model rollout cost; and (iii) a Reliability Gate, which adjusts how much the planner trusts intuition in each setting. Across four pixel-based goal-reaching tasks (Two-Room, Reacher, Push-T, and OGBench-Cube), IMWM has higher mean success than the world-model-only planner on all four, with the largest gains on Two-Room (99.2%, +11.5 percentage points) and OGBench-Cube (94.7%, +28.5 percentage points).
Baoqi Gao, Ruize Han, Miao Wang +1
Beihang University · Shenzhen University of Advanced Technology