World models such as DINO-WM and LeWM specify the goal with an image, which is difficult to obtain in advance for novel tasks. We present the Grounded World Model (GWM), a latent world model that enables zero-shot planning in the real world from language goals alone. Given a candidate action sequence and the current observation, GWM predicts the future in the visual space of a pretrained video-language embedding model. The frozen readout of this embedding model maps this imagined future and the task description into the same embedding space, where their negative cosine similarity serves as the planning cost. Training GWM requires only offline and task-agnostic video-action pairs and no language labels. In simulated experiments on WISER, planning with GWM, which executes the candidate action of lowest cost, solves 87% of 288 tasks with unseen instructions and visual signals, while ten fine-tuned VLAs average 22%. We then scale GWM up with real robot data, and use it for zero-shot planning in realistic simulation and real scenes. In the IsaacSim evaluation, planning with GWM completes all 70 trials across 14 tasks that require reasoning over referring expressions, matching a modular planner grounded by a frontier VLM, while pi0.5 reaches 37/70. Deployed on a real Franka, the same stack completes 55/60 separately evaluated pick-and-place sub-tasks, comparable to the modular planner's 52/60, with the full system running locally on a single consumer GPU. Project website: https://quanyili.github.io/gwm-wiser/.
Figures & tables
Figure 1: World Model planning with image goals and with language goals. Left : the image interface, where candidate futures are imagined and compared to the goal in a pure visual latent space. Right : planning with GWM, where the same imagination step runs inside the visual latent space of a video-language embedding model and a fixed readout compares the imagined future to the task description.
Figure 2: Training GWM and reading out its predictions for planning, shown for one candidate of a bowl pick-up task. Blue flow : the candidate trajectory at:t+c is rendered by R into robot-only frames and encoded together with the current observation into et by the frozen vision encoder E ; the learned predictor Pθ outputs the imagined future latent pt , which the frozen readout B projects to zt for comparison with the instruction embedding zg , where 1−cos ranks candidates as Eq. 2 does. Pink flow : during training the ground-truth future clip is encoded by the same E into the target eˉt of Eq. 1 .
Figure 3: The WISER testbed. Each observation combines the instruction ℓ , joint positions, gripper state, and camera input ot . Each of the 24 world-knowledge categories has a training scene and a paired test scene. The two splits share the same 12 motions but have disjoint images, referring expressions, and cube colors. Cube order is randomized across categories, so even a fixed position, such as the second cube from the left, carries different visual and semantic cues across the splits.
Method
Train Success
Test Success
GWM (Ours)
0.92
0.87
GWM w/ action encoder
0.74
0.24
GWM w/ pooled target
0.92
0.61
GWM w/ WEMM
0.92
0.76
GWM w/ xArm6
0.87
0.83
Oracle
0.90
0.93
Table 1: WISER success of GWM and its variants.
Figure 4: WISER success of planning with GWM and of ten fine-tuned VLAs on training tasks (yellow) and semantically disjoint test tasks (green). Dashed lines show the VLA averages.
Figure 5: GWM-based scoring landscapes after scaling GWM to real-world data. Each row begins with an external scoring view. Top : scores over end-effector hover positions for four pointing instructions. The orange path follows the CEM mean, and the star marks the selected final gripper pose. Darker blue denotes a higher value of the probe’s scoring objective. Bottom : scores over slide endpoints for three pushing instructions, followed by the named cube’s final positions under each instruction. The outcome legend distinguishes “moved” (at least 1 cm) from “pushed” (at least 3 cm).
Figure 6: Zero-shot planning in DROID-sim. Left : The scoring view, containing a banana, a cube, a bowl, and two colored bins. Middle : 16 language-blind grasp candidates, visualized as gripper outlines from the wrist camera at the home pose. Right : 16 placement candidates, shown as stars after the block has been grasped. Bottom : The 14 instructions and the target.
Method
Success
Grounding failures
Planning failures
GWM (Ours)
70/70
0
0
Opposite View
65/70
5
0
TiPToP
66/70
1
3
π0.5
37/70
-
V-JEPA 2-AC
25/70
45
0
Table 2: Success over 14 tasks and five trials per task, with failures attributed where the method exposes the distinction. π0.5 is scored under a relaxed pick criterion (Appendix D ).
Figure 9
Method
Scene 1
Scene 2
Total
Ground. fail.
Plan. fail.
GWM (Ours)
36/40
19/20
55/60
2
3
TiPToP
37/40
15/20
52/60
0
8
Table 3: Real-robot success over 60 separately scored pick-and-place sub-task trials per method.
Figure 7: The two real-robot scenes from the external scoring camera, each captured while executing the pick instruction in its title. The figure also lists the twelve instructions and their object or container referents. Scene 1 was recorded under artificial light in the evening, and scene 2 under daylight.
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: GWM (left) and its raw-action-conditioned variant (right). With RAT, actions are rendered into robot-only frames and read by the same frozen vision encoder as the observation. The variant instead encodes numeric actions with a learnable module. Captured images are exemplary.
Hyperparameter
Value
Architecture
Hidden dimension ( dmodel )
4096
FFN intermediate dimension ( dffn )
8192
Attention head dimension ( dhead )
128
Number of layers
5
Number of attention heads
32
Appendix
Table 4: Transformer configuration and training hyperparameters of GWM.
Figure 9: Held-out MSE and cosine similarity between predicted and ground-truth future latents over the real-world training run, evaluated every 2k steps on 256 held-out windows at time scales 1.0, 0.5, and 1.5. The gray curve is the training-batch value smoothed over 200 steps.
Figure 10: Oracle sweeps of the replanning interval, prediction horizon, and keyframe count on the WISER training and test splits, and the Oracle in the 8B and 2B embedding spaces. The training split selects a replanning interval of 20, c=60 , and K=6 , and the test split selects the same optimum in each sweep.
Prompt
Training Set
Test Set
Grasp
Reach
Success
Grasp
Reach
Success
GWM ℓpick + ℓplace
0.97
0.95
0.92
0.99
0.88
0.87
GWM ℓ
0.93
0.88
0.82
0.92
0.76
0.73
InstructVLA ℓpick + ℓplace
0.98
0.53
0.52
0.80
0.30
0.30
InstructVLA ℓ
0.98
0.92
0.89
0.79
0.51
0.47
Appendix
Table 5: Prompt decomposition for GWM (Ours) and InstructVLA on WISER.
Embedding space
Grasp
Reach
Success
Qwen3-VL-Embedding-8B
1.00
0.93
0.93
WEMM-Embedding-9B [ Zhou et al., 2026 ]
0.94
0.90
0.88
Qwen3-VL-Embedding-2B
0.67
0.40
0.33
Perception Encoder [ Bolya et al., 2025 ]
-
-
0.00
Appendix
Table 6: The same Oracle planner built in different embedding spaces (WISER test split).
Figure 18
Model
c
Replan Interval
Dataloader
Implementation
Action
Motus
48
40
LeRobot
Official
Absolute
XVLA
40
20
LeRobot
LeRobot
Absolute
GR00T-N1.6
40
40
LeRobot
Official
Relative
InstructVLA
16
16
RLDS
Official
Absolute
SmolVLA
20
20
LeRobot
LeRobot
Absolute
Wall-OSS
20
20
LeRobot
LeRobot
Absolute
Appendix
Table 7: Configuration of the evaluated VLAs.
Method
H100 GPU Hours
Training Set
Test Set
Grasp
Reach
Success
Grasp
Reach
Success
GWM (Ours)
20
0.97
0.95
0.92
0.99
0.88
0.87
Fine-tuned VLAs
InstructVLA [ Yang et al., 2025 ]
70
0.98
0.92
0.89
0.79
0.51
0.47
SmolVLA [ Shukor et al., 2025 ]
75
0.99
1.00
0.99
0.29
0.31
0.08
Wall-OSS [ Zhai et al., 2025 ]
80
1.00
1.00
1.00
0.68
0.50
0.40
Appendix
Table 8: Full WISER results for the ten fine-tuned VLAs of Fig. 4 and for the GWM variants and diagnostics, with Grasp and Reach alongside Success. Trainable models use the full demonstration corpus except the reduced-data variant, whose category coverage and demonstration count are specified in Appendix B.7 . The GWM rows use N=12 retrieved candidates and the same cosine comparison. They use the same frozen embedding model except for the WEMM variant. GWM (Ours) is repeated from Table 1 for reference.
Task (referring expression)
Target
GWM cam2
GWM cam1
GWM fusion
TiPToP
π0.5
V-JEPA 2-AC
“pick up the fruit”
banana
5/5
5/5
5/5
4/5
0/5
0/5
“…the yellow object”
banana
5/5
0/5
5/5
5/5
5/5
0/5
“…the thing you could eat”
banana
5/5
5/5
5/5
5/5
1/5
0/5
“…neither a toy nor a container”
banana
5/5
5/5
5/5
5/5
2/5
0/5
“…the puzzle toy”
cube
5/5
5/5
5/5
4/5
5/5
0/5
“…the most colorful object”
cube
5/5
5/5
5/5
4/5
5/5
0/5
Appendix
Table 9: Per-task results in the IsaacSim evaluation (successes out of 5 trials). Cam2 is the shipped scoring view used in § 6 . Cam1 is the symmetric opposite view. Fusion averages the two views’ scores per candidate (Appendix F ). The π0.5 column uses the relaxed pick criterion of this appendix.
Scene
Task
GWM best view
GWM shoulder view
TiPToP
1
pick the cylinder
5/5
5/5
5/5
1
place it inside the box
4/5
5/5
5/5
1
pick the object that can eat
5/5
5/5
5/5
1
place it inside the cup
4/5
3/5
5/5
1
pick the yellow object
5/5
4/5
5/5
1
place it in the container in the sky color
5/5
3/5
3/5
Appendix
Table 10: Per-task real-robot results. Each row reports successes out of 5 trials of one pick or place sub-task; totals count the two stages separately.
Figure 11: Two scoring views of hardware scene 1, shown at the start of separate “pick up the cylinder” trials. Success counts cover all eight scene-1 pick and place instructions, with five trials each. Each evaluation uses one camera.
Task
cam2
cam1
fusion
fruit
+0.104
+0.014
+0.057
yellow
+0.105
+0.002 ✗
+0.053
eat
+0.091
+0.018
+0.056
negation
+0.094
+0.019
+0.065
puzzle
+0.052
+0.076
+0.069
colorful
+0.054
+0.057
+0.062
Appendix
Table 11: Object-selection margins per scoring view. The margin is the aggregate cosine of the best object minus the second best, and ✗ marks a wrong selection.