Latent world models plan toward goal images with a frozen pretrained predictor, without task rewards or extra trained heads. However, their planners struggle with long-range goals, and prior work addresses this by training extra components such as value functions or subgoal models. We show that the planning target itself can cause this failure: even with exact dynamics and globally optimal short-horizon search, scoring predictions by their distance to the final goal rejects the first steps of a route that initially moves away from the goal. Building on this insight, we propose Anchored Planning (AP), a training-free method that reuses the world model's own offline trajectories. AP retrieves a segment that leads from the current observation toward the goal and aims the frozen planner at an observation shortly after the segment's start. Across four diverse tasks, AP substantially improves frozen LeWM planners for both action synthesis and action ranking, and it outperforms both additional final-goal search and the LeWM planner on long-range goals.
Figures & tables
Figure 1: Changing the target improves control with the same frozen predictor. (a) Both target rules score the same endpoints pi=F(zt,ui) . Final-goal scoring selects u1 , whereas observed-target scoring selects u2 . Solid curves show predictions, and dashed segments show scoring distances. The target inset shows a recorded successor. Positions and observations are schematic. (b) On PushT, intermediate targets sustain control as the recorded goal offset grows. Curves use paired standard-start queries and the same frozen LeWM weights.
No memory
Observation-only
Recorded actions
Task
Start
LeWM
CEM final
CEM learned
AP-CEM
Direct
Rank final
Rank learned
AP-rank
Cube
Standard
15.6
7.0
62.5
46.9
82.0
62.5
91.4
73.4
Perturbed
12.5
2.7
38.7
29.7
47.7
41.8
50.0
45.7
PushT
Standard
7.0
2.3
52.3
69.5
51.6
25.8
54.7
66.4
Perturbed
9.4
2.7
48.8
64.8
18.0
12.9
35.5
47.7
Reacher
Standard
66.4
25.8
54.7
75.8
99.2
90.6
95.3
96.9
Table 1: Intermediate targets improve control with the same frozen model. We report success (%) on 128 paired queries per task. Column groups indicate what experience each controller can use. Learned targets are trained on memory observations. For perturbed starts, we average within each query, and task means weight all tasks equally. Bold marks the highest success within each group of controllers. AP-CEM and AP-rank use observed targets.
Successor error
CEM success (%)
Task
Learned
Observed
Learned
Observed
Cube
19.708
63.056
62.5
46.9
PushT
3.774
10.668
52.3
69.5
Reacher
74.724
154.351
54.7
75.8
TwoRoom
133.857
266.676
53.1
46.9
Table 2: Training-free observed targets achieve higher mean success than a more accurate learned target. We measure the mean squared distance to the recorded successor on held-out memory transitions and CEM success on standard-start evaluation queries. Bold marks higher success. Appendix C.1 explains how target accuracy is measured.
(a) Target offset and memory size
AP-rank
AP-CEM
Variant
Standard
Perturbed
Standard
Perturbed
Target ℓ=1
69.5
61.3
43.0
37.9
Target ℓ=3
78.5
68.5
69.9
62.6
Target ℓ=5 (default)
83.8
72.3
59.8
55.1
Target ℓ=10
81.6
68.4
54.1
52.7
Table 3: Target placement and retrieval timing determine how memory guides control. We average success (%) across the four tasks. Predictions span L=5 actions, and ℓ sets the target’s offset in the recorded trajectory. When varying memory size, we keep the default target offset. In the fixed-span setting, we use h=H throughout execution. Bold marks the highest mean per column within each panel.
Endpoint error
Selection regret
Task
Start
Predictor
Displ.
Predictor
Displ.
Direct
Cube
Standard
7.538
35.344
5.332
7.535
11.375
Perturbed
37.726
66.860
8.472
11.731
14.903
PushT
Standard
1.532
3.078
0.200
0.580
0.849
Perturbed
2.980
9.562
0.753
2.418
4.449
Reacher
Standard
25.243
51.917
8.644
15.058
18.640
Table 4: Live-state prediction improves endpoint accuracy and action selection. We report mean squared latent distances in task-specific units, with lower values indicating better performance. Bold marks each metric’s minimum. Displacement transfers recorded motion, whereas Direct executes the closest record’s action. We exclude starts if any candidate ends before five actions (at most 3 of 128 standard and 6 of 256 perturbed starts per task).
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Cube
PushT
Reacher
TwoRoom
H
150
140
150
100
B
150
280
300
200
Memory
8,000
16,688
8,000
8,000
Appendix
Table 5: Evaluation inputs. H is the recorded goal offset, B the primitive-action allowance, and memory the number of eligible episodes.
Task
Width
Output
Best step
Validation MSE
Cube
512
Absolute
48,500
0.104926
512
Residual
48,500
0.104579
1024
Absolute
41,000
0.104422
1024
Residual
48,500
0.102646
PushT
512
Absolute
48,000
0.023950
512
Residual
48,000
0.024007
Appendix
Table 6: Selecting learned targets on held-out memory. We average squared error over latent coordinates to compute validation MSE. Bold marks the selected configuration.
Task
Controller
Standard
Perturbed
Stalling
Detour
Stalling
Detour
Cube
LeWM
39.8 (43/108)
10.5 (2/19)
48.2 (108/224)
6.7 (2/30)
CEM (final goal)
100.0 (119/119)
37.5 (3/8)
100.0 (249/249)
40.0 (2/5)
CEM (learned)
79.2 (38/48)
60.8 (48/79)
89.2 (140/157)
68.0 (66/97)
AP-CEM
86.8 (59/68)
66.1 (39/59)
92.2 (166/180)
73.0 (54/74)
CEM (transported)
91.3 (73/80)
68.1 (32/47)
95.4 (186/195)
71.2 (42/59)
Appendix
Table 7: Physical behavior at the original goal offsets. Each percentage is followed by the number of positive cases and eligible rollouts. With W=10 , we measure stalling over failed rollouts in which the policy made at least one decision. With δ=0.5 , we measure detours over successful rollouts that entered policy control. For perturbed behavior, we pool eligible rollouts from both prefixes.
Controller
Start
Stalling window
Detour threshold
W=5
10
20
δ=0.25
0.5
1.0
CEM (final goal)
Standard
55.1
53.3
49.9
68.6
30.6
29.1
CEM (final goal)
Perturbed
54.9
52.7
49.7
46.4
32.5
19.0
CEM (learned)
Standard
46.1
39.9
31.4
77.6
70.0
53.7
CEM (learned)
Perturbed
45.6
42.7
32.1
79.3
72.0
58.9
AP-CEM
Standard
51.7
49.5
40.5
90.6
84.1
70.7
Appendix
Table 8: Behavior across thresholds. We average rates (%) equally across the four tasks, computing each rate from that task’s eligible rollout counts.
Task
Iterations
Final goal
Learned target
Observed target
Standard
Perturbed
Standard
Perturbed
Standard
Perturbed
Cube
1
3.9
2.0
11.7
6.6
10.9
7.0
2
3.1
2.3
40.6
26.2
26.6
15.6
5
3.1
2.3
52.3
34.0
36.7
26.6
10
4.7
3.9
59.4
35.2
44.5
29.7
30
7.0
2.7
62.5
38.7
46.9
29.7
Appendix
Table 9: Success across search budgets. We report success (%) on the main queries at the original offsets. Only the number of CEM iterations changes, and each decision evaluates 300I+2 predicted blocks. For perturbed starts, we average the two starts of each query.
Task
Paper report
LeWM evaluator
Our implementation
Cube
74.0
74.7
74.7
PushT
96.0
92.0
91.3
Reacher
86.0
80.0
80.0
TwoRoom
87.0
87.3
87.3
Appendix
Table 10: LeWM under its evaluation protocol. We compare published success (%) from Figure 6 of Maes et al. [2026] with measurements from the LeWM evaluator and our implementation on the same queries. Measurements are averaged over three seeds with 50 queries each, including the seed specified in the LeWM configuration.
Success (%)
Work ( 103 blocks)
Task
Standard
Perturbed
Standard
Perturbed
Cube
13.3
13.3
1,189
1,197
PushT
7.0
8.2
2,402
2,374
Reacher
73.4
77.3
1,388
1,322
TwoRoom
27.3
25.8
1,500
1,509
Appendix
Table 11: LeWM with five-action replanning. The variant keeps the 25-action lookahead and the LeWM CEM settings and changes only execution, which now covers one five-action block. Evaluation uses the main queries and original goal offsets, with 128 standard starts and 256 perturbed starts per task. Work is the mean number of predicted five-action blocks per episode, in thousands as in Table 13 .
Task
H
LeWM
Final goal
Learned
AP-CEM
Std.
Pert.
Std.
Pert.
Std.
Pert.
Std.
Pert.
Cube
25
32.0
33.6
28.9
26.2
75.0
53.9
63.3
49.6
50
16.4
12.1
7.0
4.3
58.6
37.9
47.7
30.5
100
18.8
13.3
8.6
9.0
55.5
34.8
32.0
25.4
150
15.6
12.5
7.0
2.7
62.5
38.7
46.9
29.7
PushT
25
60.9
57.8
51.6
46.9
91.4
75.4
91.4
87.5
Appendix
Table 12: Success across recorded goal offsets. To compare success (%) across offsets, we keep query sources fixed and move the goal to the observation H actions later, preserving each task’s B/H . Prediction lookahead and target-model training remain unchanged. Results are shown for standard starts (Std.) and for the average of each query’s perturbed starts (Pert.).
Controller
Cube
PushT
Reacher
TwoRoom
Std.
Pert.
Std.
Pert.
Std.
Pert.
Std.
Pert.
LeWM
235.20
241.88
514.69
505.37
255.59
244.51
320.98
321.50
CEM (final goal)
252.62
263.17
493.56
491.10
407.76
425.31
354.95
354.95
CEM (learned)
142.27
191.54
375.06
387.40
393.91
379.98
250.65
247.94
AP-CEM
170.26
206.84
326.11
338.63
323.58
313.17
262.61
265.07
CEM (transported)
193.40
223.26
347.84
393.80
306.56
310.50
289.19
291.37
Appendix
Table 13: Mean episode prediction work at the original offsets. Values are thousands of predicted five-action blocks per episode. Columns marked Std. show standard starts, and those marked Pert. show the average of each query’s perturbed starts.
Setting
Cube
PushT
Reacher
TwoRoom
Std.
Pert.
Std.
Pert.
Std.
Pert.
Std.
Pert.
AP-CEM
Default
46.9
29.7
69.5
64.8
75.8
78.9
46.9
46.9
ℓ=1
43.8
31.3
34.4
29.3
50.0
44.9
43.8
46.1
ℓ=3
61.7
42.6
68.8
60.9
91.4
90.6
57.8
56.3
ℓ=10
19.5
15.6
58.6
57.4
87.5
89.8
50.8
48.0
Appendix
Table 14: Target offsets and memory size. On the main queries, we measure success (%) while varying the target offset or memory size. By default, the target is the five-step successor and retrieval uses the full memory. Prediction and execution retain five-action blocks throughout.
Setting
Cube
PushT
Reacher
TwoRoom
Std.
Pert.
Std.
Pert.
Std.
Pert.
Std.
Pert.
Full key
73.4
45.7
66.4
47.7
96.9
95.7
98.4
100.0
Without distant endpoint
81.3
46.5
67.2
50.4
94.5
94.5
100.0
100.0
Without displacement
73.4
45.7
63.3
44.9
96.1
97.7
99.2
100.0
Fixed span h=H
4.7
3.1
0.0
0.0
12.5
19.1
21.1
23.0
Appendix
Table 15: Retrieval keys and temporal span. For AP-rank, we compare success (%) on the main queries. When omitting a key component, we keep the other components unchanged. Fixing h=H replaces the adaptive span max(5,H−t) .
Task
Start
Observed target
Transported target
Cube
Standard
46.9
37.5
Perturbed
29.7
23.8
PushT
Standard
69.5
61.7
Perturbed
64.8
44.5
Reacher
Standard
75.8
81.3
Perturbed
78.9
80.5
Appendix
Table 16: Anchoring and displacement transport on the main queries. Under the same CEM search, we compare success (%) when using the observed endpoint or transported motion. At perturbed starts, we average the two starts of each query, and the mean weights all tasks equally.
First-block selector
Cube
PushT
Reacher
TwoRoom
Std.
Pert.
Std.
Pert.
Std.
Pert.
Std.
Pert.
Direct
82.0
47.7
51.6
18.0
99.2
96.9
99.2
100.0
Predictor / Final
67.2
43.0
39.1
18.4
96.1
97.3
100.0
100.0
Simulator / Final
73.4
44.5
43.0
19.1
96.1
97.3
99.2
99.6
Predictor / Learned
85.9
48.8
45.3
18.8
96.9
97.7
100.0
99.6
Simulator / Learned
86.7
49.2
42.2
22.3
96.9
97.7
100.0
99.6
Appendix
Table 17: Changing only the first action block. After selecting the first block, we use the same Direct continuation to measure success (%). Predictor and simulator selectors score the same eight candidates against the indicated target. All 128 assigned queries per task are included.
Task
CEM: Standard
Ranking: Standard
Ranking: Perturbed
Final
Observed
Final
Observed
Final
Observed
Cube
8.6
39.1
64.8
71.1
41.4
43.4
PushT
1.6
81.3
25.8
78.1
9.0
49.2
Reacher
25.0
75.8
92.2
95.3
88.7
95.7
TwoRoom
0.8
39.1
85.2
100.0
88.3
99.6
Appendix
Table 18: Target effects on independent supporting queries. For each action rule, we compare success (%) with final-goal and observed targets. Ranking is evaluated at standard and perturbed starts, whereas CEM is evaluated at standard starts. Within each query, we average the two assigned perturbed starts.