Latent world models plan by rolling a frozen predictor forward under candidate action sequences and ranking the candidates by the latent distance between their imagined end state and the goal. However, this ranking breaks down when the goal lies several plans away, because the latent distance measures how closely an end state resembles the goal rather than how far it remains from reaching it. To address this, we propose TEMPO, a temporal-distance planning objective that leaves the world model untouched, learns only from the demonstrations already used to train it, and adds negligible cost to the planner's search. TEMPO learns a small map of the frozen latent in which the distance between two states of an episode reflects the number of environment steps between them, and blends this distance into the planner's cost. It requires no rewards, policies or success labels and, being a cost rather than a model, applies to frozen world models with one latent vector per state that plan by a latent distance. We evaluate TEMPO on eleven simulated environments (e.g., maze navigation, tabletop pushing, robotic arm control and three-dimensional manipulation) with the LeWM and PLDM planners. With a small MLP that adds at most 0.3% to a plan's arithmetic, TEMPO improves both planners at every goal distance, including the one-plan setting of their evaluations, raises LeWM from 36% to 99% on TwoRoom three plans from the goal, and remains competitive on a broad range of 2D and 3D navigation, reaching and manipulation tasks.
Figures & tables
Figure 1: How long, not how close. Left: three plans from the goal, the latent distance prefers the plan whose end state looks most like the goal frame, toward the wall, whereas a distance that counts steps prefers the plan through the door. Right: TwoRoom success as the goal moves from one to three plans away.
Figure 2: TEMPO in the planner. LeWM’s frozen encoder E and predictor WM (blue) are unchanged, and nothing is decoded. The planner rolls WM out under each CEM candidate plan and executes the one that minimises J . TEMPO adds the map Fϕ (orange) and scores each candidate plan by ∥Fϕ(z^t+25)−Fϕ(G)∥2 , the steps its end state still is from the goal, added to LeWM’s cost. Top right: Fϕ places the frames of one episode one unit apart per 25 steps.
Planner
TwoRoom
Reacher
Push-T
Cube
Scene
Avg.
Δ
LeWM
one plan
87.7 ± 0.8
94.7 ± 0.7
93.6 ± 1.0
74.8 ± 2.4
69.6 ± 1.1
84.1
–
+ TEMPO (ours)
97.2 ± 0.6
95.5 ± 0.7
95.1 ± 0.6
77.1 ± 1.7
71.1 ± 1.1
87.2
+3.1
two plans
58.5 ± 1.7
95.7 ± 0.6
49.7 ± 1.3
57.9 ± 2.4
56.4 ± 1.2
63.7
–
+ TEMPO (ours)
99.2 ± 0.3
96.8 ± 0.5
55.7 ± 1.2
58.7 ± 2.0
57.9 ± 1.3
73.7
+10.0
three plans
36.3 ± 2.3
89.7 ± 0.4
22.0 ± 0.6
68.7 ± 1.6
48.7 ± 1.3
53.1
–
Table 1: Success rates (%) on five pixel-based goal-conditioned control tasks, one to three plans from the goal. Bold marks the better of each pair. Δ is the gain of TEMPO over the same planner, averaged over environments. Every world model plans from pixels alone.
Planner
Wall †
TwoRoom cross-room
PointMaze M
PointMaze L
Ant-U
Avg.
Δ
LeWM
one plan
67.6 ± 1.4
12.0 ± 1.0
32.1 ± 1.4
28.0 ± 2.0
19.7 ± 1.0
31.9
–
+ TEMPO (ours)
75.7 ± 2.2
46.5 ± 1.2
75.7 ± 1.7
54.8 ± 1.6
21.9 ± 1.1
54.9
+23.0
two plans
57.7 ± 2.4
15.5 ± 1.0
17.6 ± 1.0
2.9 ± 0.6
25.5 ± 1.5
23.8
–
+ TEMPO (ours)
65.9 ± 2.2
69.6 ± 1.8
41.1 ± 2.0
21.3 ± 1.2
29.6 ± 1.0
45.5
+21.7
three plans
–
17.1 ± 1.1
10.8 ± 0.8
0.3 ± 0.2
28.1 ± 1.6
14.1
–
Table 2: Success rates (%) on more complex goal-conditioned control tasks, one to three plans from the goal. Reporting, bold and Δ as in Table 1 . † Wall episodes end at 50 steps, so its two-plan cell uses goal 40 and it has no three-plan cell.
WallRandom
PushObj
TwoRoom
Planner
unseen 1 plan
unseen cross-wall
unseen 1 plan
unseen 3 plans
door moved 1 plan
door moved 3 plans
LeWM ℓ planner
67.5 ± 1.7
15.1 ± 1.9
53.6 ± 2.2
18.8 ± 1.4
65.3 ± 1.9
9.9 ± 1.0
+ TEMPO (ours)
76.9 ± 1.4
22.9 ± 2.5
54.8 ± 2.0
20.5 ± 1.4
74.5 ± 1.6
30.1 ± 0.8
PLDM planner
75.6 ± 2.0
6.0 ± 1.0
44.1 ± 1.6
11.1 ± 0.8
73.9 ± 1.7
18.9 ± 1.0
+ TEMPO (ours)
77.9 ± 2.0
6.5 ± 1.1
48.8 ± 1.7
15.3 ± 1.0
76.3 ± 1.5
30.4 ± 1.0
Table 3: Generalisation to unseen configurations: a layout shift (WallRandom), a shape shift (PushObj) and a moved door (TwoRoom). World models and TEMPO are trained on the training layouts, block shapes or door position; in-distribution results are in Tables 1 and 2 . Cross-wall: the test of Zhou et al. [37] , with start and goal on opposite sides of an unseen wall. ℓ : LeWM trained with single-layout batches. Bold marks the better of each pair.
Figure 3: Training and test setups for the three generalisation families of Table 3 . Unseen configurations are outlined in orange ( ⋆ ). WallRandom holds out wall and door positions outside every training range, PushObj holds out the L, I and + shapes, and TwoRoom moves the door to the opposite end of the wall.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: The eleven environments and two RoboCasa kitchen tasks , left to right, top to bottom: TwoRoom, Reacher, Push-T, Cube, Wall, Scene, PushObj, PointMaze medium and large, Ant-U-Maze, WallRandom, and the RoboCasa kitchen tasks Pick-and-Place and Kitchen-Nav. Each panel is one 224×224 training frame, the only observation the world models and the planner receive.
Environment
Episodes (trained on)
Mean length
Action dim.
TwoRoom
10,000
92
2
Push-T
18,685
125
2
Reacher
10,000
200
2
Cube
10,000
200
5
Scene
10,000
200
5
PushObj
20,000
101
2
Appendix
Table 4: Training sets. Episodes the world models were trained on, mean episode length in environment steps, and action dimension. Where a 90:10 episode split defines the protocol (Section 4.1 ) the count is the 90% training split.
Ablation
TwoRoom
Push-T
Reacher
Cube
Scene
Avg.
LeWM
One plan from the goal
Fϕ alone ( w=1 )
98.3 ± 0.3
89.1 ± 0.9
93.5 ± 0.6
75.9 ± 2.1
68.5 ± 1.5
85.0
Latent +Fϕ (ours, w=0.5 )
97.2 ± 0.6
95.1 ± 0.6
95.5 ± 0.7
77.1 ± 1.7
71.1 ± 1.1
87.2
Two plans
Fϕ alone ( w=1 )
97.2 ± 0.7
30.4 ± 1.9
96.5 ± 0.5
52.3 ± 2.1
52.5 ± 1.5
65.8
Appendix
Table 5: Ablation of the planning cost on each frozen world model. The world model and planner are fixed and only the cost changes: latent +Fϕ is the cost of Eq. ( 5 ) used in every other table, and Fϕ alone drops the latent term ( w=1 ). Bold marks the better of each pair. Avg.: mean over the five environments.
Condition
TwoRoom
Push-T
Reacher
Cube
Scene
Avg.
in-distribution — LeWM planner
36.5 ± 2.1
21.1 ± 1.3
89.2 ± 0.5
69.5 ± 1.3
47.7 ± 1.0
52.8
in-distribution — + TEMPO (ours)
98.4 ± 0.3
26.1 ± 1.4
97.7 ± 0.4
70.0 ± 1.0
49.5 ± 1.7
68.3
blur — LeWM planner
25.9 ± 2.2
20.7 ± 1.1
88.5 ± 0.5
68.0 ± 1.6
42.0 ± 1.4
49.0
blur — + TEMPO (ours)
38.1 ± 2.4
26.5 ± 1.7
96.0 ± 0.6
68.9 ± 1.5
40.5 ± 1.5
54.0
colour jitter — LeWM planner
34.1 ± 3.3
6.3 ± 1.3
24.5 ± 5.2
59.2 ± 4.3
41.1 ± 2.1
33.0
colour jitter — + TEMPO (ours)
74.3 ± 8.0
12.8 ± 3.2
23.2 ± 6.5
61.5 ± 4.5
42.1 ± 2.2
42.8
Appendix
Table 6: Same out-of-distribution protocol on every environment (goal 75 steps ahead). The pixel corruptions are generic wrappers on the rendered frame, applied with identical parameters in every environment (values as in 35 ); the reference row runs the same protocol with no corruption, so the first frame is re-rendered live in every row. World model frozen, Fϕ trained online in the corrupted environment, blend cost of Eq. ( 5 ). Bold: the better of each pair.
Condition
Wall †
TwoRoom cross-room
PointMaze M
PointMaze L
Ant-U
Avg.
in-distribution — LeWM planner
57.7 ± 2.4
17.1 ± 1.1
10.8 ± 0.8
0.3 ± 0.2
28.1 ± 1.6
22.8
in-distribution — + TEMPO (ours)
65.9 ± 2.2
74.5 ± 1.7
31.7 ± 1.0
6.5 ± 1.0
34.0 ± 2.1
42.5
blur — LeWM planner
52.1 ± 2.8
8.9 ± 1.4
10.4 ± 0.9
0.1 ± 0.1
12.5 ± 1.4
16.8
blur — + TEMPO (ours)
60.9 ± 2.4
23.3 ± 2.5
17.9 ± 1.7
4.1 ± 1.0
12.4 ± 1.5
23.7
colour jitter — LeWM planner
56.4 ± 2.6
16.1 ± 1.5
9.6 ± 1.0
1.6 ± 0.6
17.7 ± 1.7
20.3
colour jitter — + TEMPO (ours)
68.0 ± 2.4
59.7 ± 6.9
10.4 ± 1.3
1.7 ± 0.7
19.7 ± 1.3
31.9
Appendix
Table 7: The same pixel-corruption protocol on the more complex environments of Table 2 (goal 75 steps ahead). The pixel corruptions are generic wrappers on the rendered frame, applied with identical parameters in every environment (values as in 35 ); the in-distribution rows repeat the corresponding cells of Table 2 . † Wall episodes end at 50 steps, so its goal is 40 steps ahead; cross-room uses its 150-step budget. World model frozen, Fϕ trained online in the corrupted environment, blend cost of Eq. ( 5 ). Bold: the better of each pair.
Figure 5: Test-time pixel corruptions: cumulative success versus MPC step, goal 75 steps ahead (30 steps = budget 150). Rows: condition; columns: environment. Dashed blue: the LeWM planner with its own cost; orange: the same planner + TEMPO (ours), on identical windows.
Environment
Shift
Planner
+ TEMPO
Push-T
none (in distribution)
21.1
26.1
larger agent
18.3
26.3
TwoRoom
none (in distribution)
36.5
98.4
door moved
9.9
30.3
larger agent
36.4
94.9
Reacher
none (in distribution)
89.2
97.7
Appendix
Table 8: Factor shifts on Push-T, TwoRoom, Reacher, Cube and Scene, goal 75 steps ahead. Protocol of Wang et al. [35] : the first frame is rendered live under the shift, the goal frame stays the clean dataset frame, the world model is frozen and Fϕ is trained online in the shifted environment. Shifts change the geometry, the dynamics or the camera. Success (%), LeWM planner and the same planner with TEMPO ; mean over three seeds; bold: the better of the pair.
Planner
Pick-and-Place (SR ↑ )
Kitchen-Nav (SR ↑ )
LeWM
one plan
0.32
0.16
+ TEMPO (ours)
0.36
0.22
two plans
0.20
0.19
+ TEMPO (ours)
0.21
0.27
three plans
0.14
0.23
Appendix
Table 9: Planning results for offline world models on two RoboCasa kitchen tasks, one to three plans from the goal. Success rate as a fraction of episodes. Bold marks the better of each pair.
Fϕ -distance
R2 of frame index
Environment
same index, other episode
25 steps, same episode
from Fϕ(z)
from z
TwoRoom
1.48
1.03
0.11
0.13
Push-T
2.27
1.09
0.54
0.73
Reacher
2.27
1.95
0.01
0.03
Cube
1.63
1.07
0.03
0.13
Scene
2.45
1.03
0.21
0.32
Appendix
Table 10: Fϕ does not encode the frame index. Median Fϕ -distance (units of 25 steps) between states with the same frame index in different episodes, against states 25 steps apart in the same episode, and R2 of a linear probe predicting the frame index from Fϕ(z) and from the raw latent z .
Training episodes
Held-out episodes
Environment
ρ
error
ρ
error
PointMaze M
0.83
0.28
0.83
0.28
PointMaze L
0.99
0.06
0.99
0.06
Ant-U
0.56
0.56
0.51
0.59
Appendix
Table 11: Fϕ generalises to held-out episodes. Spearman ρ between the Fϕ -distance and the step gap k , and mean absolute error of the Fϕ -distance against its target k/25 (units of 25 steps), over about 24,000 same-episode pairs with k≤75 (LeWM, mean over three seeds).
benchmark / offset
windows
window success
ρ (score, realised dist)
ρ (score, success)
top-5 success
oracle top-5
PushT / 25
59
0.74
+0.83
+0.32
0.89
1.00
TwoRoom / 25
48
0.80
+0.49
+0.21
0.84
0.91
PushT / 75
60
0.28
+0.66
+0.02
0.29
0.65
Reacher / 75
60
0.95
+0.85
+0.22
0.97
0.99
TwoRoom / 75
60
0.42
+0.66
−0.07
0.41
0.75
Appendix
Table 12: What the latent distance ranks. We diagnose the two tasks where the flat planner collapses three plans from the goal (TwoRoom from 88 to 36% , Push-T from 94 to 22% ); Reacher, where it does not collapse, is the control, and the one-plan rows show the same cost working close to the goal. 50 executed action sequences per window, scored by the planner’s own latent distance to the target; ρ (score, realised distance) and ρ (score, eventual success) over the sequences of a window; top-5 = success of the five best-scored sequences, oracle = best five in hindsight.
ρ (cost, success)
top-5 success
Environment
latent
TEMPO
latent
TEMPO
best five
TwoRoom
-0.08
+0.17
26
35
57
Push-T
+0.06
+0.13
30
30
51
Reacher
+0.25
+0.33
95
98
99
Cube
+0.39
+0.38
32
32
42
Scene
+0.02
+0.04
18
21
60
Appendix
Table 13: Which executed plans each cost ranks first, three plans from the goal. For 20 windows per seed, 50 candidate plans are executed from the same state for 25 steps and then finished by the flat planner; each candidate is scored by LeWM’s latent distance and by TEMPO ’s cost. ρ : Spearman correlation between the (negated) cost and eventual success; top-5: success rate (%) of the five best-scored candidates; best five: the success rate of the five candidates that actually succeed most. Mean over three seeds; bold: the better of the pair.