Latent world models plan by rolling a frozen predictor forward under candidate action sequences and ranking the candidates by the latent distance between their imagined end state and the goal. However, this ranking breaks down when the goal lies several plans away, because the latent distance measures how closely an end state resembles the goal rather than how far it remains from reaching it. To address this, we propose TEMPO, a temporal-distance planning objective that leaves the world model untouched, learns only from the demonstrations already used to train it, and adds negligible cost to the planner's search. TEMPO learns a small map of the frozen latent in which the distance between two states of an episode reflects the number of environment steps between them, and blends this distance into the planner's cost. It requires no rewards, policies or success labels and, being a cost rather than a model, applies to frozen world models with one latent vector per state that plan by a latent distance. We evaluate TEMPO on eleven simulated environments (e.g., maze navigation, tabletop pushing, robotic arm control and three-dimensional manipulation) with the LeWM and PLDM planners. With a small MLP that adds at most 0.3% to a plan's arithmetic, TEMPO improves both planners at every goal distance, including the one-plan setting of their evaluations, raises LeWM from 36% to 99% on TwoRoom three plans from the goal, and remains competitive on a broad range of 2D and 3D navigation, reaching and manipulation tasks.
Figures & tables
Figure 1: How long, not how close. Left: three plans from the goal, the latent distance prefers the plan whose end state looks most like the goal frame, toward the wall, whereas a distance that counts steps prefers the plan through the door. Right: TwoRoom success as the goal moves from one to three plans away.
Figure 2: TEMPO in the planner. LeWM’s frozen encoder E and predictor WM (blue) are unchanged, and nothing is decoded. The planner rolls WM out under each CEM candidate plan and executes the one that minimises J . TEMPO adds the map Fϕ (orange) and scores each candidate plan by ∥Fϕ(z^t+25)−Fϕ(G)∥2 , the steps its end state still is from the goal, added to LeWM’s cost. Top right: Fϕ places the frames of one episode one unit apart per 25 steps.
Planner
TwoRoom
Reacher
Push-T
Cube
Scene
Avg.
Δ
LeWM
one plan
87.7 ± 0.8
94.7 ± 0.7
93.6 ± 1.0
74.8 ± 2.4
69.6 ± 1.1
84.1
–
+ TEMPO (ours)
97.2 ± 0.6
95.5 ± 0.7
95.1 ± 0.6
77.1 ± 1.7
71.1 ± 1.1
87.2
+3.1
two plans
58.5 ± 1.7
95.7 ± 0.6
49.7 ± 1.3
57.9 ± 2.4
56.4 ± 1.2
63.7
–
+ TEMPO (ours)
99.2 ± 0.3
96.8 ± 0.5
55.7 ± 1.2
58.7 ± 2.0
57.9 ± 1.3
73.7
+10.0
three plans
36.3 ± 2.3
89.7 ± 0.4
22.0 ± 0.6
68.7 ± 1.6
48.7 ± 1.3
53.1
–
Table 1: Success rates (%) on five pixel-based goal-conditioned control tasks, one to three plans from the goal. Bold marks the better of each pair. Δ is the gain of TEMPO over the same planner, averaged over environments. Every world model plans from pixels alone.
Planner
Wall †
TwoRoom cross-room
PointMaze M
PointMaze L
Ant-U
Avg.
Δ
LeWM
one plan
67.6 ± 1.4
12.0 ± 1.0
32.1 ± 1.4
28.0 ± 2.0
19.7 ± 1.0
31.9
–
+ TEMPO (ours)
75.7 ± 2.2
46.5 ± 1.2
75.7 ± 1.7
54.8 ± 1.6
21.9 ± 1.1
54.9
+23.0
two plans
57.7 ± 2.4
15.5 ± 1.0
17.6 ± 1.0
2.9 ± 0.6
25.5 ± 1.5
23.8
–
+ TEMPO (ours)
65.9 ± 2.2
69.6 ± 1.8
41.1 ± 2.0
21.3 ± 1.2
29.6 ± 1.0
45.5
+21.7
three plans
–
17.1 ± 1.1
10.8 ± 0.8
0.3 ± 0.2
28.1 ± 1.6
14.1
–
Table 2: Success rates (%) on more complex goal-conditioned control tasks, one to three plans from the goal. Reporting, bold and Δ as in Table 1 . † Wall episodes end at 50 steps, so its two-plan cell uses goal 40 and it has no three-plan cell.
WallRandom
PushObj
TwoRoom
Planner
unseen 1 plan
unseen cross-wall
unseen 1 plan
unseen 3 plans
door moved 1 plan
door moved 3 plans
LeWM ℓ planner
67.5 ± 1.7
15.1 ± 1.9
53.6 ± 2.2
18.8 ± 1.4
65.3 ± 1.9
9.9 ± 1.0
+ TEMPO (ours)
76.9 ± 1.4
22.9 ± 2.5
54.8 ± 2.0
20.5 ± 1.4
74.5 ± 1.6
30.1 ± 0.8
PLDM planner
75.6 ± 2.0
6.0 ± 1.0
44.1 ± 1.6
11.1 ± 0.8
73.9 ± 1.7
18.9 ± 1.0
+ TEMPO (ours)
77.9 ± 2.0
6.5 ± 1.1
48.8 ± 1.7
15.3 ± 1.0
76.3 ± 1.5
30.4 ± 1.0
Table 3: Generalisation to unseen configurations: a layout shift (WallRandom), a shape shift (PushObj) and a moved door (TwoRoom). World models and TEMPO are trained on the training layouts, block shapes or door position; in-distribution results are in Tables 1 and 2 . Cross-wall: the test of Zhou et al. [37] , with start and goal on opposite sides of an unseen wall. ℓ : LeWM trained with single-layout batches. Bold marks the better of each pair.
Figure 3: Training and test setups for the three generalisation families of Table 3 . Unseen configurations are outlined in orange ( ⋆ ). WallRandom holds out wall and door positions outside every training range, PushObj holds out the L, I and + shapes, and TwoRoom moves the door to the opposite end of the wall.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: The eleven environments and two RoboCasa kitchen tasks , left to right, top to bottom: TwoRoom, Reacher, Push-T, Cube, Wall, Scene, PushObj, PointMaze medium and large, Ant-U-Maze, WallRandom, and the RoboCasa kitchen tasks Pick-and-Place and Kitchen-Nav. Each panel is one 224×224 training frame, the only observation the world models and the planner receive.
Environment
Episodes (trained on)
Mean length
Action dim.
TwoRoom
10,000
92
2
Push-T
18,685
125
2
Reacher
10,000
200
2
Cube
10,000
200
5
Scene
10,000
200
5
PushObj
20,000
101
2
Appendix
Table 4: Training sets. Episodes the world models were trained on, mean episode length in environment steps, and action dimension. Where a 90:10 episode split defines the protocol (Section 4.1 ) the count is the 90% training split.
Ablation
TwoRoom
Push-T
Reacher
Cube
Scene
Avg.
LeWM
One plan from the goal
Fϕ alone ( w=1 )
98.3 ± 0.3
89.1 ± 0.9
93.5 ± 0.6
75.9 ± 2.1
68.5 ± 1.5
85.0
Latent +Fϕ (ours, w=0.5 )
97.2 ± 0.6
95.1 ± 0.6
95.5 ± 0.7
77.1 ± 1.7
71.1 ± 1.1
87.2
Two plans
Fϕ alone ( w=1 )
97.2 ± 0.7
30.4 ± 1.9
96.5 ± 0.5
52.3 ± 2.1
52.5 ± 1.5
65.8
Appendix
Table 5: Ablation of the planning cost on each frozen world model. The world model and planner are fixed and only the cost changes: latent +Fϕ is the cost of Eq. ( 5 ) used in every other table, and Fϕ alone drops the latent term ( w=1 ). Bold marks the better of each pair. Avg.: mean over the five environments.
Condition
TwoRoom
Push-T
Reacher
Cube
Scene
Avg.
in-distribution — LeWM planner
36.5 ± 2.1
21.1 ± 1.3
89.2 ± 0.5
69.5 ± 1.3
47.7 ± 1.0
52.8
in-distribution — + TEMPO (ours)
98.4 ± 0.3
26.1 ± 1.4
97.7 ± 0.4
70.0 ± 1.0
49.5 ± 1.7
68.3
blur — LeWM planner
25.9 ± 2.2
20.7 ± 1.1
88.5 ± 0.5
68.0 ± 1.6
42.0 ± 1.4
49.0
blur — + TEMPO (ours)
38.1 ± 2.4
26.5 ± 1.7
96.0 ± 0.6
68.9 ± 1.5
40.5 ± 1.5
54.0
colour jitter — LeWM planner
34.1 ± 3.3
6.3 ± 1.3
24.5 ± 5.2
59.2 ± 4.3
41.1 ± 2.1
33.0
colour jitter — + TEMPO (ours)
74.3 ± 8.0
12.8 ± 3.2
23.2 ± 6.5
61.5 ± 4.5
42.1 ± 2.2
42.8
Appendix
Table 6: Same out-of-distribution protocol on every environment (goal 75 steps ahead). The pixel corruptions are generic wrappers on the rendered frame, applied with identical parameters in every environment (values as in 35 ); the reference row runs the same protocol with no corruption, so the first frame is re-rendered live in every row. World model frozen, Fϕ trained online in the corrupted environment, blend cost of Eq. ( 5 ). Bold: the better of each pair.
Condition
Wall †
TwoRoom cross-room
PointMaze M
PointMaze L
Ant-U
Avg.
in-distribution — LeWM planner
57.7 ± 2.4
17.1 ± 1.1
10.8 ± 0.8
0.3 ± 0.2
28.1 ± 1.6
22.8
in-distribution — + TEMPO (ours)
65.9 ± 2.2
74.5 ± 1.7
31.7 ± 1.0
6.5 ± 1.0
34.0 ± 2.1
42.5
blur — LeWM planner
52.1 ± 2.8
8.9 ± 1.4
10.4 ± 0.9
0.1 ± 0.1
12.5 ± 1.4
16.8
blur — + TEMPO (ours)
60.9 ± 2.4
23.3 ± 2.5
17.9 ± 1.7
4.1 ± 1.0
12.4 ± 1.5
23.7
colour jitter — LeWM planner
56.4 ± 2.6
16.1 ± 1.5
9.6 ± 1.0
1.6 ± 0.6
17.7 ± 1.7
20.3
colour jitter — + TEMPO (ours)
68.0 ± 2.4
59.7 ± 6.9
10.4 ± 1.3
1.7 ± 0.7
19.7 ± 1.3
31.9
Appendix
Table 7: The same pixel-corruption protocol on the more complex environments of Table 2 (goal 75 steps ahead). The pixel corruptions are generic wrappers on the rendered frame, applied with identical parameters in every environment (values as in 35 ); the in-distribution rows repeat the corresponding cells of Table 2 . † Wall episodes end at 50 steps, so its goal is 40 steps ahead; cross-room uses its 150-step budget. World model frozen, Fϕ trained online in the corrupted environment, blend cost of Eq. ( 5 ). Bold: the better of each pair.
Figure 5: Test-time pixel corruptions: cumulative success versus MPC step, goal 75 steps ahead (30 steps = budget 150). Rows: condition; columns: environment. Dashed blue: the LeWM planner with its own cost; orange: the same planner + TEMPO (ours), on identical windows.
Environment
Shift
Planner
+ TEMPO
Push-T
none (in distribution)
21.1
26.1
larger agent
18.3
26.3
TwoRoom
none (in distribution)
36.5
98.4
door moved
9.9
30.3
larger agent
36.4
94.9
Reacher
none (in distribution)
89.2
97.7
Appendix
Table 8: Factor shifts on Push-T, TwoRoom, Reacher, Cube and Scene, goal 75 steps ahead. Protocol of Wang et al. [35] : the first frame is rendered live under the shift, the goal frame stays the clean dataset frame, the world model is frozen and Fϕ is trained online in the shifted environment. Shifts change the geometry, the dynamics or the camera. Success (%), LeWM planner and the same planner with TEMPO ; mean over three seeds; bold: the better of the pair.
Planner
Pick-and-Place (SR ↑ )
Kitchen-Nav (SR ↑ )
LeWM
one plan
0.32
0.16
+ TEMPO (ours)
0.36
0.22
two plans
0.20
0.19
+ TEMPO (ours)
0.21
0.27
three plans
0.14
0.23
Appendix
Table 9: Planning results for offline world models on two RoboCasa kitchen tasks, one to three plans from the goal. Success rate as a fraction of episodes. Bold marks the better of each pair.
Fϕ -distance
R2 of frame index
Environment
same index, other episode
25 steps, same episode
from Fϕ(z)
from z
TwoRoom
1.48
1.03
0.11
0.13
Push-T
2.27
1.09
0.54
0.73
Reacher
2.27
1.95
0.01
0.03
Cube
1.63
1.07
0.03
0.13
Scene
2.45
1.03
0.21
0.32
Appendix
Table 10: Fϕ does not encode the frame index. Median Fϕ -distance (units of 25 steps) between states with the same frame index in different episodes, against states 25 steps apart in the same episode, and R2 of a linear probe predicting the frame index from Fϕ(z) and from the raw latent z .
Training episodes
Held-out episodes
Environment
ρ
error
ρ
error
PointMaze M
0.83
0.28
0.83
0.28
PointMaze L
0.99
0.06
0.99
0.06
Ant-U
0.56
0.56
0.51
0.59
Appendix
Table 11: Fϕ generalises to held-out episodes. Spearman ρ between the Fϕ -distance and the step gap k , and mean absolute error of the Fϕ -distance against its target k/25 (units of 25 steps), over about 24,000 same-episode pairs with k≤75 (LeWM, mean over three seeds).
benchmark / offset
windows
window success
ρ (score, realised dist)
ρ (score, success)
top-5 success
oracle top-5
PushT / 25
59
0.74
+0.83
+0.32
0.89
1.00
TwoRoom / 25
48
0.80
+0.49
+0.21
0.84
0.91
PushT / 75
60
0.28
+0.66
+0.02
0.29
0.65
Reacher / 75
60
0.95
+0.85
+0.22
0.97
0.99
TwoRoom / 75
60
0.42
+0.66
−0.07
0.41
0.75
Appendix
Table 12: What the latent distance ranks. We diagnose the two tasks where the flat planner collapses three plans from the goal (TwoRoom from 88 to 36% , Push-T from 94 to 22% ); Reacher, where it does not collapse, is the control, and the one-plan rows show the same cost working close to the goal. 50 executed action sequences per window, scored by the planner’s own latent distance to the target; ρ (score, realised distance) and ρ (score, eventual success) over the sequences of a window; top-5 = success of the five best-scored sequences, oracle = best five in hindsight.
ρ (cost, success)
top-5 success
Environment
latent
TEMPO
latent
TEMPO
best five
TwoRoom
-0.08
+0.17
26
35
57
Push-T
+0.06
+0.13
30
30
51
Reacher
+0.25
+0.33
95
98
99
Cube
+0.39
+0.38
32
32
42
Scene
+0.02
+0.04
18
21
60
Appendix
Table 13: Which executed plans each cost ranks first, three plans from the goal. For 20 windows per seed, 50 candidate plans are executed from the same state for 25 steps and then finished by the flat planner; each candidate is scored by LeWM’s latent distance and by TEMPO ’s cost. ρ : Spearman correlation between the (negated) cost and eventual success; top-5: success rate (%) of the five best-scored candidates; best five: the success rate of the five candidates that actually succeed most. Mean over three seeds; bold: the better of the pair.
Latent world models are judged by how well they predict, so when planning fails at long horizons the natural reading is that the predictor degrades. On a reproduction of LeWorldModel on TwoRoom we show the binding constraint is the planner's objective instead. The predictor is not the limit: its imagined state seventy-five environment steps ahead is still only 0.189 as wrong as assuming the world froze, while the planner never imagines beyond twenty-five. The objective is. Cross-entropy-method planning minimises squared latent distance, which tracks true distance at r = 0.426, saturates by about eighty arena units and decreases beyond a hundred and twenty, so moving away from the goal can lower the cost. The information is present throughout: a ridge probe recovers position from the frozen embedding at R^2 0.9922. The pathology is the method's, not one reimplementation's. It is present in the authors' released weights, and across four checkpoints long-horizon success rank-orders exactly with metric quality and inversely with prediction accuracy. Replacing only the objective, with nothing retrained and no GPU, lifts goals reached at offset 100 from 26.0% to 98.0%, equals the 98.0% at offset 25, and reaches 92.0% under a third of the budget: planning stops depending on the horizon. The best cost is not the most accurate. A head learned from frame separation alone predicts spatial distance worse than a position probe (r = 0.819 against 0.9897) yet plans better, charging 24% more to cross the environment's dividing wall where squared latent distance charges 4% less. It has learned reachability, not proximity.
Joint-Embedding Predictive Architectures (JEPAs) learn world models by predicting in representation space rather than reconstructing pixels, making them a natural backbone for latent model predictive control from offline demonstration logs. JEPA-style training optimizes short-horizon latent prediction, whereas planning requires a multi-step ranking of imagined futures by goal progress. Prior JEPA planners often inherit that ranking from embedding geometry, typically latent Euclidean distance, which arises as a byproduct of representation learning rather than as a progress cost mined from the logs. We propose Temporal-Distance-JEPA, which retains the LeWM encoder--predictor backbone and mines a directed temporal cost from reward-free trajectories: same-trajectory step order supplies positive targets, cross-trajectory pairs act as heuristic negatives, and a rollout-consistency term matches the planner horizon. The mined supervision serves two roles: as the deployed planning cost when progress is topological, and as a representation signal that improves Euclidean planning when contact geometry dominates. Under locked evaluation, deploying the mined cost raises Two-Room success to 100.0% versus LeWM's 97.4%, while shared Euclidean planning on the same temporally trained checkpoint raises OGB-Cube by 14.2 points over LeWM and improves Push-T. Against LeWM and the concurrent RC-aux baseline under locked evaluation, Temporal-Distance-JEPA matches or exceeds both methods on every environment. Ablations show that the directed head, cross-trajectory negatives, and rollout consistency each contribute. Temporal-Distance-JEPA narrows the train--plan gap for JEPA world-model planners by discovering temporal progress structure in offline logs and co-designing cost form with plan-time deployment. Code is available at https://github.com/HKBU-KnowComp/Temporal-Distance-JEPA.
A latent world model may achieve accurate short-horizon prediction while still inducing a latent space that is poorly aligned with planning. A key issue is spatiotemporal mismatch: these models are often trained with local predictive supervision, but deployed for long-horizon goal-directed search in latent spaces where Euclidean distance may not reflect what is reachable within a finite action budget. We present the Reachability-Correction auxiliary objective (RC-aux), a lightweight correction for this mismatch in reconstruction-free latent world models. RC-aux keeps the world-model backbone unchanged and adds planning-aligned supervision along two axes. Along the time axis, multi-horizon open-loop prediction trains the model beyond one-step consistency. Along the space axis, budget-conditioned reachability supervision, together with temporal hard negatives, encourages the latent space to distinguish states that are eventually reachable from those reachable within the current planning horizon. At test time, the learned reachability signal can also be used by a reachability-aware planner to favor trajectories that are both goal-directed and attainable under the available budget. We instantiate RC-aux on LeWorldModel and evaluate it under both continuation-training and matched-from-scratch settings. Across goal-conditioned pixel-control tasks and a LIBERO-Goal extension, RC-aux improves LeWM-style planning with modest additional cost. These results suggest that planning with latent world models depends not only on predictive accuracy, but also on whether the learned representation encodes the temporal and geometric structure required by downstream search. The code is available at https://github.com/Guang000/RC-aux.