World models allow agents to plan in latent space by choosing a sequence of actions that most reduces the distance to a given goal state. Thus, planning can benefit from latent representations whose distances mirror commute-times in the environment. The spectral embedding space of the graph Laplacian provides such a representation, if it obeys a specific eigenvalue-dependent scaling. Unfortunately, instantiating the graph Laplacian is intractable in large, continuous environments. Self-supervised learning offers a natural route to such commute-time-preserving embeddings at scale. However, here we show that existing methods, which commonly encourage isotropic representations to prevent representational collapse, tend to degrade the "correct" eigenvalue-dependent scaling, leading to an inaccurate representation of commute times. To address this problem, we introduce Commute-Time-Preserving World Models (CTWMs), combining a latent displacement predictor and a log-determinant regularizer that prevents collapse, which provably recover the correctly scaled Laplacian representation under reversible deterministic dynamics and at the predictor's fixed point. In numerical simulations, CTWM matches or outperforms LeWM, a task-agnostic baseline, on several complex, continuous goal-reaching benchmarks, while using half the parameters.
Figures & tables
Figure 1: (a) Overview of CTWM: a shared encoder ϕ maps observations s and s′ to the representations z and z′ respectively. The transition model T conditioned on action a , predicts a residual displacement Δ that is added to the current embedding via a residual connection. The loss L(ϕ,T) is computed from the predicted and actual next-state embeddings. (b) Geometric interpretation of the loss L(ϕ,T⋆) in the latent space; arrow colors refer to terms in equation 6 . The loss terms encourage predictable transitions (teal), penalize large distances between consecutive points (blue), and prevent collapse of representation (magenta). (c) Latent distances reflect commute-time distance rather than spatial proximity. Left: Two-room connected by a door; s1 and s3 are adjacent but separated by a wall (dashed line). Right: in the CTWM embedding the rooms unfold around the door, placing s1 near s2 and far from s3 .
Figure 2: Torus environment and empirical verification of the theoretical predictions. (a) The torus graph: each vertex carries a unique three-digit label, randomly shuffled across the graph, observed as three MNIST images sampled independently from the corresponding classes at each visit. (b) Subspace similarity to the ground-truth Laplacian eigenbasis for the four methods ± SD ( n=4 ; see Appendix F ); the dashed line indicates the ceiling imposed by the degenerate eigenspectrum. (c) Variance ratio var(zk)/var(z1) against eigenvalue index k (cf. Corollary 1.2 ), with the ground-truth scaling λ1/λk in dashed black. Only CTWM tracks it, confirming that the logdet regularizer recovers the λk−1/2 scaling; VICReg and WM-GDO show more isotropic spectra, as encouraged by their respective regularizers. (d) Squared L2 latent distance versus ground-truth commute-time distance, one point per pair of graph nodes. Black line: least-squares linear fit, with its coefficient of determination R2 . CTWM achieves the highest R2 .
Figure 3: Success rates of CTWM and LeWM on the six benchmark tasks: Reacher, Cube, Two-Room, PushT, Scene and Pointmaze. CTWM (red) matches or outperforms LeWM (blue) on every environment using half the number of parameters. Shaded regions indicate the SD ( n=3 ).
Model
Reacher
Cube
Two-Room
PushT
Scene
Pointmaze
Random
13.0 2.3
45.2 0.9
24.3 2.7
4.3 2.1
49.8 3.1
14.3 1.7
LeWM
83.3 2.1
64.8 1.3
88.8 1.6
91.5 1.1
69.3 2.6
68.5 1.8
CTWM
82.3 2.6
89.8 1.0
100.0 0.0
92.2 0.6
82.3 1.4
100.0 0.0
Random-100
7.3 1.9
45.5 2.5
1.0 0.4
0.3 0.2
24.2 1.7
11.3 1.0
LeWM -100
81.0 0.8
61.7 1.4
17.3 0.8
15.3 1.4
40.3 2.5
28.2 1.2
CTWM -100
98.2 0.8
73.0 1.2
98.3 0.9
26.5 2.6
57.5 1.6
98.8 0.2
Table 1: Planning success rate in percent for each model trained on three seeds after ten epochs using CEM for the standard setting by Maes et al. (2026a) with a goal distance of 25 steps and planning budget of 50 steps (top half) and for an extended goal distance of 100 steps with planning budget of 125 steps (bottom half). A baseline floor is given by the random policy for each setting and environment.
Figure 4: Commute-time distance and variance scaling on Pointmaze. (a) The view of the Pointmaze environment. (b) Ground truth commute-time distance from the blue star to any other point on the map computed from the discretized continuous state space. (c) Top row: Relative L2 distance (normalised) in the embedding space of LeWM (left) and CTWM (right). Bottom row: difference to the ground-truth commute-time distance. CTWM reproduces the ground-truth distance more closely. (d) Variance of the first 64 latent features for both models. CTWM exhibits the λk−1/2 decay predicted by Corollary 1.2 (dashed line), whereas LeWM produces a flatter spectrum.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
Value
Architecture
encoder ϕ
ViT-Tiny (5.5 M)
projector
MLP with LayerNorm
predictor T
Transformer architecture of ( Maes et al., 2026b ) of width D (3.3 M)
Data
image size
224×224
context length H
3
frame skip k
5
Appendix
Table 2: Main hyperparameters. ( Green ) Changes from LeWM.
β
βest
0.005
0.0074
0.1
0.108
0.25
0.26
0.5
0.52
1.0
1.04
Appendix
Table 3: Comparison of β and βest on PushT.
Model
Reacher
Cube
Two-Room
PushT
Scene
Pointmaze
Random
13.0 2.3
45.2 0.9
24.3 2.7
4.3 2.1
49.8 3.1
14.3 1.7
LeWM
73.7 2.7
60.2 0.5
93.7 2.5
86.5 1.2
66.7 2.4
88.3 1.2
CTWM
77.0 3.3
78.2 2.1
98.7 0.8
87.5 1.1
72.8 2.5
100.0 0.0
Appendix
Table 4: Planning success rate in percent for each model trained on three seeds after ten epochs using Adam. A floor is given by the random policy for each environment.
Latent world models plan by rolling a frozen predictor forward under candidate action sequences and ranking the candidates by the latent distance between their imagined end state and the goal. However, this ranking breaks down when the goal lies several plans away, because the latent distance measures how closely an end state resembles the goal rather than how far it remains from reaching it. To address this, we propose TEMPO, a temporal-distance planning objective that leaves the world model untouched, learns only from the demonstrations already used to train it, and adds negligible cost to the planner's search. TEMPO learns a small map of the frozen latent in which the distance between two states of an episode reflects the number of environment steps between them, and blends this distance into the planner's cost. It requires no rewards, policies or success labels and, being a cost rather than a model, applies to frozen world models with one latent vector per state that plan by a latent distance. We evaluate TEMPO on eleven simulated environments (e.g., maze navigation, tabletop pushing, robotic arm control and three-dimensional manipulation) with the LeWM and PLDM planners. With a small MLP that adds at most 0.3% to a plan's arithmetic, TEMPO improves both planners at every goal distance, including the one-plan setting of their evaluations, raises LeWM from 36% to 99% on TwoRoom three plans from the goal, and remains competitive on a broad range of 2D and 3D navigation, reaching and manipulation tasks.
Modern vision-based world models can represent observations as compact yet expressive latent manifolds, but fast goal-oriented planning in these spaces remains challenging. This raises a central question: when does a learned representation simplify control, rather than merely enabling prediction? We study this question in a pretrained LeWorldModel, whose latent geometry is regularized for smoothness and uniformity. Our key insight is that, under such geometry, planning can be amortized into a latent inverse-dynamics mapping instead of requiring online search. We therefore replace iterative planning with a lightweight Goal-Conditioned Inverse Dynamics Model (GC-IDM) that maps the current latent state, goal latent state, and remaining horizon directly to the next action. Empirically, across four benchmark environments spanning navigation, contact-rich manipulation, and continuous control, our controller matches or exceeds CEM in seven of eight environment-protocol settings while reducing per-decision cost by 100-130x. A broader sweep over test-time planners (CEM, MPPI, iCEM, and gradient-based methods) shows that this result is not specific to a particular optimizer. These findings suggest that much of the structure recovered by test-time planning is already locally encoded in the latent representation. More broadly, our results indicate that sufficiently structured latent spaces can shift part of the planning burden from online optimization to learned inference. Our code is publicly available at https://github.com/hdnndh/Latent-Geometry-Beyond-Search-Amortizing-Planning-in-World-Models .
Hoang Nguyen, Xiaohao Xu, Xiaonan Huang
Department of Robotics, University of Michigan, Ann Arbor
Latent world models learn action-conditioned dynamics in representation space and often score candidate actions by Euclidean distance to a goal representation. Joint training typically regularizes the representation to prevent collapse, but the resulting representation geometry also determines how terminal errors are weighted during planning. We show that accurate prediction and noncollapsed representations do not guarantee a task-aligned latent planning cost: isotropic Gaussian regularization can induce a geometry that ranks feasible outcomes differently from the task cost. To address this mismatch, we introduce AnisoWM with ΛReg, which replaces the fixed isotropic Gaussian target with a learnable diagonal covariance under fixed-trace and anisotropy constraints. The prediction objective, predictor architecture, and Euclidean planner remain unchanged; the target is used only during training. Our analysis characterizes the prediction-driven allocation of target variance, its dependence on the training distribution, and the conditions under which the induced metric reduces planning regret. Across four visual control environments, AnisoWM improves planning success over LeWorldModel in all four. Its latent planning cost also shows better agreement with task outcomes. Project website: https://rkdrn79.github.io/AnisoWM-page/
Mingu Kang, Yoori Oh, Sookyung Kim +1
Seoul National University · Ewha Womans University