Latent world models learn action-conditioned dynamics in representation space and often score candidate actions by Euclidean distance to a goal representation. Joint training typically regularizes the representation to prevent collapse, but the resulting representation geometry also determines how terminal errors are weighted during planning. We show that accurate prediction and noncollapsed representations do not guarantee a task-aligned latent planning cost: isotropic Gaussian regularization can induce a geometry that ranks feasible outcomes differently from the task cost. To address this mismatch, we introduce AnisoWM with ΛReg, which replaces the fixed isotropic Gaussian target with a learnable diagonal covariance under fixed-trace and anisotropy constraints. The prediction objective, predictor architecture, and Euclidean planner remain unchanged; the target is used only during training. Our analysis characterizes the prediction-driven allocation of target variance, its dependence on the training distribution, and the conditions under which the induced metric reduces planning regret. Across four visual control environments, AnisoWM improves planning success over LeWorldModel in all four. Its latent planning cost also shows better agreement with task outcomes. Project website: https://rkdrn79.github.io/AnisoWM-page/
Figures & tables
Figure 1: Prediction–planning separation with a nonlinear encoder. (a) The Euclidean task cost and the SIGReg latent cost select different outcomes from the same reachable set; the learned latent geometry closely follows the Σ−1 reference. (b) As process noise decreases, held-out conditional-mean prediction error decreases while normalized physical planning regret remains nearly unchanged. Points and error bars show the mean and 95%t -confidence interval over ten training seeds.
Figure 2: Planning success across environments. AnisoWM with Λ Reg at κ=2 in every environment, compared with LeWM and the reference baselines. Reported AnisoWM and LeWM values are means over three training seeds.
Jenc (representation only)
Jpred (planning cost)
Task
LeWM
AnisoWM
Gap
LeWM
AnisoWM
Gap
TwoRoom
0.564
0.661
+0.097
0.575
0.646
+0.071
Reacher
0.790
0.893
+0.103
0.676
0.723
+0.048
PushT
0.647
0.641
−0.005
0.594
0.621
+0.028
Cube
0.555
0.576
+0.021
0.537
0.553
+0.016
Table 1: Ordering recorded action sequences by the outcome they reached. For each initial–goal pair, both models score the same sequences, and the entry is the fraction of sequence pairs whose cost ordering agrees with their outcome ordering. Jenc encodes the observation a sequence actually reached; Jpred is the cost used for planning and additionally includes the predictor rollout.
Figure 3: Latent cost neighborhoods around the goal. For Cube and PushT, the left panels show the rendered arena and evaluated region, and the right panels show positions in the lowest 10% of latent planning cost around the goal ( ⋆ ) for LeWM (blue) and AnisoWM ( κ=2 , red). The dashed circle shows the corresponding neighborhood under physical Euclidean distance.
Figure 4: Learned target spectra. Left: learned targets at κ=2 , sorted by variance, with the isotropic target marked. The trace and maximum variance ratio are fixed, while the allocation is learned; the legend shows the number of coordinates above the mean variance. Right: target spectrum during one TwoRoom run, showing continued evolution after the condition-number bound is reached.
Figure 5: Planning success across anisotropy bounds. Each point with κ>1 reports one training run evaluated on the same set of initial–goal pairs. The κ=1 marker shows the isotropic LeWM baseline from the primary comparison.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Notation
Description
ot,og
Observation at time t and goal observation.
fθ,gϕ
Neural encoder and action-conditioned latent predictor.
zt,zg
Latent representation at time t and goal representation.
at,U
Per-step action and finite-horizon action sequence U=(a0,…,aH−1) .
H
Planning horizon.
z^H(U)
Predicted terminal representation after rolling out U .
Appendix
Table 2: Notation used throughout the paper.
Figure 6: Evaluation environments. From left to right: TwoRoom, PushT, Reacher, and Cube. All four environments have continuous action spaces, and visual goal planning is performed directly from image observations.
Setting
Value
Epochs
15
Training seeds
3072, 1234, 42
Optimizer / batch size
AdamW / 128
Peak learning rate
5×10−5
Encoder and predictor weight decay
10−3
Target-logit learning-rate schedule
Same as the model
Appendix
Table 3: Implementation settings. Target logits use zero weight decay and are clipped after each update as described in Section 4.2 .
Task
Seed 3072
Seed 1234
Seed 42
Mean
TwoRoom
90
94
94
92.7
Reacher
90
88
88
88.7
PushT
98
96
96
96.7
Cube
78
80
80
79.3
Appendix
Table 4: Per-seed AnisoWM planning success (%). Each entry reports AnisoWM planning success under the same evaluation protocol used for Figure 2 . The final column reports the mean across the three training seeds, which the main text rounds to an integer.
Task
Outcome scalar
Threshold
H
Cases
TwoRoom
agent–goal distance
16 px
20
45
Reacher
largest joint-position error
0.05 rad
25
50
PushT
max(position/20px,angle/(π/9))
1
25
48
Cube
block–target distance
0.04 m
15
33
Appendix
Table 5: Outcome scalar and diagnostic horizon. Each scalar is normalized by its threshold, so one is the success boundary and lower is better. Cases are counted out of the fifty initial–goal pairs of the evaluation protocol.
Figure 7: Ranking accuracy, pair by pair. Each point is one initial–goal pair: how well LeWM’s cost orders its sequences on the horizontal axis against how well AnisoWM’s does on the vertical one, both scoring the same sequences. The shaded half-plane above the diagonal is where AnisoWM orders that pair better. Top row: Jenc , which leaves the representation on its own. Bottom row: Jpred , the cost the planner minimizes. Rows of a column share their axis limits; each panel shows the seed Table 1 reports.
Figure 8: Additional latent-cost neighborhoods around the goal. Four additional Cube/PushT example pairs, shown using the same visualization as Figure 3 . The dashed circle denotes the task-cost neighborhood in physical space; blue and red contours denote the lowest-cost regions under LeWM and AnisoWM ( κ=2 ), respectively.
Figure 9: Evolution of the learned target condition number. Each curve shows one TwoRoom training run with the indicated anisotropy bound, and the dotted line at each bound is the ratio that bound permits. Filled markers indicate that the learned target reaches its bound.
κ
TwoRoom
Reacher
PushT
Cube
1.25
90
90
94
82
2
90
90
98
78
4
92
84
92
78
8
92
84
92
78
16
94
86
92
78
Appendix
Table 6: Planning success across anisotropic target bounds (%). Each entry corresponds to the training run used in the sweep of Figure 5 .
Figure 10: Measured feature-covariance spectra in TwoRoom. Eigenvalues of the centered feature covariance are normalized by their mean and sorted independently for each displayed anisotropy bound; the dashed line marks index 100 , the value quoted in the text. These empirical feature covariances are distinct from the learned target covariance Λ .
Figure 11: Latent prediction loss and planning success. One marker per trained run, coloured by its anisotropy bound. A run further left predicts better and a run higher plans better. The red ring marks the best-planning run and the grey ring the lowest-loss one; where a single run is both, the two rings are concentric. The MSE is measured in each model’s own learned representation space.
Figure 12: Rollouts in TwoRoom and Reacher. Two rollout pairs per environment compare AnisoWM at κ=2 with LeWM under the same evaluation planner and seed. Columns show the indicated environment steps and the goal observation; red outlines indicate recorded goal attainment.
Figure 13: Rollouts in PushT and Cube. Two rollout pairs per environment compare AnisoWM at κ=2 with LeWM under the same evaluation planner and seed. Columns show the indicated environment steps and the goal observation; red outlines indicate recorded goal attainment.
Modern vision-based world models can represent observations as compact yet expressive latent manifolds, but fast goal-oriented planning in these spaces remains challenging. This raises a central question: when does a learned representation simplify control, rather than merely enabling prediction? We study this question in a pretrained LeWorldModel, whose latent geometry is regularized for smoothness and uniformity. Our key insight is that, under such geometry, planning can be amortized into a latent inverse-dynamics mapping instead of requiring online search. We therefore replace iterative planning with a lightweight Goal-Conditioned Inverse Dynamics Model (GC-IDM) that maps the current latent state, goal latent state, and remaining horizon directly to the next action. Empirically, across four benchmark environments spanning navigation, contact-rich manipulation, and continuous control, our controller matches or exceeds CEM in seven of eight environment-protocol settings while reducing per-decision cost by 100-130x. A broader sweep over test-time planners (CEM, MPPI, iCEM, and gradient-based methods) shows that this result is not specific to a particular optimizer. These findings suggest that much of the structure recovered by test-time planning is already locally encoded in the latent representation. More broadly, our results indicate that sufficiently structured latent spaces can shift part of the planning burden from online optimization to learned inference. Our code is publicly available at https://github.com/hdnndh/Latent-Geometry-Beyond-Search-Amortizing-Planning-in-World-Models .
Hoang Nguyen, Xiaohao Xu, Xiaonan Huang
Department of Robotics, University of Michigan, Ann Arbor
Latent world models rely on representation geometry for planning, yet regularizing the latent marginal alone does not determine the state-to-state relationships used for action selection. We show that this can cause planning-relevant novelty structure to be weakened as representations are transformed into the final latent used by the planner. We introduce Aligned Transport of Latent Structure (ATLAS), a training objective that explicitly preserves relational geometry while calibrating the global latent distribution. ATLAS transfers normalized pairwise structure from an informative encoder representation to the planning latent and uses Wasserstein embedding matching (WEMReg) to calibrate its marginal through one-dimensional Wasserstein-2 transport. Our analysis shows that relational preservation and marginal calibration impose non-redundant constraints, and connects finite-candidate planning stability to relational distortion, latent-scale mismatch, and prediction error. Instantiated in LeWM, ATLAS improves mean goal-reaching success across PushT, TwoRoom, and OGBench-Cube on both lower- and higher-novelty evaluation subsets, with the largest gain on higher-novelty TwoRoom episodes. Representation and rollout diagnostics further show stronger novelty-related structure in the planning latent, improved marginal calibration, and lower multi-step prediction error. Together, these results highlight preservation of planning-relevant latent geometry as an important ingredient for reliable world-model planning. Code is available at https://anonymous.4open.science/r/atlas-world-model-72C4/.
Joint-Embedding Predictive Architectures (JEPAs) learn world models by predicting in representation space rather than reconstructing pixels, making them a natural backbone for latent model predictive control from offline demonstration logs. JEPA-style training optimizes short-horizon latent prediction, whereas planning requires a multi-step ranking of imagined futures by goal progress. Prior JEPA planners often inherit that ranking from embedding geometry, typically latent Euclidean distance, which arises as a byproduct of representation learning rather than as a progress cost mined from the logs. We propose Temporal-Distance-JEPA, which retains the LeWM encoder--predictor backbone and mines a directed temporal cost from reward-free trajectories: same-trajectory step order supplies positive targets, cross-trajectory pairs act as heuristic negatives, and a rollout-consistency term matches the planner horizon. The mined supervision serves two roles: as the deployed planning cost when progress is topological, and as a representation signal that improves Euclidean planning when contact geometry dominates. Under locked evaluation, deploying the mined cost raises Two-Room success to 100.0% versus LeWM's 97.4%, while shared Euclidean planning on the same temporally trained checkpoint raises OGB-Cube by 14.2 points over LeWM and improves Push-T. Against LeWM and the concurrent RC-aux baseline under locked evaluation, Temporal-Distance-JEPA matches or exceeds both methods on every environment. Ablations show that the directed head, cross-trajectory negatives, and rollout consistency each contribute. Temporal-Distance-JEPA narrows the train--plan gap for JEPA world-model planners by discovering temporal progress structure in offline logs and co-designing cost form with plan-time deployment. Code is available at https://github.com/HKBU-KnowComp/Temporal-Distance-JEPA.