Latent world models plan by scoring candidate action sequences with distances in latent space. However, task success is judged by physical quantities, which we call the success-criterion quantities. In all four latent world models we examine, the end-effector position is encoded in the latent state with an error larger than the success criterion allows. Such a latent state cannot separate successful candidates from failing ones. We propose an auxiliary loss that uses success-criterion quantities as training targets, whereas existing latent world models take them only as inputs. During training, a linear head on the encoder and predictor outputs regresses the success-criterion quantities, and the regression error is added to the training loss. The head is discarded after training, so the model, its cost, and its inputs at test time are unchanged. This loss alone improves the success rate on PushT and cube by 3.5% and 3.4% (absolute), respectively, and both improvements are statistically significant. A success criterion thus specifies what a world model must retain in its latent state, and we show that it can serve directly as a training target.
Figures & tables
Figure 1: Planning with a latent world model and the auxiliary loss . (a) Planning with a latent world model, which is the same with and without the auxiliary loss. Red marks the auxiliary-loss path, which acts only at training time: a shared linear head reads the success-criterion quantities out of the latent states of the encoder and the predictor, and the error is added to the training loss. (b) Readout error of the hand position of LeWM on stack, the median over 5000 images not used for training; the dashed line is the tolerance of the success criterion, 2 cm. For the model supervised with the hand ( λ=1 , Table 7 ) the error falls inside the tolerance.
Task
Success criterion
Origin
PushT
∥agent−goal∥2+∥block−goal∥2<20∧∣Δθ∣<20∘
Prior evaluation
cube
∥block−goal∥<4cm
Prior evaluation
Reacher
∣q−qgoal∣<0.05rad (all joints)
Prior evaluation
stack
∥obj−goal∥<4cm∧∥hand−goal∥<2cm
Ours
coffee
∥pod−goal∥<4cm∧∥hand−goal∥<2cm
Ours
Table 1: Form of the success criteria. The top three rows are the criteria of the evaluation used by LeWM; the bottom two rows are criteria we defined and fixed before the auxiliary-loss experiments. agent is the disk moved by the policy, block the pushed object, obj the stacked object, hand the end effector, q the joint angles of Reacher, pod the coffee pod, goal the value in the respective goal state, and Δθ the difference of the block angle from the goal.
Task (baseline success rate)
Supervised quantity
Precondition
Diff. [%]
stack (30.1%)
hand only
◯
−1.30
object only
◯
−0.40
both
◯
+5.05∗
coffee (34.0%)
hand only
◯
+0.00
object only
◯
+3.90∗
both
◯
−0.60
Table 2: Change in success rate . Success rate of the baseline model and the difference of the auxiliary-loss model from it (absolute, %). LeWM uses identical pairs of 2000 trials; the DINO-WM type identical pairs of 1000 trials. Criteria as in Table 1 ; test: paired two-sided McNemar; ∗ marks p<0.05 . “Precondition” is whether the supervised quantity is in the success criterion and visible in the observed image; the three arms below the double rule do not satisfy it (the hand in cube and the fingertip in Reacher are not in the criterion, and the joint angles of Reacher are not visible in the image). One training seed per model; auxiliary-loss weights in Appendix A .
Task
Supervised quantity
Overall
Reach-only
Manipulation
stack
baseline
30.1%
44.8% ( n=980 )
16.1% ( n=1020 )
hand only
−1.30
+5.92∗
−8.24∗
object only
−0.40
−7.86∗
+6.76∗
both
+5.05∗
+9.08∗
+1.18
coffee
baseline
34.0%
41.7% ( n=1296 )
19.9% ( n=704 )
hand only
+0.00
+2.01
−3.69∗
Table 3: Effect per trial stratum (four tasks) . Absolute difference in success rate from the baseline model [%]. Trial pairs, criteria, test, seeds, and weights as in Table 2 . Baseline rows give the baseline success rate and the n of each stratum. A trial is a manipulation trial if the object displacement between start and goal is at least the object tolerance of the criterion (stack and coffee 4 cm, PushT 20 units), and a reach-only trial otherwise. Reach-only trials of cube satisfy the criterion at the start because it has no hand term, so they are not compared (—). The 8 stack models with varied weights and seeds are in Appendix D .
Figure 2: Examples of the trial strata (stack) . The start frame and the goal frame of the demonstration for a reach-only trial and a manipulation trial. The strata are those of Table 3 ; a trial with an object displacement of at least 4 cm is a manipulation trial. Solid circles mark the cube positions, dashed circles the positions at the start, and the number by the arrow is the displacement from start to goal. In the reach-only trial the dashed circle coincides with the solid one. Examples for all five tasks are in Figure 3 of Appendix C .
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
System
Task
Supervised quantity
λ hand
λ object
Added parameters
LeWM
PushT
hand only
1
—
386
object only
—
0.1
579
both
1
0.1
965
LeWM
cube
hand only
0.03
—
579
object only
—
0.1
579
both
1
0.1
1,158
Appendix
Table 4: Auxiliary-loss weights and added parameters . For the auxiliary-loss models in Tables 2 and 3 , the weights λ for the hand and the object and the number of parameters added by the linear heads. — means the head is not used. The λ of the 8 stack models with varied weights and seeds are given in Figure 5 .
Task
Source of demonstrations
Demos
Total frames
Action dim.
PushT
Diffusion Policy (redistributed by LeWM)
18,685
2,336,736
2
cube
OGBench (redistributed by LeWM)
10,000
2,010,000
5
Reacher
collected by the LeWM authors
10,000
2,010,000
2
stack
synthesized by MimicGen
20,000
2,196,344
7
coffee
synthesized by MimicGen
20,000
4,371,670
7
Appendix
Table 5: Tasks and training data . Total frames is the number of images over all time steps of all demonstrations in the training data. Numbers of demonstrations and total frames are measured on the training data; action dimensions follow the environment definitions.
Task
Tolerance of the criterion
Baseline [%]
Δ of auxiliary-loss models [%]
False match [%]
stack (object ∧ hand)
both
hand only
object only
object 1 cm (hand 2 cm)
16.6
+2.45∗
−0.05
−3.95∗
—
object 2 cm (hand 2 cm)
24.3
+4.20∗
−1.60
−1.45
—
object 3 cm (hand 2 cm)
28.1
+4.75∗
−1.90
−0.60
—
object 4 cm (hand 2 cm) †
30.1
+5.05∗
−1.30
−0.40
0.15
hand 1 cm (object 4 cm)
0.7
+0.15
+0.10
+0.40∗
—
Appendix
Table 6: Validity of the success criteria . Baseline success rate and effect of the auxiliary-loss models as the tolerance of the stack and coffee criteria is varied (identical 2000 trial pairs; ∗ : p<0.05 ), and the false-match rate, the fraction of states from other episodes that are judged a success, shown only for the tolerance used in the main text ( † ). The auxiliary-loss models are those of Table 2 .
Figure 3: Examples of the trial strata (five tasks) . For each task, the start frame and the goal frame of the demonstration for a reach-only trial and a manipulation trial. The strata are those of Table 3 ; a trial whose object displacement is at least the object tolerance (4 cm for stack, coffee, and cube; 20 units for PushT) is a manipulation trial. Solid circles mark the object position and dashed circles the position at the start; the number next to the arrow is the displacement from start to goal. In reach-only trials the dashed circle coincides with the solid one. The circle in PushT marks the object position, not the outline of the T. Reacher has no object, all of its trials are reach-only, and no marks are drawn. Because the criterion of cube has no hand term, its reach-only trials satisfy the criterion at the start.
Model
Hand
Carried cube
Base cube
baseline
3.15
1.82
0.96
supervised by the hand only ( λ=1 )
0.14
1.05
3.49
supervised by the object only ( λ=0.1 )
1.96
0.29
0.18
Appendix
Table 7: Readout error from the latent state (stack) . 5000 frames not used for training, linear ridge regression, median of the 3D distance [cm]. Bold marks the supervised quantity. Tolerances are 2 cm for the hand and 4 cm for the cubes. The error of a constant prediction is 8.93 for the hand, 7.56 for the carried cube, and 6.96 cm for the base cube.
Figure 4: Ranking of candidates, shape of the search, and success rate (stack) . The horizontal axis orders the baseline and three auxiliary-loss models by the weight on the hand (the model supervised with both quantities has object λ=0.1 ). (a) Demo top-5 rate: the fraction of trials in which the demonstration action sequence ranks in the top five by cost among itself and 63 noisy copies (400 trials). (b) Search spread: the standard deviation of the actions among the top 30 candidates of the final CEM iteration in the actual planning; the smaller, the more focused the search (first plan of 2000 trials). (c) Difference in success rate from the baseline [%] (2000 trials). Intervals are 95%; filled markers mark a difference from the baseline whose interval excludes 0 (in (c), two-sided McNemar p<0.05 ). (a) and (b) follow the weight on the hand; (c) does not.
Supervised quantity
Official criterion
Agent component
Block component
baseline
91.5%
23.5%
93.3%
hand only
93.5(+2.0∗)
28.5(+5.0∗)
93.4(+0.1)
object only
91.6(+0.2)
24.1(+0.6)
93.6(+0.3)
both
95.0(+3.5∗)
26.1(+2.6∗)
94.8(+1.5∗)
Appendix
Table 8: Success rate of PushT by component of the criterion . Identical 2000 trial pairs. The improvements stated in the main text are the numbers in the official-criterion column; the component columns show which component changed. We set the component thresholds to agent <10 , block <15 , and angle <20∘ .
Stratum
Model
n
Hand error
Object error
Hand inside tolerance
Object inside tolerance
Manipulation
baseline
794
5.07
9.83
20.8%
26.7%
hand only
794
3.19
11.07
37.0%
14.9%
Reach-only
baseline
382
4.36
0.09
0.8%
100%
object only
382
5.05
0.04
0.0%
100%
Appendix
Table 9: End states of the trials that both models fail (stack) . Only the two combinations in which the complementary degradation occurs are shown: the hand-only model on the manipulation trials, and the object-only model on the reach-only trials. The end state is at step 50. Errors are medians of the 3D distance [cm]. Tolerances are 2 cm for the hand and 4 cm for the object, where the object error is the larger error of the two cubes. “Inside tolerance” is the fraction of trials in which the quantity entered its tolerance at some point along the path. Strata are as in Table 3 .
Figure 5: Effect by trial stratum for eight auxiliary-loss models on stack . Difference in success rate from the baseline [%] on identical 2000 trial pairs; the strata have 980 reach-only and 1020 manipulation trials. Includes the hand-only ( λ=1 ) and the both ( λ=1 , seed 2) models of Table 3 . Seeds 1–3 are three training seeds, the same three for every arm. Bars are 95% intervals from a paired bootstrap over trials; filled markers are significant by two-sided paired McNemar ( p<0.05 ). The criterion is as in Table 3 .
Figure 6: Demanded object displacement and the effect of the auxiliary loss (stack). Four auxiliary-loss models on stack, namely hand only with λ=0.03 , object only with λ=0.1 , both with λ=(1,0.1) , and both with λ=(0.03,0.1) and seed 2, evaluated on the same 2000 trials binned by displacement. The horizontal axis is the bin of the start-to-goal displacement of the object, the larger of the two cubes. The vertical axis is the difference in success rate from the baseline [%]; bars are 95% intervals from a paired bootstrap over trials, and filled markers are significant by two-sided paired McNemar ( p<0.05 ). The five bins contain 645, 335, 236, 449, and 335 trials.
Plans with latent WM
Auxiliary loss
Supervised target
Target fixed by the success criterion
Applied to predicted latent
Inference unchanged
Per quantity and stratum
LeWM, DINO-WM, V-JEPA 2-AC
✓
×
—
—
—
—
—
RC-aux
✓
✓
reachability
×
×
×
×
PhyLatent
✓
✓
full simulator state
×
✓
✓
×
PSG-JEPA
×
✓
proprioceptive state
×
×
✓
×
Ours
✓
✓
only success-criterion quantities
✓
✓
✓
✓
Appendix
Table 10: Position relative to the closest work . “Auxiliary loss” is a supervised loss added to the main objective of latent prediction; the rollout loss of V-JEPA 2-AC is not counted because its target is the latent state itself. “Applied to predicted latent” is whether the auxiliary loss shapes the latent state output by the predictor, which is what the planning cost scores: PhyLatent and this work apply the loss to both the encoder and the predictor outputs, PSG-JEPA only to the encoder output, and RC-aux stops the gradient on the predictor side. “Per quantity and stratum” is whether the effect is measured separately for each supervised quantity and for each trial stratum. “—” means the design element is absent. PhyLatent and PSG-JEPA are concurrent work released in August 2026. The planner of PSG-JEPA is an amortized inverse-dynamics model that uses neither the predictor nor the latent-distance cost, so it is not counted as planning with a latent world model.
Latent world models have become increasingly popular as a method to predict and plan in latent space rather than pixel space. Recent architectures, such as LeWorldModel (LeWM), jointly train the encoder and predictor using regularization techniques like SIGReg to prevent representation collapse. Even with such regularization preventing representation collapse, we identify a new world model failure mode of physical representation laziness, particularly noted in highly dynamic environments. For these lazy cases, the learned latent states do not collapse but nonetheless fail to represent key physical properties, causing ubiquitous downstream planning failure. To resolve this issue, we propose training-time auxiliary supervision with a lightweight "Fourier auxiliary head", which enforces physically-informed structuring of the latent space with no additional inference-time cost and can be generalized to any environment. Experimentally, we show that the auxiliary head substantially improves planning success rates in dynamic environments where the baseline LeWM exhibits physical representation laziness. It also leads to modest improvements in other environments, even when the baseline does not exhibit physical representation laziness. We further observe superior planning performance being accompanied by higher latent space correlations with key physical properties, indicating both the ability of our method to physically structure latent states and the potential planning-side benefit to the learned representation being physically structured. We also see in low-data regimes, auxiliary supervision is particularly impactful in increasing success rate. These findings support the use of our Fourier auxiliary head method to improve both overall success rate and data efficiency, while avoiding representation laziness in latent world models.
A latent world model may achieve accurate short-horizon prediction while still inducing a latent space that is poorly aligned with planning. A key issue is spatiotemporal mismatch: these models are often trained with local predictive supervision, but deployed for long-horizon goal-directed search in latent spaces where Euclidean distance may not reflect what is reachable within a finite action budget. We present the Reachability-Correction auxiliary objective (RC-aux), a lightweight correction for this mismatch in reconstruction-free latent world models. RC-aux keeps the world-model backbone unchanged and adds planning-aligned supervision along two axes. Along the time axis, multi-horizon open-loop prediction trains the model beyond one-step consistency. Along the space axis, budget-conditioned reachability supervision, together with temporal hard negatives, encourages the latent space to distinguish states that are eventually reachable from those reachable within the current planning horizon. At test time, the learned reachability signal can also be used by a reachability-aware planner to favor trajectories that are both goal-directed and attainable under the available budget. We instantiate RC-aux on LeWorldModel and evaluate it under both continuation-training and matched-from-scratch settings. Across goal-conditioned pixel-control tasks and a LIBERO-Goal extension, RC-aux improves LeWM-style planning with modest additional cost. These results suggest that planning with latent world models depends not only on predictive accuracy, but also on whether the learned representation encodes the temporal and geometric structure required by downstream search. The code is available at https://github.com/Guang000/RC-aux.
Latent world models are judged by how well they predict, so when planning fails at long horizons the natural reading is that the predictor degrades. On a reproduction of LeWorldModel on TwoRoom we show the binding constraint is the planner's objective instead. The predictor is not the limit: its imagined state seventy-five environment steps ahead is still only 0.189 as wrong as assuming the world froze, while the planner never imagines beyond twenty-five. The objective is. Cross-entropy-method planning minimises squared latent distance, which tracks true distance at r = 0.426, saturates by about eighty arena units and decreases beyond a hundred and twenty, so moving away from the goal can lower the cost. The information is present throughout: a ridge probe recovers position from the frozen embedding at R^2 0.9922. The pathology is the method's, not one reimplementation's. It is present in the authors' released weights, and across four checkpoints long-horizon success rank-orders exactly with metric quality and inversely with prediction accuracy. Replacing only the objective, with nothing retrained and no GPU, lifts goals reached at offset 100 from 26.0% to 98.0%, equals the 98.0% at offset 25, and reaches 92.0% under a third of the budget: planning stops depending on the horizon. The best cost is not the most accurate. A head learned from frame separation alone predicts spatial distance worse than a position probe (r = 0.819 against 0.9897) yet plans better, charging 24% more to cross the environment's dividing wall where squared latent distance charges 4% less. It has learned reachability, not proximity.