World models are typically trained and evaluated by prediction error, assuming that more accurate predictions lead to better decisions. We show that this assumption can fail because models with similar total error can differ substantially in planning performance when their errors occur on different state dimensions. We introduce Decision-Relevant Prediction Error (DRPE), which measures prediction error on the state dimensions that affect decisions. We also develop an iso-error evaluation protocol that varies error allocation while keeping total error fixed. In a factored gridworld with known state relevance and a standardized planner, we evaluate 55 controlled and learned models across different error levels and allocations. Total prediction error is weakly related to planning success (Spearman ρ=−0.25), whereas DRPE is strongly predictive (ρ=−0.84; −0.98 within the controlled family). Models with only a 1% difference in total error can differ by 60 percentage points in planning success (97% vs 37%). The relevant error also depends on the task, with model rankings reversing across tasks at the same total error. Deeper imagination further amplifies decision-relevant errors, while learned models exhibit systematic bias on rare but decision-critical events. We formalize sufficient conditions under which DRPE correctly ranks models and total prediction error cannot.
Figures & tables
Figure 1: Environment illustration of the controlled laboratory. Gray cells are walls; the shaded column holds a single door; the key (diamond) and agent (circle) are placed at random. Lines show true closed-loop trajectories, shaded from start to end; the plus marker denotes key pickup, the cross the final cell. Under irrelevant noise (b), behavior matches the oracle (a). Under comparable noise on relevant dimensions (c), the agent never picks up the key and times out at a total error close to model (b), which succeeds 97% of the time (Table 2 ).
Figure 2: The decoupling, established causally. Each panel shows the (σR,σI) plane of the controlled model family, where σR injects prediction noise on decision-relevant dimensions and σI on irrelevant ones. (a) Planning success depends almost only on σR . (b) Total error Etot depends on both. (c) Iso- Etot contours overlaid on success. Along each contour, success varies by up to 60 points at fixed total error.
Figure 3: The same two models, two tasks, one world. In Task A both c1 and c2 are irrelevant and both noise models are harmless. In Task B, c1 selects the active goal and the other goal becomes a terminal hazard.
Figure 4: Planning success against total error (left) and DRPE (right) for the 48 controlled grid models ( 8×6 sweep over (σR,σI) ), colored by σR . Against total error the cloud is uninformative because the most accurate models include both the best and almost worst planners. Against DRPE, success is close to deterministic.
Model pool
n
Total Error
DRPE
All models
55
−0.25
−0.84
Controlled grid
48
−0.32 [ −0.56,−0.04 ]
−0.98 [ −1.00,−0.94 ]
Relevance sweep, graded
24
−0.24 [ −0.62,0.19 ]
−0.99 [ −1.00,−0.94 ]
Relevance sweep, binary
24
−0.24
−0.99
Table 1: Spearman correlation between prediction error and closed-loop planning success. DRPE consistently outperforms total prediction error, with the strongest relationship in the controlled relevance sweep (a 24-model depth-4 sweep, σR∈{0,0.1,0.2,0.25,0.3,0.35} , σI∈{0,0.3,0.6,1.2} , scored with graded and with binary weights). Brackets denote bootstrap 95% CIs. For the iso-error pairs, the absolute change in DRPE strongly mirrors the absolute change in planning success ( ρ=0.97 , 77 pairs).
Model
σR
σI
Etot
DRPE
success
oracle M(0,0)
0
0
0.020
0.000
0.97±0.01
M(0,0.3) (irrelevant error)
0
0.3
0.151
0.000
0.97±0.01
M(0.35,0.1) (relevant error)
0.35
0.1
0.152
0.234
0.37±0.02
Table 2: Total error does not determine decision quality. The oracle and σI -only model both plan perfectly. The σR model with almost matched total error (rows 2 and 3 differ by 1%) loses 60 points of success.
Figure 5: Closed-loop trajectories in Task B with the same start state. Stars mark the two goal cells; the cross marks the terminal cell. (a) With noise on the irrelevant bit c2 , the model reaches the active goal. (b) With matched noise on c1 , the model misreads which goal is active and walks into the hazard.
Figure 6: Planner depth and horizon change what an error is worth. (a) Planning success vs. planner depth (depth factorial, 5 seeds × 20 episodes). Irrelevant noise is harmless at every depth, while relevant error hurts more the deeper the planner imagines. (b) Error compounding over rollout horizon H . Irrelevant error grows without bound (to 1.38 ) with zero decision impact, while the relevant-noise model’s DRPE grows to 0.28 . Round markers show Etot while square markers show DRPE.
Variant
data
success
Etot
DRPE
uniform
20k
0.75±0.03
0.145
0.076
event-weighted (pickup/door)
20k
0.79±0.07
0.128
0.043
anti-weighted (relevant dims down-weighted)
20k
0.40±0.04
0.168
0.105
rel-heavy (relevant dims up-weighted)
20k
0.35±0.10
0.111
0.020
rel-only (irrelevant dims dropped)
20k
0.04±0.04
0.102
0.011
uniform
12k
0.70±0.05
0.355
0.383
Table 3: Learned world models. Does total error or DRPE rank them? No, and the failures are informative. The 12k data model has the worst total error ( 0.355 ) yet plans well ( 0.70 ). The rel-only model has nearly the lowest total error ( 0.102 ) yet fails completely ( 0.04 ). Bold marks the best success and underline the second best. Success is mean ± SE over 5 training seeds × 40 episodes.
Figure 7: Imagined rollouts versus reality in the decision-relevant projection. From one start state and one shared action sequence, each panel shows 12 imagined rollouts (colored) and the true trajectory (black) over H=10 steps, plotted in the agent-position plane that the planner reads. (a) Irrelevant-noise model: imagined positions coincide with the truth. Its large total error lives entirely on dimensions the planner never reads. (b) Relevant-noise model: imagined positions fan out within a few steps, and the value heuristic is evaluated at the wrong cells.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Mask
dims
ρ (DRPE, success)
all 6 relevant (exact)
6
−0.98
FN drop agent- x (top weight)
5
−0.98
FN drop 3 relevant dims
3
−0.98
FN carrying flag only
1
−0.98
FP add static bit c1
7
−0.45
FP add c1,c2
8
−0.39
Appendix
Table 4: Rank correlation between DRPE and success for n=48 controlled models under perturbed relevance masks. Here FN marks a mask that removes relevant dimensions and FP marks one that adds irrelevant dimensions.
Dimension
role
weight wj
agent x
navigation, reward reachability
0.29
agent y
navigation, reward reachability
0.21
key x
fetch target of the plan
0.15
key y
fetch target of the plan
0.11
carrying
switches the plan target
0.15
door-open
gates access to the goal chamber
0.10
Appendix
Table 5: Ground truth relevance weights (normalized).
Figure 8: The failure of the learned model failure is invisible to averages. (a) Closed-loop trajectory of the model trained on uniform-random data, showing 60 steps of oscillation near the wall column with the key untouched and the goal unreached. (b) Per-dimension MSE of the same model against a controlled model with noise on irrelevant dimensions only. The learned model’s errors are small on every dimension. The failure is the event-level one of Section 5.4 , not an excess in the average profile.
σR
σI
Etot
DRPE
success
σR
σI
Etot
DRPE
success
0.00
0.00
0.020
0.000
0.97±0.01
0.20
0.00
0.059
0.077
0.85±0.02
0.00
0.10
0.035
0.000
0.97±0.01
0.20
0.10
0.074
0.077
0.85±0.02
0.00
0.20
0.080
0.000
0.97±0.01
0.20
0.20
0.119
0.077
0.85±0.02
0.00
0.30
0.151
0.000
0.97±0.01
0.20
0.30
0.189
0.077
0.85±0.02
0.00
0.60
0.503
0.000
0.97±0.01
0.20
0.60
0.541
0.077
0.85±0.02
0.00
1.20
1.864
0.000
0.97±0.01
0.20
1.20
1.903
0.077
0.85±0.02
Appendix
Table 6: Full iso-error grid listing total error Etot , DRPE and planning success (mean ± SE over 10 seeds × 30 episodes) for every (σR,σI) cell. σR in grid units, σI in normalized units.
Figure 9: Planning success against total multi-step prediction error for the 48 controlled models of the grid on a log scale colored by σR . At matched Etot , both the best and the worst planners appear. Only the color, which shows where each model’s error falls, separates the two groups.
Dimension
oracle
irr. noise
rel. noise
learned (rand.)
learned (bal.)
ax
0.0000
0.0000
0.0147
0.0044
0.0044
ay
0.0000
0.0000
0.0131
0.0060
0.0055
kx
0.0000
0.0000
0.0129
0.0029
0.0055
ky
0.0000
0.0000
0.0142
0.0008
0.0040
carry
0.0000
0.0000
0.6471
0.0700
0.0799
door
0.0000
0.0000
0.7053
0.0047
0.0238
Appendix
Table 7: Per-dimension rollout MSE averaged over H=10 steps and 10 seeds × 60 start states, for five representative models. The controlled models are the oracle M(0,0) , the irrelevant-noise model M(0,0.6) and the relevant-noise model M(0.35,0) . The learned models are the 12 k random-data and 20 k balanced-data MLPs of Section 5.4 .
Figure 10: Four failing episodes of the c1 -noise model in Task B , labelled by evaluation seed and episode. A green star marks the active goal and a red cross marks the hazard. In panels (a) to (c) the corrupted belief about c1 commits the model to the inactive goal, so it walks into the hazard, while in panel (d) it never commits and circles until the horizon expires. Across the 8×30 -episode pool, 35 episodes end in the hazard and 8 time out, aggregating to the 0.82 success rate of Section 5.3 .
Figure 11: Imagined rollouts versus the true trajectory for four model families . Same start state, action sequence, and H=10 horizon as Figure 7 . K=12 imagined rollouts per model (colored), true trajectory in black, key as diamond. Divergence from the true path grows with σR and is zero for the oracle and the irrelevant-noise model.
model A
model B
Δ DRPE
σRA
σIA
EtotA
σRB
σIB
EtotB
success gap
0.234
0.00
0.30
0.151
0.35
0.00
0.138
0.60
0.234
0.00
0.30
0.151
0.35
0.10
0.152
0.60
0.234
0.00
1.20
1.864
0.35
1.20
1.982
0.60
0.171
0.00
0.30
0.151
0.30
0.20
0.166
0.51
0.171
0.00
1.20
1.864
0.30
1.20
1.950
0.51
Appendix
Table 8: The eight iso-error pairs with the largest success gaps, from the 77 controlled pairs matched to 10% in total error. Model A carries noise on irrelevant dimensions only, and model B carries relevant-dimension noise.
Latent world models predict the consequences of actions, but accurate prediction does not guarantee that latent distance reflects which candidate will execute successfully. We identify a decision-local prediction gap: among the few futures competing for execution, a candidate predicted closer to the goal can produce a worse realized outcome than an available alternative. We introduce D-JEPA, a decision-aligned latent world model that learns decision-relevant relations among candidate futures from executed outcomes. A bounded, permutation-equivariant operator jointly reasons over goal-relative predictive features and ordinal evidence, refining pretrained predictive geometry where action choices are most consequential. Restricted predictor adaptation and a shared ordinal interface extend this alignment across complementary predictive geometries. D-JEPA further realizes the learned decision structure in JEPA-compatible future representations, enabling deployment through native latent-distance planning. Evaluations across latent control, manipulation, pretrained action-producing models, physical robots and autonomous driving demonstrate improved action selection, including 87.89% success on PushT, a 15.04-point average gain on RoboTwin, and a 17-point gain on physical robot tasks. These results establish decision-relevant relational structure as a direct bridge between predictive world modeling and effective control.
Shuaijun Liu, Chengyu Wu, Qifu Wen +5
The Hong Kong University of Science and Technology (Guangzhou) · Boston University · Shanghai Jiao Tong University
Neural world models coupled with model predictive control (MPC) replan at every environment step to bound accumulated prediction error, but this incurs substantial computational overhead. Reusing a cached plan reduces this overhead, yet its effectiveness depends on how prediction mismatch propagates through the local dynamics. We analyze this trade-off with a perturbation-based dynamic-regret framework and show that stale-plan penalties scale with the reuse tolerance, the accumulated mismatch since the last replanning step, and the local dynamics sensitivity. Based on this structure, we propose AdaReP, a training-free wrapper that adapts the replanning tolerance online using the current deviation from the cached rollout and a local sensitivity estimate, without modifying the learned world model or planner. Across image-space planning, latent-space control, and real-world robotic manipulation, AdaReP substantially reduces planner-side computation while maintaining comparable task performance, including over 80% fewer queries on a 50-trial physical robot study.
Yutian Cheng, Xiaojian Ma, Xianhao Wang +6
Shanghai Jiao Tong University · Beijing Institute for General Artificial Intelligence (BIGAI) · University of Science and Technology of China
Latent world models are trained to predict future states in a learned representation and are then deployed inside a planner that selects actions by simulating them forward. Current practice adopts the prediction error, the single- or multi-step rollout loss on held-out data, as the training and model-selection objective, on the assumption that a lower prediction error yields better control. We show that this assumption is unreliable for a structural reason: a planner does not query the model on the training distribution but on the states that its candidate actions reach, which generally leave the data manifold, so an error averaged over the data cannot by itself govern control. We therefore reframe the objective as the discrepancy between the predicted and the true plan-cost at the plan the planner commits to, and prove that the planner's suboptimality is bounded by twice this discrepancy, whereas the data-averaged prediction error neither bounds nor tracks it. Under a linear-control premise the discrepancy separates into two terms. The first is a small on-manifold residual, on which the predicted and true dynamics agree and which a spectral tax prices through the non-normality of the latent transition operator. The second is an off-manifold divergence, on which an action carries the state off the manifold and the two dynamics diverge; this divergence is the binding term and is bounded by no data-averaged error. Synthetic operators confirm the pricing formulas, and latent model-predictive control experiments confirm the decoupling: across seeds, the single-step validation error is essentially uncorrelated with control success, whereas a fidelity score on the planner-reachable measure tracks it.
Hanzhe You, Yonggang Zhang, Maohao Ran +6
University of Science and Technology of China · The Hong Kong University of Science and Technology, HKGAI · Hong Kong Baptist University