World models are typically trained and evaluated by prediction error, assuming that more accurate predictions lead to better decisions. We show that this assumption can fail because models with similar total error can differ substantially in planning performance when their errors occur on different state dimensions. We introduce Decision-Relevant Prediction Error (DRPE), which measures prediction error on the state dimensions that affect decisions. We also develop an iso-error evaluation protocol that varies error allocation while keeping total error fixed. In a factored gridworld with known state relevance and a standardized planner, we evaluate 55 controlled and learned models across different error levels and allocations. Total prediction error is weakly related to planning success (Spearman ρ=−0.25), whereas DRPE is strongly predictive (ρ=−0.84; −0.98 within the controlled family). Models with only a 1% difference in total error can differ by 60 percentage points in planning success (97% vs 37%). The relevant error also depends on the task, with model rankings reversing across tasks at the same total error. Deeper imagination further amplifies decision-relevant errors, while learned models exhibit systematic bias on rare but decision-critical events. We formalize sufficient conditions under which DRPE correctly ranks models and total prediction error cannot.
Figures & tables
Figure 1: Environment illustration of the controlled laboratory. Gray cells are walls; the shaded column holds a single door; the key (diamond) and agent (circle) are placed at random. Lines show true closed-loop trajectories, shaded from start to end; the plus marker denotes key pickup, the cross the final cell. Under irrelevant noise (b), behavior matches the oracle (a). Under comparable noise on relevant dimensions (c), the agent never picks up the key and times out at a total error close to model (b), which succeeds 97% of the time (Table 2 ).
Figure 2: The decoupling, established causally. Each panel shows the (σR,σI) plane of the controlled model family, where σR injects prediction noise on decision-relevant dimensions and σI on irrelevant ones. (a) Planning success depends almost only on σR . (b) Total error Etot depends on both. (c) Iso- Etot contours overlaid on success. Along each contour, success varies by up to 60 points at fixed total error.
Figure 3: The same two models, two tasks, one world. In Task A both c1 and c2 are irrelevant and both noise models are harmless. In Task B, c1 selects the active goal and the other goal becomes a terminal hazard.
Figure 4: Planning success against total error (left) and DRPE (right) for the 48 controlled grid models ( 8×6 sweep over (σR,σI) ), colored by σR . Against total error the cloud is uninformative because the most accurate models include both the best and almost worst planners. Against DRPE, success is close to deterministic.
Model pool
n
Total Error
DRPE
All models
55
−0.25
−0.84
Controlled grid
48
−0.32 [ −0.56,−0.04 ]
−0.98 [ −1.00,−0.94 ]
Relevance sweep, graded
24
−0.24 [ −0.62,0.19 ]
−0.99 [ −1.00,−0.94 ]
Relevance sweep, binary
24
−0.24
−0.99
Table 1: Spearman correlation between prediction error and closed-loop planning success. DRPE consistently outperforms total prediction error, with the strongest relationship in the controlled relevance sweep (a 24-model depth-4 sweep, σR∈{0,0.1,0.2,0.25,0.3,0.35} , σI∈{0,0.3,0.6,1.2} , scored with graded and with binary weights). Brackets denote bootstrap 95% CIs. For the iso-error pairs, the absolute change in DRPE strongly mirrors the absolute change in planning success ( ρ=0.97 , 77 pairs).
Model
σR
σI
Etot
DRPE
success
oracle M(0,0)
0
0
0.020
0.000
0.97±0.01
M(0,0.3) (irrelevant error)
0
0.3
0.151
0.000
0.97±0.01
M(0.35,0.1) (relevant error)
0.35
0.1
0.152
0.234
0.37±0.02
Table 2: Total error does not determine decision quality. The oracle and σI -only model both plan perfectly. The σR model with almost matched total error (rows 2 and 3 differ by 1%) loses 60 points of success.
Figure 5: Closed-loop trajectories in Task B with the same start state. Stars mark the two goal cells; the cross marks the terminal cell. (a) With noise on the irrelevant bit c2 , the model reaches the active goal. (b) With matched noise on c1 , the model misreads which goal is active and walks into the hazard.
Figure 6: Planner depth and horizon change what an error is worth. (a) Planning success vs. planner depth (depth factorial, 5 seeds × 20 episodes). Irrelevant noise is harmless at every depth, while relevant error hurts more the deeper the planner imagines. (b) Error compounding over rollout horizon H . Irrelevant error grows without bound (to 1.38 ) with zero decision impact, while the relevant-noise model’s DRPE grows to 0.28 . Round markers show Etot while square markers show DRPE.
Variant
data
success
Etot
DRPE
uniform
20k
0.75±0.03
0.145
0.076
event-weighted (pickup/door)
20k
0.79±0.07
0.128
0.043
anti-weighted (relevant dims down-weighted)
20k
0.40±0.04
0.168
0.105
rel-heavy (relevant dims up-weighted)
20k
0.35±0.10
0.111
0.020
rel-only (irrelevant dims dropped)
20k
0.04±0.04
0.102
0.011
uniform
12k
0.70±0.05
0.355
0.383
Table 3: Learned world models. Does total error or DRPE rank them? No, and the failures are informative. The 12k data model has the worst total error ( 0.355 ) yet plans well ( 0.70 ). The rel-only model has nearly the lowest total error ( 0.102 ) yet fails completely ( 0.04 ). Bold marks the best success and underline the second best. Success is mean ± SE over 5 training seeds × 40 episodes.
Figure 7: Imagined rollouts versus reality in the decision-relevant projection. From one start state and one shared action sequence, each panel shows 12 imagined rollouts (colored) and the true trajectory (black) over H=10 steps, plotted in the agent-position plane that the planner reads. (a) Irrelevant-noise model: imagined positions coincide with the truth. Its large total error lives entirely on dimensions the planner never reads. (b) Relevant-noise model: imagined positions fan out within a few steps, and the value heuristic is evaluated at the wrong cells.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Mask
dims
ρ (DRPE, success)
all 6 relevant (exact)
6
−0.98
FN drop agent- x (top weight)
5
−0.98
FN drop 3 relevant dims
3
−0.98
FN carrying flag only
1
−0.98
FP add static bit c1
7
−0.45
FP add c1,c2
8
−0.39
Appendix
Table 4: Rank correlation between DRPE and success for n=48 controlled models under perturbed relevance masks. Here FN marks a mask that removes relevant dimensions and FP marks one that adds irrelevant dimensions.
Dimension
role
weight wj
agent x
navigation, reward reachability
0.29
agent y
navigation, reward reachability
0.21
key x
fetch target of the plan
0.15
key y
fetch target of the plan
0.11
carrying
switches the plan target
0.15
door-open
gates access to the goal chamber
0.10
Appendix
Table 5: Ground truth relevance weights (normalized).
Figure 8: The failure of the learned model failure is invisible to averages. (a) Closed-loop trajectory of the model trained on uniform-random data, showing 60 steps of oscillation near the wall column with the key untouched and the goal unreached. (b) Per-dimension MSE of the same model against a controlled model with noise on irrelevant dimensions only. The learned model’s errors are small on every dimension. The failure is the event-level one of Section 5.4 , not an excess in the average profile.
σR
σI
Etot
DRPE
success
σR
σI
Etot
DRPE
success
0.00
0.00
0.020
0.000
0.97±0.01
0.20
0.00
0.059
0.077
0.85±0.02
0.00
0.10
0.035
0.000
0.97±0.01
0.20
0.10
0.074
0.077
0.85±0.02
0.00
0.20
0.080
0.000
0.97±0.01
0.20
0.20
0.119
0.077
0.85±0.02
0.00
0.30
0.151
0.000
0.97±0.01
0.20
0.30
0.189
0.077
0.85±0.02
0.00
0.60
0.503
0.000
0.97±0.01
0.20
0.60
0.541
0.077
0.85±0.02
0.00
1.20
1.864
0.000
0.97±0.01
0.20
1.20
1.903
0.077
0.85±0.02
Appendix
Table 6: Full iso-error grid listing total error Etot , DRPE and planning success (mean ± SE over 10 seeds × 30 episodes) for every (σR,σI) cell. σR in grid units, σI in normalized units.
Figure 9: Planning success against total multi-step prediction error for the 48 controlled models of the grid on a log scale colored by σR . At matched Etot , both the best and the worst planners appear. Only the color, which shows where each model’s error falls, separates the two groups.
Dimension
oracle
irr. noise
rel. noise
learned (rand.)
learned (bal.)
ax
0.0000
0.0000
0.0147
0.0044
0.0044
ay
0.0000
0.0000
0.0131
0.0060
0.0055
kx
0.0000
0.0000
0.0129
0.0029
0.0055
ky
0.0000
0.0000
0.0142
0.0008
0.0040
carry
0.0000
0.0000
0.6471
0.0700
0.0799
door
0.0000
0.0000
0.7053
0.0047
0.0238
Appendix
Table 7: Per-dimension rollout MSE averaged over H=10 steps and 10 seeds × 60 start states, for five representative models. The controlled models are the oracle M(0,0) , the irrelevant-noise model M(0,0.6) and the relevant-noise model M(0.35,0) . The learned models are the 12 k random-data and 20 k balanced-data MLPs of Section 5.4 .
Figure 10: Four failing episodes of the c1 -noise model in Task B , labelled by evaluation seed and episode. A green star marks the active goal and a red cross marks the hazard. In panels (a) to (c) the corrupted belief about c1 commits the model to the inactive goal, so it walks into the hazard, while in panel (d) it never commits and circles until the horizon expires. Across the 8×30 -episode pool, 35 episodes end in the hazard and 8 time out, aggregating to the 0.82 success rate of Section 5.3 .
Figure 11: Imagined rollouts versus the true trajectory for four model families . Same start state, action sequence, and H=10 horizon as Figure 7 . K=12 imagined rollouts per model (colored), true trajectory in black, key as diamond. Divergence from the true path grows with σR and is zero for the oracle and the irrelevant-noise model.
model A
model B
Δ DRPE
σRA
σIA
EtotA
σRB
σIB
EtotB
success gap
0.234
0.00
0.30
0.151
0.35
0.00
0.138
0.60
0.234
0.00
0.30
0.151
0.35
0.10
0.152
0.60
0.234
0.00
1.20
1.864
0.35
1.20
1.982
0.60
0.171
0.00
0.30
0.151
0.30
0.20
0.166
0.51
0.171
0.00
1.20
1.864
0.30
1.20
1.950
0.51
Appendix
Table 8: The eight iso-error pairs with the largest success gaps, from the 77 controlled pairs matched to 10% in total error. Model A carries noise on irrelevant dimensions only, and model B carries relevant-dimension noise.