World Models are appearing as the next major frontier in computer vision. However, their robustness is currently largely unexplored. We identify the phenomenon of hallucination in latent World Models: given a state and an action, the predicted next latent can decode to a scene that never occurs. Because the prediction is statistically ordinary and is fed back autoregressively by the model, the error is both silent and compounding. We study whether such latent hallucination can be detected, localised, and corrected at inference time, on a frozen self-supervised world model in the absence of ground-truth error labels. We introduce Masked Empirical-Bayes Neural Denoising (MEND), a single conditional score network trained by denoising score matching on real transitions, whose score field serves three roles: its magnitude detects hallucination, its per-token field localises it to specific image patches, and it defines an inference-time correction direction. On two navigation environments MEND detects hallucination with an AUROC of up to 0.80 without using actions, exceeding a single-Gaussian density baseline while also localising the error (per-token AUPRC up to 0.87) and correcting it, all from one score field. Our correction reliably reduces single-step latent error and improves predictions. We identify that a part of the error is tangent to the data manifold, hence, we focus on detection and localisation while highlighting promises of the correction.
Figures & tables
Fig. 1: Latent hallucination grows with imagination. Free-running a frozen world model on PointMaze : the top row is ground truth, the bottom row is the model’s imagined rollout decoded to images. Early on they agree, but by step 17 the imagined agent has drifted to a part of the maze it never visits. The prediction is still a plausible latent; nothing flags the error until the true future is known.
Fig. 2: MEND’s single-iteration pipeline. The world model predicts z^t+1 from (zt,at) and the detector scores it, D(z^t+1) . Two mitigation paths follow. Option 1 (inference-time correction) uses the localiser to select the suspicious tokens ( m ) and iteratively move them toward the valid manifold while freezing the confident tokens; Option 2 (data-driven repair) logs the flagged predictions for offline fine-tuning of the world model. Bottom: over K iterations the correction restores the hallucinated tokens (red) to trusted ones (grey).
Detection AUROC
Det.
Loc.
Environment
Ours
Gauss.
AUPRC
AUPRC
Wall
0.801
0.648
0.802
0.712
PointMaze
0.691
0.631
0.687
0.874
TABLE I: MEND ( Ours ) against a diagonal-Gaussian density reference [ 15 ] . Detection is AUROC; the last two columns are detection and localisation per-token AUPRC (random 0.50 ).
Fig. 3: Localising deep-rollout hallucinations on PointMaze (step 17 of a free-running rollout, where the imagined agent has drifted far from reality). Columns: decoded ground truth, decoded prediction, ground-truth per-token error, and the D1 detector score field (per-token AUPRC annotated). Even when the whole state is far off the manifold the detector still points at the tokens that are wrong. Cyan circles mark the nine most-erroneous tokens.
Detector ( Wall )
AUROC
Conditional score net (no action)
0.801
Directly conditioned score, +action (AdaLN)
0.799
Cross-attention, +action (15M)
0.782
Inverse model — D2 action factor
0.490
Directly conditioned score, − action
0.507
TABLE II: Ablation of detector variants on Wall . Variant architectures are detailed in the supplementary material.
Environment
Δ error ↓
% improved ↑
Wall
−6.4%
98.5
PointMaze
−3.0%
99.5
TABLE III: Single-step correction. Δ error is the relative change in latent error, so a negative value indicates improvement; arrows mark the better direction.
Depth
No corr.
MEND (D1) ↓
1
1.770
1.626 ( −8.2% )
5
2.370
2.237 ( −5.6% )
9
2.830
2.732 ( −3.5% )
TABLE IV: Per-step correction on Wall : mean per-token latent error versus rollout depth (lower is better). Parentheses show the reduction relative to No corr.
Modern generative world models render increasingly realistic action-controllable futures, yet they frequently hallucinate: rollouts remain visually fluent while drifting from the ground-truth dynamics. We hypothesize that hallucination concentrates in low-coverage regions of the state-action space, where lightweight data-centric signals can both detect it and guide mitigation. To test this, we introduce MMBench2, a 427-hour, 210-task dataset for visual world modeling with ground-truth actions, rewards, and live simulators, and train a 350M-parameter world model on it. We identify three distinct hallucination modes: perceptual, action-marginalized, and scene-diverging -- each anchored to a different stage of the pipeline, and develop three signals that accurately predict where the model will fail. To close coverage gaps at training time, we develop a coverage-aware sampling technique; to close them online, our hallucination predictors serve as curiosity rewards for targeted data collection, yielding a data-efficient finetuning recipe that adapts the pretrained world model to entirely unseen environments with as few as 50 real environment trajectories. Overall, our findings reveal that hallucination in world models is inherently a data coverage issue, and that the same signals used to detect it can also be used for mitigation. An interactive web version of our paper is available at https://www.nicklashansen.com/mmbench2
A latent world model may achieve accurate short-horizon prediction while still inducing a latent space that is poorly aligned with planning. A key issue is spatiotemporal mismatch: these models are often trained with local predictive supervision, but deployed for long-horizon goal-directed search in latent spaces where Euclidean distance may not reflect what is reachable within a finite action budget. We present the Reachability-Correction auxiliary objective (RC-aux), a lightweight correction for this mismatch in reconstruction-free latent world models. RC-aux keeps the world-model backbone unchanged and adds planning-aligned supervision along two axes. Along the time axis, multi-horizon open-loop prediction trains the model beyond one-step consistency. Along the space axis, budget-conditioned reachability supervision, together with temporal hard negatives, encourages the latent space to distinguish states that are eventually reachable from those reachable within the current planning horizon. At test time, the learned reachability signal can also be used by a reachability-aware planner to favor trajectories that are both goal-directed and attainable under the available budget. We instantiate RC-aux on LeWorldModel and evaluate it under both continuation-training and matched-from-scratch settings. Across goal-conditioned pixel-control tasks and a LIBERO-Goal extension, RC-aux improves LeWM-style planning with modest additional cost. These results suggest that planning with latent world models depends not only on predictive accuracy, but also on whether the learned representation encodes the temporal and geometric structure required by downstream search. The code is available at https://github.com/Guang000/RC-aux.
The component that dominates a latent world model's prediction error need not be the one whose repair most improves action selection. We show this by comparing action sequences from identical physical starts and separating endpoint error into a candidate-pool center and action-relative responses. Across four model families and four tasks, a confirmation pool of 256 new starts per task and 300 shared candidates per start shows that center error dominates MSE in 14/16 model-task cells. Yet in six of these cells, an oracle that corrects only the action-relative responses yields better physical rank correlation and top-30 elite quality than one that corrects only the center, while leaving more latent MSE (family-wise corrected intervals). The preference differs across the evaluated settings: a separate LeWorldModel (LeWM) study that executes oracle-selected actions favors center repair on PushT and on Reacher with a render-matched goal. Matched-candidate tests localize ordering loss: for LeWM, encoding realized endpoints raises physical Spearman from 0.464 to 0.975 on that Reacher setting and from 0.193 to 0.631 on PushT (64 starts per task), while Cube's encoded-goal cost remains uninformative. A 72-run objective study improves selected response diagnostics, while incremental closed-loop planning gains remain unconfirmed. These results separate error magnitude from the decision effects of oracle correction and motivate evaluating representation, prediction, and planning as separate stages.
Rui Min, Xianyao Li, Fang Xu +1
Department of Mechanical and Aerospace Engineering, University of Florida, Gainesville, FL 32611, USA · Department of Civil and Coastal Engineering, University of Florida, Gainesville, FL 32611, USA