World Models Dream of Success: Diagnosing and Repairing Failure Insensitivity in Robot World Models
Authors: Jiuyi Xu, Xiao Hu, Meida Chen, Peng Gao, Yang Ye, Yangming Shi
Organizations: Colorado School of Mines · Northeastern University · Institute for Creative Technologies, University of Southern California · North Carolina State University
Robot world models support policy evaluation, planning, and synthetic data generation, but these applications require predictions that distinguish successful actions from failures. Across four released checkpoints from two architecture families, we observe weak sensitivity to action changes and success-like predictions on verified failures. Although recent work incorporates failures into model training, which data can repair released checkpoints without changing their architecture or training objective still remains underexplored. To this end, we introduce CureWM, which constructs alternative actions from successful demonstrations across a severity grid, verifies their outcomes through execution in simulation or on hardware, and fine-tunes released models on the resulting failures and surviving successes alongside nominal demonstrations. This construction provides controlled action contrasts from shared starting contexts. On 484 held-out LIBERO failure counterfactuals, optimism falls from 80% after fine-tuning on the official data to 30--43% across four independently fine-tuned CureWM models (38% mean). In two separate evaluations on a physical robot arm, failure predictions scored as success-like by a latent-distance diagnostic decrease from 90% after fine-tuning on successful demonstrations alone to 33% with CureWM. With failure counts per task, successful replay data, and training budget matched, counterfactual failures yield a success--failure value gap of 0.124, compared with 0.014 for freshly collected on-policy failures. These findings support execution-verified counterfactual replay for post-hoc repair and show why reduced optimism must be evaluated alongside success--failure discrimination. Code is available at https://github.com/jiuyixu25/CureWM.
Figures & tables
Figure 1: Failure insensitivity on a held-out Franka counterfactual. From the same starting context, the original action sequence succeeds and a perturbed sequence fails under physical execution, yet the model predicts a success-like future for the failing sequence.
Figure 2: The overview of CureWM. We perturb successful demonstrations at different severities, verify outcomes through execution, and use both successful and failing replays for post-training alongside the original training data. Severity values, outcome assignments, and predicted-value bars are used for illustration rather than measured results.
Suite
Model
Value optimism ↓
ΔSF↑
AUROC ↑
False-positive rate ↓
Task success ↑
Goal (484/1196)
Released
79.13
+0.002
0.497
30.5
98.4
Baseline
79.55
+0.003
0.495
31.0
96.4
CureWM
30.17
+0.282
0.743
38.0
94.8
Spatial (367/443)
Released
65.67
−0.014
0.518
36.6
98.4
Baseline
64.03
−0.011
0.522
36.3
95.6
CureWM
47.68
+0.081
0.629
35.0
96.2
Table 1: LIBERO prediction and robot task success. Baseline and CureWM use the same fine-tuning steps. Suite counts indicate failed/successful held-out replays. False-positive rate measures successful replays assigned a value at or below 0.5. Means weight the four suites equally.
Training data / setting
Steps
AUROC ↑
ΔSF↑
OpenDrawer
CloseDoubleDoor
Released
—
0.634
0.0170
28/30
30/30
Grasp replays
5k
0.674
0.0672
17/30
23/30
Grasp replays
20k
0.677
0.0342
26/30
29/30
Expanded replays
5k
0.839
0.0302
25/30
26/30
Expanded replays
20k
0.780
0.0193
25/30
28/30
Grasp replays, action loss masked
5k
0.662
0.0308
20/30
30/30
Table 2: RoboCasa prediction sensitivity and task success. All ΔSF values use the same 120 modified-action examples, whose modified actions are not re-executed. AUROC is held-out articulated failure detection on 225 rollouts (111 successes, 114 failures) scored at 75% of episode length. Task success is out of 30 episodes (three evaluation seeds, ten episodes each). Expanded replays add drawer and door tasks. Masking removes action-prediction loss on failed replays.
Full pool: 30 pairs
Held-out subset: 10 pairs
Model
Optimism
Mean s(a−)
Optimism
Mean s(a−)
Adapted
27/30
+0.080
9/10
+0.076
Control
26/30
+0.069
8/10
+0.071
CureWM
21/30
+0.016
6/10
+0.005
Table 3: Ctrl-World: full-pool and held-out evaluation.
Model
Optimism
Pair-ranking accuracy (%)
ΔSF
AUROC
First hardware evaluation
Adapted
13/15
80.0
+0.077
0.75
Baseline
15/15
66.7
+0.066
0.65
CureWM, model 1
6/15
100.0
+0.251
0.95
CureWM, model 2
6/15
100.0
+0.271
1.00
Second hardware evaluation
Table 4: Cup pick-and-place evaluation on Franka Research 3. AUROC uses both successful and failed reference outcomes.
Figure 3: Predicted futures for a failing cup-task action sequence. Columns show successive time steps. Rows show recorded execution, the scene-adapted model, and CureWM. CureWM predicts the observed release; the adapted model predicts continued grasping.
Training setting
Value optimism (%) ↓
Mean V(a−)
Mean V(a+)
ΔSF↑
Released model
79.1
0.601
0.603
+0.002
Matched failure-source comparison: 7,500 steps
Official-data baseline
79.3
0.594
0.597
+0.003
Fresh policy failures
66.9
0.545
0.559
+0.014
Matched counterfactual failures
46.9
0.463
0.587
+0.124
Training-data comparisons: 17,500 steps
Table 5: Failure sources and successful replays on LIBERO. All models are evaluated on the same 484 held-out failing replays and their successful counterparts. The 7,500-step comparison matches failure counts and successful training replays. The 17,500-step comparisons report the primary models. The two groups use different failure pools and training budgets. ΔSF is computed before rounding.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Steps
Training A
Training B
Held-out
Primary
17,500
16.5
17.3
30.2
Supplementary 1
20,000
28.7
29.9
43.0
Supplementary 2
20,000
18.3
20.5
37.6
Supplementary 3
17,500
21.7
32.3
40.9
Appendix
Table 6: Four independently fine-tuned CureWM models on LIBERO. Values are optimism percentages. Training groups A and B contain 115 and 127 failures, while held-out evaluation uses the same 484 failures for every model.
AUROC
False-positive rate
Perturbation
Failed
Successful
Baseline
CureWM
Baseline
CureWM
Approach overshoot
33
247
0.88
0.87
14.6
19.0
Carry slip
12
268
0.68
0.70
27.6
29.5
Contact oscillation
74
206
0.52
0.51
34.5
37.9
Insufficient grip
122
158
0.45
0.86
43.7
44.3
Premature release
140
140
0.57
0.78
42.1
42.9
Appendix
Table 7: LIBERO-Goal discrimination by perturbation type. AUROC and false-positive rates compare the fine-tuning baseline with CureWM on identical replays. False-positive rates are percentages of successful replays assigned a value at or below 0.5.
Perturbation
Model
0.4
0.6
0.8
1.0
Approach overshoot
Released
0.85
0.94
0.89
0.96
CureWM
0.85
0.95
0.88
0.95
Carry slip
Released
—
0.55
0.72
0.57
CureWM
—
0.56
0.75
0.57
Contact oscillation
Released
0.67
0.52
0.55
0.38
CureWM
0.66
0.52
0.56
0.39
Appendix
Table 8: AUROC at fixed perturbation type and severity. Each group contains 56 replays. Dashes indicate no failures. Severity 0.2 is omitted because every type contains at most two failures at that severity. Carry slip at severity 0.6 contains only two failures and is excluded from the 21-group mean. Means use values before rounding.
Excluded type
nA
nB
A: Released / models 1, 2
B: Released / models 1, 2
Carry slip
1
6
100 / 0, 0
16.7 / 16.7, 16.7
Contact oscillation
13
19
92.3 / 53.8, 69.2
57.9 / 57.9, 57.9
Insufficient grip
31
25
80.6 / 45.2, 19.4
60.0 / 24.0, 40.0
Premature release
36
31
88.9 / 41.7, 55.6
51.6 / 22.6, 22.6
Wrist tilt
24
33
83.3 / 75.0, 75.0
57.6 / 69.7, 69.7
Appendix
Table 9: Value optimism after excluding a perturbation type from counterfactual training. Each result lists the released model followed by two independently fine-tuned models.
Figure 4: Hardware tasks and executed outcomes. Each row shows the starting scene, a successful execution, and a failing execution. Top: cup pick-and-place, with failure caused by insufficient grip. Bottom: red-cube picking, with failure caused by grasping a distractor instead of the target cube. All images show physical executions, not model predictions.
Data
Training
Validation
Evaluation
Failed perturbed replays
90
2
15
Successful perturbed replays
24
4
5
Unperturbed successful trajectories
31
2
10
Appendix
Table 10: Cup-task data split. The 31 unperturbed training trajectories contain 28 demonstrations and three additional replays. Successful perturbed replays all use wrist tilt.
Executed action
Model
Human: successful
Diagnostic: success-like
Failing
Adapted
8/15
13/15
Failing
CureWM
0/15
6/15
Successful
Adapted
3/5
—
Successful
CureWM
5/5
—
Appendix
Table 11: Majority judgments from three raters. The diagnostic classifies predictions as success-like when s(a)>0 .
Repair setting
Failure-score difference
Cube-task repair
+0.025
Cup-task repair before cube adaptation
−0.007
Appendix
Table 12: Cube picking on 29 held-out pairs from ten scenes. Differences are baseline minus repaired-model mean s(a−) ; positive values indicate lower failure-action scores after repair.
Post-training is essential for turning pretrained generalist robot policies into reliable task-specific controllers, but existing human-in-the-loop pipelines remain tied to physical execution: each correction requires robot time, scene setup, resets, and operator supervision in the real world. Meanwhile, action-conditioned world models have been studied mainly for imagination, synthetic data generation, and policy evaluation. We propose \textbf{Human-in-the-World-Model (Hi-WM)}, a post-training framework that uses a learned world model as a reusable corrective substrate for failure-targeted policy improvement. A policy is first rolled out in closed loop inside the world model; when the rollout becomes incorrect or failure-prone, a human intervenes directly in the model to provide short corrective actions. Hi-WM caches intermediate states and supports rollback and branching, allowing a single failure state to be reused for multiple corrective continuations and yielding dense supervision around behaviors that the base policy handles poorly. The resulting corrective trajectories are then added back to the training set for post-training. We evaluate Hi-WM on three real-world manipulation tasks spanning both rigid and deformable object interaction, and on two policy backbones. Hi-WM improves real-world success by 37.9 points on average over the base policy and by 19.0 points over a world-model closed-loop baseline, while world-model evaluation correlates strongly with real-world performance (r = 0.953). These results suggest that world models can serve not only as generators or evaluators, but also as effective corrective substrates for scalable robot post-training.
Yaxuan Li, Zhongyi Zhou, Yefei Chen +5
Current Robotics · Tsinghua University · Peking University +1
Action-conditioned world models are increasingly used as scalable simulators for robot learning, yet current evaluations provide limited evidence that their predictions are reliable under the actions they condition on. Existing benchmarks largely emphasize visual fidelity, leaving unclear whether predicted futures are physically plausible, faithful to commanded actions, and calibrated to failure when actions should not succeed. We introduce \textsc{MiraBench}, a hierarchical benchmark that defines \emph{action-conditioned reliability} as a core evaluation target for robotic world models. MiraBench decomposes this target into three progressively demanding levels: \emph{Physics Adherence}, which evaluates reference-free physical consistency; \emph{Action-Following Fidelity}, which measures whether predictions respect task-relevant action inputs; and \emph{Optimism Bias Detection}, which probes the tendency to predict successful outcomes under failure-inducing actions. To support this evaluation, we curate a human-annotated corpus with over 16,000 judgments across tasks, failure categories, and leading world models. We evaluate 12 representative model configurations spanning vector-conditioned robotic world models, text-conditioned generative world models, open-weight systems, closed-source systems, and multiple model scales. Across this broad model landscape, MiraBench reveals three central findings: visual fidelity is a poor proxy for action fidelity; increasing model scale does not reliably improve action following; and optimism bias is pervasive across current systems. By shifting evaluation from appearance to action-conditioned reliability, MiraBench provides a diagnostic foundation for assessing and improving robotic world models as faithful simulators.
Tianzhuo Yang, Zihan Shen, Zirui Mi +7
Institute for Artificial Intelligence, Peking University · 2Physis Lab
World-action models (WAMs) have emerged as a promising paradigm for robot manipulation by jointly modeling future visual dynamics and robot actions. However, existing WAMs are trained predominantly on successful trajectories, making them prone to failure when real-world execution diverges from the learned dynamics. This issue is amplified in autoregressive WAMs, where execution errors become part of the causal history and continue to influence subsequent predictions. To this end, we introduce \method{}, a training-free framework that reformulates failure recovery as \emph{test-time scaling over causal histories}. This formulation decomposes recovery into three coupled decisions: \emph{when} to revise the causal history, \emph{where} to recover a reliable history prefix, and \emph{which} history configuration best supports subsequent execution. Specifically, \method{} realizes these decisions through three stages: 1) \textbf{Progress-Aware Recovery Trigger} detects persistent non-progress and triggers recovery only when the current execution state permits intervention; 2) \textbf{History-Prefix Recovery} identifies the unreliable history suffix, retrieves a historical anchor matching the current physical state, and reconstructs the causal KV state from the retained prefix while conditioning on the latest real observation; and 3) \textbf{Hypothesis Verification} compares the future continuations induced by complete-history, recovered-prefix, and full-reset hypotheses, and commits the best-supported hypothesis. Experiments in both simulated and real-world manipulation settings demonstrate consistent improvements in task success, while ablations confirm the contribution of each recovery stage.
Lin Li, Long Chen, Kwunhang +10
Edwin · The Hong Kong University of Science and Technology (HKUST), ACCESS – AI Chip Center for Emerging Smart Systems, and Zhejiang University. · Zhejiang University. +1