World Models Dream of Success: Diagnosing and Repairing Failure Insensitivity in Robot World Models
Authors: Jiuyi Xu, Xiao Hu, Meida Chen, Peng Gao, Yang Ye, Yangming Shi
Organizations: Colorado School of Mines · Northeastern University · Institute for Creative Technologies, University of Southern California · North Carolina State University
Robot world models support policy evaluation, planning, and synthetic data generation, but these applications require predictions that distinguish successful actions from failures. Across four released checkpoints from two architecture families, we observe weak sensitivity to action changes and success-like predictions on verified failures. Although recent work incorporates failures into model training, which data can repair released checkpoints without changing their architecture or training objective still remains underexplored. To this end, we introduce CureWM, which constructs alternative actions from successful demonstrations across a severity grid, verifies their outcomes through execution in simulation or on hardware, and fine-tunes released models on the resulting failures and surviving successes alongside nominal demonstrations. This construction provides controlled action contrasts from shared starting contexts. On 484 held-out LIBERO failure counterfactuals, optimism falls from 80% after fine-tuning on the official data to 30--43% across four independently fine-tuned CureWM models (38% mean). In two separate evaluations on a physical robot arm, failure predictions scored as success-like by a latent-distance diagnostic decrease from 90% after fine-tuning on successful demonstrations alone to 33% with CureWM. With failure counts per task, successful replay data, and training budget matched, counterfactual failures yield a success--failure value gap of 0.124, compared with 0.014 for freshly collected on-policy failures. These findings support execution-verified counterfactual replay for post-hoc repair and show why reduced optimism must be evaluated alongside success--failure discrimination. Code is available at https://github.com/jiuyixu25/CureWM.
Figures & tables
Figure 1: Failure insensitivity on a held-out Franka counterfactual. From the same starting context, the original action sequence succeeds and a perturbed sequence fails under physical execution, yet the model predicts a success-like future for the failing sequence.
Figure 2: The overview of CureWM. We perturb successful demonstrations at different severities, verify outcomes through execution, and use both successful and failing replays for post-training alongside the original training data. Severity values, outcome assignments, and predicted-value bars are used for illustration rather than measured results.
Suite
Model
Value optimism ↓
ΔSF↑
AUROC ↑
False-positive rate ↓
Task success ↑
Goal (484/1196)
Released
79.13
+0.002
0.497
30.5
98.4
Baseline
79.55
+0.003
0.495
31.0
96.4
CureWM
30.17
+0.282
0.743
38.0
94.8
Spatial (367/443)
Released
65.67
−0.014
0.518
36.6
98.4
Baseline
64.03
−0.011
0.522
36.3
95.6
CureWM
47.68
+0.081
0.629
35.0
96.2
Table 1: LIBERO prediction and robot task success. Baseline and CureWM use the same fine-tuning steps. Suite counts indicate failed/successful held-out replays. False-positive rate measures successful replays assigned a value at or below 0.5. Means weight the four suites equally.
Training data / setting
Steps
AUROC ↑
ΔSF↑
OpenDrawer
CloseDoubleDoor
Released
—
0.634
0.0170
28/30
30/30
Grasp replays
5k
0.674
0.0672
17/30
23/30
Grasp replays
20k
0.677
0.0342
26/30
29/30
Expanded replays
5k
0.839
0.0302
25/30
26/30
Expanded replays
20k
0.780
0.0193
25/30
28/30
Grasp replays, action loss masked
5k
0.662
0.0308
20/30
30/30
Table 2: RoboCasa prediction sensitivity and task success. All ΔSF values use the same 120 modified-action examples, whose modified actions are not re-executed. AUROC is held-out articulated failure detection on 225 rollouts (111 successes, 114 failures) scored at 75% of episode length. Task success is out of 30 episodes (three evaluation seeds, ten episodes each). Expanded replays add drawer and door tasks. Masking removes action-prediction loss on failed replays.
Full pool: 30 pairs
Held-out subset: 10 pairs
Model
Optimism
Mean s(a−)
Optimism
Mean s(a−)
Adapted
27/30
+0.080
9/10
+0.076
Control
26/30
+0.069
8/10
+0.071
CureWM
21/30
+0.016
6/10
+0.005
Table 3: Ctrl-World: full-pool and held-out evaluation.
Model
Optimism
Pair-ranking accuracy (%)
ΔSF
AUROC
First hardware evaluation
Adapted
13/15
80.0
+0.077
0.75
Baseline
15/15
66.7
+0.066
0.65
CureWM, model 1
6/15
100.0
+0.251
0.95
CureWM, model 2
6/15
100.0
+0.271
1.00
Second hardware evaluation
Table 4: Cup pick-and-place evaluation on Franka Research 3. AUROC uses both successful and failed reference outcomes.
Figure 3: Predicted futures for a failing cup-task action sequence. Columns show successive time steps. Rows show recorded execution, the scene-adapted model, and CureWM. CureWM predicts the observed release; the adapted model predicts continued grasping.
Training setting
Value optimism (%) ↓
Mean V(a−)
Mean V(a+)
ΔSF↑
Released model
79.1
0.601
0.603
+0.002
Matched failure-source comparison: 7,500 steps
Official-data baseline
79.3
0.594
0.597
+0.003
Fresh policy failures
66.9
0.545
0.559
+0.014
Matched counterfactual failures
46.9
0.463
0.587
+0.124
Training-data comparisons: 17,500 steps
Table 5: Failure sources and successful replays on LIBERO. All models are evaluated on the same 484 held-out failing replays and their successful counterparts. The 7,500-step comparison matches failure counts and successful training replays. The 17,500-step comparisons report the primary models. The two groups use different failure pools and training budgets. ΔSF is computed before rounding.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Steps
Training A
Training B
Held-out
Primary
17,500
16.5
17.3
30.2
Supplementary 1
20,000
28.7
29.9
43.0
Supplementary 2
20,000
18.3
20.5
37.6
Supplementary 3
17,500
21.7
32.3
40.9
Appendix
Table 6: Four independently fine-tuned CureWM models on LIBERO. Values are optimism percentages. Training groups A and B contain 115 and 127 failures, while held-out evaluation uses the same 484 failures for every model.
AUROC
False-positive rate
Perturbation
Failed
Successful
Baseline
CureWM
Baseline
CureWM
Approach overshoot
33
247
0.88
0.87
14.6
19.0
Carry slip
12
268
0.68
0.70
27.6
29.5
Contact oscillation
74
206
0.52
0.51
34.5
37.9
Insufficient grip
122
158
0.45
0.86
43.7
44.3
Premature release
140
140
0.57
0.78
42.1
42.9
Appendix
Table 7: LIBERO-Goal discrimination by perturbation type. AUROC and false-positive rates compare the fine-tuning baseline with CureWM on identical replays. False-positive rates are percentages of successful replays assigned a value at or below 0.5.
Perturbation
Model
0.4
0.6
0.8
1.0
Approach overshoot
Released
0.85
0.94
0.89
0.96
CureWM
0.85
0.95
0.88
0.95
Carry slip
Released
—
0.55
0.72
0.57
CureWM
—
0.56
0.75
0.57
Contact oscillation
Released
0.67
0.52
0.55
0.38
CureWM
0.66
0.52
0.56
0.39
Appendix
Table 8: AUROC at fixed perturbation type and severity. Each group contains 56 replays. Dashes indicate no failures. Severity 0.2 is omitted because every type contains at most two failures at that severity. Carry slip at severity 0.6 contains only two failures and is excluded from the 21-group mean. Means use values before rounding.
Excluded type
nA
nB
A: Released / models 1, 2
B: Released / models 1, 2
Carry slip
1
6
100 / 0, 0
16.7 / 16.7, 16.7
Contact oscillation
13
19
92.3 / 53.8, 69.2
57.9 / 57.9, 57.9
Insufficient grip
31
25
80.6 / 45.2, 19.4
60.0 / 24.0, 40.0
Premature release
36
31
88.9 / 41.7, 55.6
51.6 / 22.6, 22.6
Wrist tilt
24
33
83.3 / 75.0, 75.0
57.6 / 69.7, 69.7
Appendix
Table 9: Value optimism after excluding a perturbation type from counterfactual training. Each result lists the released model followed by two independently fine-tuned models.
Figure 4: Hardware tasks and executed outcomes. Each row shows the starting scene, a successful execution, and a failing execution. Top: cup pick-and-place, with failure caused by insufficient grip. Bottom: red-cube picking, with failure caused by grasping a distractor instead of the target cube. All images show physical executions, not model predictions.
Data
Training
Validation
Evaluation
Failed perturbed replays
90
2
15
Successful perturbed replays
24
4
5
Unperturbed successful trajectories
31
2
10
Appendix
Table 10: Cup-task data split. The 31 unperturbed training trajectories contain 28 demonstrations and three additional replays. Successful perturbed replays all use wrist tilt.
Executed action
Model
Human: successful
Diagnostic: success-like
Failing
Adapted
8/15
13/15
Failing
CureWM
0/15
6/15
Successful
Adapted
3/5
—
Successful
CureWM
5/5
—
Appendix
Table 11: Majority judgments from three raters. The diagnostic classifies predictions as success-like when s(a)>0 .
Repair setting
Failure-score difference
Cube-task repair
+0.025
Cup-task repair before cube adaptation
−0.007
Appendix
Table 12: Cube picking on 29 held-out pairs from ten scenes. Differences are baseline minus repaired-model mean s(a−) ; positive values indicate lower failure-action scores after repair.
Edwin · The Hong Kong University of Science and Technology (HKUST), ACCESS – AI Chip Center for Emerging Smart Systems, and Zhejiang University. · Zhejiang University. +1