World Action Models (WAMs) emerge as a promising foundation for policy refinement by predicting the consequences of sampled actions. However, visually plausible predictions can mislead policy refinement if they fail to reflect the input actions. To address this mismatch, we introduce R2-WAM, a two-stage repair-and-reject post-training framework that first improves the consistency of predicted futures with input actions, then uses these futures to select inferior action samples for negative fine-tuning. The repair stage grounds imagination in observed robot behavior through a kinematic alignment score that measures agreement between predicted and demonstrated motion, enabling the predicted video to faithfully reflect its input actions. Using the repaired video model, the rejection stage compares imagined outcomes of sampled and demonstrated actions, selectively applying negative fine-tuning to samples whose predicted task progress falls below the demonstrated reference by a prescribed margin. Together, the two stages extend video prediction from representation learning to consequence-based policy refinement without additional environment interaction or changes to the inference procedure. R2-WAM achieves 93.8% average success on RoboTwin 2.0 across clean and randomized settings. On the long-horizon real-world Fold Shirt task, it achieves 87.5% average success, compared with 0% for Fast-WAM. Our project page is available at https://r2-wam.github.io/.
Figures & tables
Figure 1: R2 -WAM extends video prediction from representation learning to consequence feedback for policy post-training. (a) Average success on RoboTwin 2.0. (b) Action negative-selection precision and recall against execution-based STEAM labels. (c) Real-world success under seen and unseen conditions. Bottom: execution sequences on Fold Shirt and Clean Table.
Figure 2: Overview of R2 -WAM . Predictive repair compares generated robot motion with observed motion under the same demonstrated action. The repaired video expert is then frozen and predicts consequences for selecting inferior action samples.
Method
Clean
Randomized
Average
Vision language action policies
π0 ( Black et al., 2024 )
65.9
58.4
62.2
X-VLA ( Zheng et al., 2026a )
72.9
72.8
72.9
π0.5 ( Intelligence et al., 2025b )
82.7
76.8
79.8
ABot-M0 ( Yang et al., 2026c )
81.2
80.4
80.8
LingBot-VLA ( Wu et al., 2026 )
86.5
85.3
85.9
Table 1: Closed-loop success rates (%) on RoboTwin 2.0.
Precision
Recall
Margin
Continued FM
Repair
Continued FM
Repair
0.02
76.9
79.6
44.4
63.7
0.05
63.2
76.8
38.7
56.7
0.08
64.6
69.8
35.5
50.6
Table 2: Action negative-selection precision and recall (%) at three margins. Reference labels are obtained by applying each margin to STEAM score gaps from actual executions; predicted selections use the same margin on imagined consequences.
Table 5
Figure 3: Real-world evaluation and qualitative analysis. (a–b) Success rates on Fold Shirt and Clean Table. (c) Fast-WAM executions and imagined consequences of actions from Fast-WAM and R2 -WAM during (i) bimanual folding, (ii) garment repositioning, and (iii) lateral folding. Human intervention in (c) enables later-stage analysis; these rollouts are excluded from success rates.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter
Simulation
Real robot
Kv
4
4
mv
0.02
0.02
τv
0.01
0.01
βv
0.1
0.1
λv
0.1
0.1
rmaxv
0.25
0.25
Appendix
Table 5: Post-training hyperparameters in simulation and on the real robot.
Predictive post-training
PSNR
SSIM
Tracking success (%)
AF failure (%)
Continued flow matching
30.69
0.9500
93.0
32.0
Repair with PSNR
31.94
0.9576
96.8
18.6
Ours w/o relative-trajectory consistency
32.25
0.9584
97.2
17.0
Ours w/o absolute-position consistency
32.52
0.9607
98.0
10.6
Ours
33.17
0.9642
99.2
6.8
Appendix
Table 6: Evaluation and component analysis of predictive repair on 500 randomized validation episodes from RoboTwin 2.0, with one generated future per episode. Action-following metrics are evaluated on episodes whose observed future provides a reliable reference trajectory.
Predictive repair score
Clean
Randomized
PSNR
92.1
91.8
Kinematic alignment (ours)
93.8
93.7
Appendix
Table 7: Policy success rates (%) with different predictive repair scores and the same action rejection procedure on RoboTwin 2.0.
Margin
Video model
TP
FP
FN
TN
Precision
Recall
F1
0.02
Continued FM
652
196
816
336
76.9
44.4
56.3
Predictive repair
935
239
533
293
79.6
63.7
70.8
0.05
Continued FM
437
254
692
617
63.2
38.7
48.0
Predictive repair
640
193
489
678
76.8
56.7
65.2
0.08
Continued FM
268
147
487
1,098
64.6
35.5
45.8
Predictive repair
382
165
373
1,080
69.8
50.6
58.7
Appendix
Table 8: Action negative selection against reference labels from actual executions. Counts are pooled over 2,000 shared candidate actions; precision, recall, and F1 are percentages. The positive class is an action whose execution-based consequence gap exceeds the listed margin.
Figure 4: Real-world experimental setup, showing the dual-arm PIPER platform, cameras, and task objects for Fold Shirt and Clean Table.
Task
Episodes
Frames
Duration (h)
Fold Shirt
605
3,110,127
28.80
Clean Table
2,426
1,526,405
14.13
Total
3,031
4,636,532
42.93
Appendix
Table 9: Offline demonstrations for the two evaluated real-world tasks.
Figure 5: Representative evaluation configurations for Fold Shirt (top) and Clean Table (bottom). Blue and pink outlines indicate training and test position regions, respectively.
Method
Seen
Unseen position
Unseen object
Unseen position + object
π0.5
65.0 [43.3, 81.9]
35.0 [18.1, 56.7]
50.0 [29.9, 70.1]
35.0 [18.1, 56.7]
X-VLA
40.0 [21.9, 61.3]
15.0 [5.2, 36.0]
20.0 [8.1, 41.6]
10.0 [2.8, 30.1]
Motus
0.0 [0.0, 16.1]
0.0 [0.0, 16.1]
0.0 [0.0, 16.1]
0.0 [0.0, 16.1]
Fast-WAM
0.0 [0.0, 16.1]
0.0 [0.0, 16.1]
0.0 [0.0, 16.1]
0.0 [0.0, 16.1]
R2 -WAM
90.0 [69.9, 97.2]
90.0 [69.9, 97.2]
85.0 [64.0, 94.8]
85.0 [64.0, 94.8]
Appendix
Table 10: Fold Shirt success rates (%) for each real-world condition. Each condition contains 20 autonomous trials. Success rates appear above their 95% Wilson score confidence intervals.
Method
Seen
Unseen position
Unseen object
Unseen position + object
π0.5
70.0 [48.1, 85.5]
45.0 [25.8, 65.8]
60.0 [38.7, 78.1]
40.0 [21.9, 61.3]
X-VLA
50.0 [29.9, 70.1]
20.0 [8.1, 41.6]
40.0 [21.9, 61.3]
20.0 [8.1, 41.6]
Motus
45.0 [25.8, 65.8]
35.0 [18.1, 56.7]
45.0 [25.8, 65.8]
20.0 [8.1, 41.6]
Fast-WAM
55.0 [34.2, 74.2]
40.0 [21.9, 61.3]
35.0 [18.1, 56.7]
25.0 [11.2, 46.9]
R2 -WAM
85.0 [64.0, 94.8]
75.0 [53.1, 88.8]
75.0 [53.1, 88.8]
70.0 [48.1, 85.5]
Appendix
Table 11: Clean Table success rates (%) for each real-world condition. Each condition contains 20 autonomous trials. Success rates appear above their 95% Wilson score confidence intervals.
Video
Action
Average success (%)
Repair
GT FM
0.0
GT FM
Rejection
0.0
Repair
Rejection
87.5
Appendix
Table 12: Component ablation on Fold Shirt. Success rates are averaged over the four evaluation conditions. GT FM denotes flow matching on demonstrations.
Figure 6: Execution sequences of R2 -WAM on Fold Shirt. Each sequence spans two consecutive rows and proceeds from left to right.
Figure 7: Execution sequences of R2 -WAM on Fold Shirt (top four rows) and Clean Table (bottom two rows). Each Fold Shirt sequence spans two rows, while each Clean Table sequence occupies one row. Frames proceed from left to right.