-WAM: Repair-and-Reject Post-Training for World Action Models
Organizations: The Chinese University of Hong Kong, Shenzhen · Manifold AI · Southeast University · Chang’an University
Abstract
World Action Models (WAMs) emerge as a promising foundation for policy refinement by predicting the consequences of sampled actions. However, visually plausible predictions can mislead policy refinement if they fail to reflect the input actions. To address this mismatch, we introduce -WAM, a two-stage repair-and-reject post-training framework that first improves the consistency of predicted futures with input actions, then uses these futures to select inferior action samples for negative fine-tuning. The repair stage grounds imagination in observed robot behavior through a kinematic alignment score that measures agreement between predicted and demonstrated motion, enabling the predicted video to faithfully reflect its input actions. Using the repaired video model, the rejection stage compares imagined outcomes of sampled and demonstrated actions, selectively applying negative fine-tuning to samples whose predicted task progress falls below the demonstrated reference by a prescribed margin. Together, the two stages extend video prediction from representation learning to consequence-based policy refinement without additional environment interaction or changes to the inference procedure. -WAM achieves 93.8% average success on RoboTwin 2.0 across clean and randomized settings. On the long-horizon real-world Fold Shirt task, it achieves 87.5% average success, compared with 0% for Fast-WAM. Our project page is available at https://r2-wam.github.io/.
Figures & tables
| Method | Clean | Randomized | Average |
|---|---|---|---|
| Vision language action policies | |||
| ( Black et al., 2024 ) | 65.9 | 58.4 | 62.2 |
| X-VLA ( Zheng et al., 2026a ) | 72.9 | 72.8 | 72.9 |
| ( Intelligence et al., 2025b ) | 82.7 | 76.8 | 79.8 |
| ABot-M0 ( Yang et al., 2026c ) | 81.2 | 80.4 | 80.8 |
| LingBot-VLA ( Wu et al., 2026 ) | 86.5 | 85.3 | 85.9 |
| Precision | Recall | |||
|---|---|---|---|---|
| Margin | Continued FM | Repair | Continued FM | Repair |
| 0.02 | 76.9 | 79.6 | 44.4 | 63.7 |
| 0.05 | 63.2 | 76.8 | 38.7 | 56.7 |
| 0.08 | 64.6 | 69.8 | 35.5 | 50.6 |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Parameter | Simulation | Real robot |
|---|---|---|
| 4 | 4 | |
| 0.02 | 0.02 | |
| 0.01 | 0.01 | |
| 0.1 | 0.1 | |
| 0.1 | 0.1 | |
| 0.25 | 0.25 |
| Predictive post-training | PSNR | SSIM | Tracking success (%) | AF failure (%) |
|---|---|---|---|---|
| Continued flow matching | 30.69 | 0.9500 | 93.0 | 32.0 |
| Repair with PSNR | 31.94 | 0.9576 | 96.8 | 18.6 |
| Ours w/o relative-trajectory consistency | 32.25 | 0.9584 | 97.2 | 17.0 |
| Ours w/o absolute-position consistency | 32.52 | 0.9607 | 98.0 | 10.6 |
| Ours | 33.17 | 0.9642 | 99.2 | 6.8 |
| Predictive repair score | Clean | Randomized |
|---|---|---|
| PSNR | 92.1 | 91.8 |
| Kinematic alignment (ours) | 93.8 | 93.7 |
| Margin | Video model | TP | FP | FN | TN | Precision | Recall | F1 |
|---|---|---|---|---|---|---|---|---|
| 0.02 | Continued FM | 652 | 196 | 816 | 336 | 76.9 | 44.4 | 56.3 |
| Predictive repair | 935 | 239 | 533 | 293 | 79.6 | 63.7 | 70.8 | |
| 0.05 | Continued FM | 437 | 254 | 692 | 617 | 63.2 | 38.7 | 48.0 |
| Predictive repair | 640 | 193 | 489 | 678 | 76.8 | 56.7 | 65.2 | |
| 0.08 | Continued FM | 268 | 147 | 487 | 1,098 | 64.6 | 35.5 | 45.8 |
| Predictive repair | 382 | 165 | 373 | 1,080 | 69.8 | 50.6 | 58.7 |
| Task | Episodes | Frames | Duration (h) |
|---|---|---|---|
| Fold Shirt | 605 | 3,110,127 | 28.80 |
| Clean Table | 2,426 | 1,526,405 | 14.13 |
| Total | 3,031 | 4,636,532 | 42.93 |
| Method | Seen | Unseen position | Unseen object | Unseen position + object |
|---|---|---|---|---|
| 65.0 [43.3, 81.9] | 35.0 [18.1, 56.7] | 50.0 [29.9, 70.1] | 35.0 [18.1, 56.7] | |
| X-VLA | 40.0 [21.9, 61.3] | 15.0 [5.2, 36.0] | 20.0 [8.1, 41.6] | 10.0 [2.8, 30.1] |
| Motus | 0.0 [0.0, 16.1] | 0.0 [0.0, 16.1] | 0.0 [0.0, 16.1] | 0.0 [0.0, 16.1] |
| Fast-WAM | 0.0 [0.0, 16.1] | 0.0 [0.0, 16.1] | 0.0 [0.0, 16.1] | 0.0 [0.0, 16.1] |
| -WAM | 90.0 [69.9, 97.2] | 90.0 [69.9, 97.2] | 85.0 [64.0, 94.8] | 85.0 [64.0, 94.8] |
| Method | Seen | Unseen position | Unseen object | Unseen position + object |
|---|---|---|---|---|
| 70.0 [48.1, 85.5] | 45.0 [25.8, 65.8] | 60.0 [38.7, 78.1] | 40.0 [21.9, 61.3] | |
| X-VLA | 50.0 [29.9, 70.1] | 20.0 [8.1, 41.6] | 40.0 [21.9, 61.3] | 20.0 [8.1, 41.6] |
| Motus | 45.0 [25.8, 65.8] | 35.0 [18.1, 56.7] | 45.0 [25.8, 65.8] | 20.0 [8.1, 41.6] |
| Fast-WAM | 55.0 [34.2, 74.2] | 40.0 [21.9, 61.3] | 35.0 [18.1, 56.7] | 25.0 [11.2, 46.9] |
| -WAM | 85.0 [64.0, 94.8] | 75.0 [53.1, 88.8] | 75.0 [53.1, 88.8] | 70.0 [48.1, 85.5] |
| Video | Action | Average success (%) |
|---|---|---|
| Repair | GT FM | 0.0 |
| GT FM | Rejection | 0.0 |
| Repair | Rejection | 87.5 |