OpenWAM: An Open Framework for Composable World-Action Models
Organizations: Stanford University
Abstract
World-action models (WAMs) couple future prediction with robot control, yet existing systems often vary the video backbone, interaction structure, supervision, and inference procedure simultaneously, making their design choices difficult to compare. We introduce OPENWAM, an open world-action modeling framework built around a common causal robot-video foundation and configurable video-action interaction. Starting from Wan2.2-5B, we perform causal robot-video pretraining on over 10,000 hours of video, then integrate an action expert through a shared Mixture-of-Transformers architecture that supports joint, video-then-action, action-then-video, and decoupled generation. OPENWAM achieves high success rates on four LIBERO suites and real-world bimanual tasks; robot-video training with causal adaptation improves VTA success on LIBERO-Long from 68.4% to 97.8%. The same configurable architecture naturally extends to inverse and forward dynamics, allowing us to study how counterfactual transitions improve independently trained dynamics models beyond demonstrations alone. When only the video predictor is adapted to a new task, a frozen local-context inverse dynamics model trained on counterfactual data and demonstrations achieves 84.0% mean success across four held-out LIBERO-90 tasks, compared with 47.0% for a full-context inverse model and 21.5% for a local-context model trained only on demonstrations. For forward dynamics, counterfactual supervision reduces RGB prediction error by 34.5% and raises outcome identification from 21.1% to 71.3% among 16 same-state outcomes. OPENWAM provides a common testbed for comparing WAM interaction designs and for studying dynamics learning from video data beyond successful demonstrations.
Figures & tables
| Task | Target behavior | Transfer challenge |
|---|---|---|
| 64 | Stack bowls; place stack in tray | Novel skill composition |
| 74 | Book into caddy’s left compartment | Retarget familiar manipulation |
| 21 | Turn on stove; place frying pan on it | Object and grasp change |
| 45 | Same objective as Task 21 | Additional scene and layout shift |
| Task | Objective | Main challenge |
|---|---|---|
| Toast Bread | Insert bread and activate toaster | Deformability and precise insertion |
| Rubik’s Cube | Restore final layer of cube | Contact-rich push and rotation |
| Sort Cups | Sort cups according to color | Variable ordering and sequencing |
| Method | Object | Goal | Spatial | Long | Mean |
|---|---|---|---|---|---|
| OpenVLA [ 38 ] | 88.4 | 79.2 | 84.7 | 53.7 | 76.5 |
| OpenVLA-OFT [ 36 ] | 98.4 | 97.9 | 97.6 | 94.5 | 97.1 |
| [ 6 ] | 98.8 | 95.8 | 96.8 | 85.2 | 94.1 |
| [ 33 ] | 98.2 | 98.0 | 98.8 | 92.4 | 96.9 |
| GR00T-N1 [ 5 ] | 97.6 | 93.0 | 94.4 | 90.6 | 93.9 |
| Motus [ 3 ] | 99.8 | 96.6 | 96.8 | 97.6 | 97.7 |
| Method | Toast | Cube | Cups | Mean |
|---|---|---|---|---|
| OpenWAM -VTA | 92.0 | 90.0 | 94.4 | 92.1 |
| OpenWAM -Joint | 90.0 | 94.0 | 91.7 | 91.9 |
| Video / policy | Action component (frozen) | 64 Composition | 74 Retarget | 21 Object/grasp | 45 Scene shift | Transfer mean |
| Task-tuned VTA (finetuned policy reference) | 94% | 86% | 92% | 100% | 93.0% | |
| Task-tuned video | Full-context IDM, demo-only | 88% | 90% | 0% | 0% | 44.5% |
| Task-tuned video | Full-context IDM, mixed | 54% | 92% | 4% | 38% | 47.0% |
| Task-tuned video | Local-context IDM, demo-only | 36% | 50% | 0% | 0% | 21.5% |
| Task-tuned video | Local-context IDM, CF-only | 80% | 92% | 90% | 76% | 84.5% |
| Task-tuned video | Local-context IDM, mixed | 74% | 86% | 88% | 88% | 84.0% |
| Action component | Supervision | Success (%) |
| Full-context IDM | Demo-only (Native VTA) | 97.8 |
| Full-context IDM | Mixed | 92.2 |
| Local-context IDM | Demo-only | 25.8 |
| Local-context IDM | CF-only | 90.6 |
| Local-context IDM | Mixed | 94.4 |
| Supervision | RGB MSE | Acc. | Acc. |
|---|---|---|---|
| Demo-only | 14.35 | 68.3 | 21.1 |
| CF-only | 9.40 | 93.6 | 71.3 |
| Mixed | 9.62 | 91.8 | 67.7 |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Trajectories / takes | Hours | Source type |
|---|---|---|---|
| Open X-Embodiment (49 manipulation datasets) | 1,300,749 | 1,911.29 | Multi-embodiment robot data |
| AgiBot World Beta | 1,003,672 | 2,976.4 | Real robot manipulation |
| Ego-Exo4D v2 | 5,035 | 221.26 | Human interaction |
| InternData-A1 | 637,498 | 7,433.91 | Synthetic robot manipulation |
| RoboCOIN | 183,157 | 1,306.83 | Real bimanual manipulation |
| RoboMIND | 107,877 | 305.5 | Multi-embodiment robot data |
| Quantity | Demonstrations | LIBERO-Long-CF |
|---|---|---|
| Tasks | 10 | 10 |
| Stored sequences | 500 | 32,000 |
| Sequences / task | 50 | 3,200 |
| Controls / sequence | 276.2 mean | 128 |
| Total controls | 138,090 | 4,096,000 |
| Duration | 1.92 h | 56.9 h |
| Intervention family | Fraction (%) |
|---|---|
| Stop / rescale arm motion | 9.4 |
| Reverse / redirect translation | 6.3 |
| Axis biases and pulses | 12.5 |
| Dedicated yaw perturbation | 3.1 |
| Noise / randomized arm controls | 12.5 |
| Dedicated gripper interventions | 31.3 |
| Measured interaction | Rate (%) |
|---|---|
| Gripper–object / fixture contact | 90.4 |
| Detected grasp | 50.0 |
| Object-configuration effect | 72.3 |
| Video initialization | VTA (%) | Joint (%) |
|---|---|---|
| Random initialization | 20.0 | 27.2 |
| Original Wan2.2 | 68.4 | 62.2 |
| Robot-video pretrained | 97.8 | 96.6 |
| Architecture | VTA (%) | Joint (%) |
|---|---|---|
| Non-MoT | 92.8 | 93.6 |
| MoT | 97.8 | 96.6 |
| Policy | Real demonstrations | Counterfactual transitions | |||||||
| Model | Succ. (%) | FDM MSE | FDM SSIM | IDM pos. (cm) | IDM rot. ( ∘ ) | FDM MSE | FDM SSIM | IDM pos. (cm) | IDM rot. ( ∘ ) |
| OpenWAM , LC-IDM, CF-only | – | – | – | 1.76 | 2.39 | – | – | 1.17 | 3.60 |
| OpenWAM , LC-IDM, mixed | – | – | – | 0.44 | 1.21 | – | – | 1.22 | 3.68 |
| OpenWAM , LC-FDM, CF-only | – | 0.00734 | 0.9126 | – | – | 0.00940 | 0.8889 | – | – |
| OpenWAM , LC-FDM, mixed | – | 0.00186 | 0.9776 | – | – | 0.00962 | 0.8859 | – | – |
| OpenWAM , multi-objective | 92.8 | 0.00058 | 0.9934 | 0.43 | 1.10 | 0.01416 | 0.8398 | 1.73 | 3.55 |