UniWAM: Unified World-Action Model
Abstract
Vision-language-action models benefit from the understanding and reasoning capabilities of pretrained vision-language models, but action-only supervision provides limited grounding in world dynamics. Conversely, world-action models inherit spatiotemporal priors from video generation models, yet remain limited in semantic understanding and reasoning under distribution shifts. We introduce UniWAM, a unified architecture that integrates a physical reasoner, a world generator, and an action predictor to jointly learn semantic understanding of the physical world, visual generation, and action prediction. To ensure the quality of the training data, we developed a rigorous data cleaning and annotation pipeline for both human egocentric data and robot data. To adapt the vision-language component to embodied tasks while preserving its inherited language capabilities, we represent low-level actions in natural language and introduce a pre-training recipe that assigns complementary supervision from visual question answering (VQA) data, human egocentric data, and robot demonstrations to the appropriate model components. During post-training, future visual noise augmentation reduces reliance on precise future predictions, while history-conditioned flow matching uses encoded action history to initialize action generation. Together, these designs significantly reduce denoising steps while maintaining performance. UniWAM achieves state-of-the-art (SOTA) performance across multiple evaluations, including in-distribution performance, robustness, generalization, instruction following, and long-horizon task execution. Furthermore, we uncover a log-linear scaling law of unified human-robot co-training, demonstrating the effectiveness of large-scale pre-training on a mixture of human and robot data.
Figures & tables
| Data Type | Embodiment Type | Data Sources | Duration (hours) |
| Robot | Dual-arm/Mobile | AgiBot World 2026 | 891 |
| Dual-arm/Mobile | AgiBot World Alpha | 595 | |
| Single-arm | Bridge | 80 | |
| Single-arm | Droid | 365 | |
| Single-arm | Fractal | 340 | |
| Dual-arm | Robocoin | 1088 |
| QA Type | Data Source | QA Pairs |
| Spatial Understanding | SAT | 172K |
| RefSpatial | 1.430M | |
| VST-P | 563K | |
| SenseNova-SI | 800K | |
| GRiD-3D | 358K | |
| Grounding | RoboPoint | 930K |
| Stage | Input / Configuration | Model |
| Global segmentation | Full-episode timestamped contact sheets; s sampling, up to 20 frames per sheet, and completed-event rules | Qwen3.6-27B |
| Local refinement | Local timestamped contact sheets, coarse hypothesis, and episode-level context; no additional window padding | Qwen3.6-27B |
| Segment annotation | Fixed temporal segments; raw-frame, FFmpeg-based, and seed- or prior-conditioned annotation passes | Qwen3.5-397B |
| Candidate selection | Candidate description texts only; no video frames | Qwen3.5-397B |
| Method | Spatial | Object | Goal | Long | Average |
| ( Black et al., 2025b ) | 98.0 | 96.8 | 94.4 | 88.4 | 94.4 |
| PD-VLA ( Song et al., 2025 ) | 95.5 | 96.7 | 94.9 | 91.7 | 94.7 |
| ( Black et al., 2025a ) | 98.8 | 98.2 | 98.0 | 92.4 | 96.9 |
| GR00T-N1.7 ( NVIDIA, 2026 ) | 97.7 | 98.5 | 97.5 | 94.4 | 97.0 |
| OpenVLA-OFT ( Kim et al., 2025a ) | 97.6 | 98.4 | 97.9 | 94.5 | 97.1 |
| Fast-WAM ( Yuan et al., 2026 ) | 98.2 | 100.0 | 97.0 | 95.2 | 97.6 |
| Method | C2C | C2R | Average |
| VLA | |||
| GR00T-N1.7 ( NVIDIA, 2026 ) | 43.6 | 20.7 | 32.2 |
| StarVLA ( StarVLA Community, 2026 ) | 58.1 | 10.6 | 34.4 |
| Xiaomi Robotics-0 ( Cai et al., 2026b ) | 62.90 | 18.20 | 40.55 |
| Abot-M0 ( Yang et al., 2026c ) | 57.40 | 30.36 | 43.88 |
| X-VLA ( Zheng et al., 2026a ) | 68.00 | 20.90 | 44.45 |
| Method | Camera | Robot | Language | Light | Background | Noise | Layout | Total |
| ( Black et al., 2025b ) | 79.6 | 21.1 | 72.5 | 84.7 | 86.2 | 68.3 | 69.4 | 67.4 |
| GR00T-N1.6 ( NVIDIA GEAR Team, 2025 ) | 92.6 | 33.5 | 80.1 | 93.6 | 95.4 | 93.6 | 75.0 | 79.4 |
| OpenVLA-OFT ( Kim et al., 2025a ) | 92.8 | 30.3 | 85.8 | 94.9 | 93.9 | 89.3 | 77.6 | 79.5 |
| Spatial Forcing ( Li et al., 2026a ) | 95.2 | 47.9 | 73.5 | 91.2 | 95.6 | 92.2 | 74.8 | 80.5 |
| MemoryVLA ( Shi et al., 2026 ) | 91.4 | 48.6 | 79.4 | 95.2 | 95.3 | 94.0 | 75.7 | 81.9 |
| ( Black et al., 2025a ) | 87.6 | 78.4 | 80.0 | 92.6 | 91.6 | 91.4 | 81.8 | 86.2 |
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
| Task | C2R | C2C | Average | Task | C2R | C2C | Average |
| adjust_bottle | 72 | 72 | 72.00 | place_can_basket | 35 | 36 | 35.50 |
| beat_block_hammer | 56 | 65 | 60.50 | place_cans_plasticbox | 86 | 97 | 91.50 |
| blocks_ranking_rgb | 67 | 64 | 65.50 | place_container_plate | 93 | 97 | 95.00 |
| blocks_ranking_size | 38 | 36 | 37.00 | place_dual_shoes | 77 | 87 | 82.00 |
| click_alarmclock | 92 | 100 | 96.00 | place_empty_cup | 94 | 98 | 96.00 |
| click_bell | 99 | 100 | 99.50 | place_fan | 69 | 76 | 72.50 |