Humanoid World Action Model With Joint State--Action Generation
Organizations: LimX Dynamics · Southern University of Science and Technology
Abstract
Humanoid robots are a promising platform for general-purpose manipulation. Recent Vision-Language-Action (VLA) policies learn actions directly from multimodal observations, while World Action Models (WAMs) further incorporate future visual prediction to improve action generation. However, in hierarchical humanoid systems, VLA and WAM policies output reference actions that are subsequently realized through whole-body control, robot dynamics, balance, and contact. This hierarchy creates an action--execution gap: the reference produced by the policy can differ from the motion realized by the robot. Without explicitly modeling the realized body state, future visual prediction must jointly explain scene evolution and discrepancies between reference actions and executed motion, making it difficult to associate an action with its physical outcome. We propose HWAM, a Humanoid World Action Model with joint state--action generation, which makes the robot's post-execution proprioceptive state an explicit prediction target. By jointly generating reference actions and their realized body states, HWAM directly incorporates supervision of executed motion into action learning. HWAM is trained through three complementary conditional paths. The Policy path jointly denoises state--action trajectories conditioned only on current observations, matching deployment conditions. Forward Dynamics Modeling (FDM) predicts future visual observations conditioned on actions and post-execution states, while Inverse Dynamics Modeling (IDM) reconstructs the joint trajectory from visual transitions. Together, these paths connect policy references, realized body motion, and visual outcomes. HWAM achieves the highest success rate among evaluated baselines on three real-robot tasks on the LimX OLI humanoid. On Candy Picking, HWAM achieves a 70.6% success rate, compared with 43.3% for Fast-WAM.
Figures & tables
| Method | Robot-data pretraining | Candy Picking | Object Collection | Plush-Toy Picking |
|---|---|---|---|---|
| Cosmos3-Edge | Yes | 15.8 | 0.0 | 56.7 |
| GR00T N1.5 | Yes | 40.5 | 42.5 | 60.0 |
| Yes | 62.5 | 45.0 | 60.7 | |
| Fast-WAM | No | 43.3 | 0.0 | 60.0 |
| DiT4DiT | No | 52.5 | 15.0 | — |
| HWAM (Ours) | No | 70.6 | 46.7 | 73.3 |
| Configuration | Prediction target | Training recipe | Success (%) |
|---|---|---|---|
| Action + VGM/Policy | Action | VGM/Policy | 60.0 |
| State–Action + VGM/Policy | Joint state–action | VGM/Policy | 45.5 |
| Action + FDM/IDM/Policy | Action | FDM/IDM/Policy | 55.0 |
| HWAM | Joint state–action | FDM/IDM/Policy | 73.3 |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Platform | Direct MAE | Steady-state MAE |
|---|---|---|
| LimX OLI (mobile humanoid) | 4.420 | 2.885 |
| ALOHA (stationary bimanual) | 0.648 | 0.402 |
| Task | Control demands | Training episodes | Completion criterion |
|---|---|---|---|
| Candy Picking | Precise grasping and spatially targeted placement | 1,783 | Pick the instructed candy and place it in its assigned tray region. |
| Object Collection | Repeated reaching and grasping across objects | 4,220 | Place all four tabletop objects in the basket. |
| Plush-Toy Picking | Mobile object handling with whole-body coordination | 404 | Complete the specified plush-toy handling sequence. |
| Configuration | PSNR | SSIM |
|---|---|---|
| VGM + Action | 20.690 | 0.7420 |
| VGM + State–Action | 20.673 | 0.7409 |
| Joint denoising + Action | 20.844 | 0.7453 |
| Joint denoising + State–Action | 20.670 | 0.7440 |
| FDM/IDM/Policy + Action | 20.345 | 0.7361 |
| HWAM | 20.501 | 0.7370 |
| Model | Condition | PSNR |
|---|---|---|
| Action-only | Ground-truth actions | 20.559 |
| HWAM | Ground-truth states and actions | 23.336 |
| Method | Successes / trials | SR (%) | 95% CI (%) |
|---|---|---|---|
| GR00T N1.5 | 49/121 | 40.5 | [32.2, 49.4] |
| 75/120 | 62.5 | [53.6, 70.6] | |
| Fast-WAM | 52/120 | 43.3 | [34.8, 52.3] |
| HWAM | 84/119 | 70.6 | [61.9, 78.0] |
| (a) Motion directions | ||||
|---|---|---|---|---|
| Method | CH | CV | X-P | X-S |
| GR00T N1.5 | 50.0 | 25.8 | 40.0 | 46.7 |
| 60.0 | 66.7 | 80.0 | 43.3 | |
| Fast-WAM | 33.3 | 43.3 | 46.7 | 50.0 |
| HWAM | 82.8 | 63.3 | 76.7 | 60.0 |
| Method | Clothes Folding score |
|---|---|
| GR00T N1.5 | 78.75 |
| GR00T N1.7 | 58.75 |
| 65.00 | |
| 77.00 | |
| Fast-WAM | 66.25 |
| SmolVLA | 51.25 |