Organizations: IIIS, Tsinghua University, Beijing, China. · Xiong’an Institute of Artificial Intelligence, Xiong’an, China. · The University of Melbourne, Melbourne, Australia.
Humanoid loco-manipulation demands coordinated body and hand behavior, while conventional robot pre-training data provide limited coverage of such whole-body motion. We present WB-WAM, a World Action Model that incorporates explicit whole-body action supervision into generative video pre-training. A shared physical action space integrates body, root, and dexterous hand annotations from heterogeneous sources, enabling joint video and action learning from 1880.2 hours of partially annotated video and motion data. The resulting priors are refined through PICO mid-training and adapted to robot tasks with auxiliary forward kinematics supervision. We construct WB-Datasets to support these stages with retargeted egocentric human demonstrations and robot trajectories, allowing task-aligned human motion to supplement limited robot data. Evaluations in simulation demonstrate strong whole-body task performance with 81.9% in HumanoidArena, while real-world experiments further validate WB-WAM with 84.0% mean success across five tasks. Moreover, task-aligned PICO mid-training improves downstream task performance while reducing the need for real-robot demonstrations. These results support heterogeneous whole-body pre-training and human motion transfer as a practical route to data-efficient humanoid loco-manipulation.
Figures & tables
Fig. 2: WB-WAM architecture and three-stage training framework. Top: Progressive training integrates heterogeneous supervision, retargeted PICO motion, and real-robot demonstrations. Bottom: Video and action experts jointly model visual dynamics and whole-body actions conditioned on vision, language, and proprioception. Body and root references are executed through SONIC, while hand references directly control the finger joints.
Fig. 3: Composition of 9-source dataset used for pre-training.
Fig. 4: Pre-training data curation with pose validity, motion continuity, collision, and video quality checks.
Fig. 5: Robot hardware and system.
Method
Football
DoubleDesk
P&PBox
OpenDoor
SitSofa
Boxing
VisNavi
Mean SR
Best baseline (per task)
DP
π0.5
DP
DP
DP
DP
FM
FM 45.5 ± 25.2%
ACT 47.6 ± 24.2%
45.0 ± 10.8%
43.3 ± 6.2%
75.0 ± 4.1%
85.0 ± 10.8%
78.3 ± 14.3%
76.7 ± 2.4%
38.3 ± 14.3%
π0.5 51.2 ± 24.8%
DP 60.0 ± 24.2%
WB-WAM
70.0 ± 8.2%
65.0 ± 4.1%
86.7 ± 2.4%
98.3 ± 2.4%
95.0 ± 4.1%
81.7 ± 2.4%
76.7 ± 4.7%
81.9 ± 12.3%
TABLE I: Comparison on HumanoidArena [ 40 ] . For each task, we report the strongest baseline from the original benchmark.
w/ locomotion
w/o locomotion
Method
Wipe the table
Close the curtain
Make the bed
Move the pillow
Tidy the cloth
Mean SR
Mean TP
A
C
S
A
C
S
A
C
S
C
S
C
S
ACT [ 41 ]
0/20
0/20
0/20
0/20
0/20
0/20
0/20
0/20
0/20
16/20
8/20
19/20
12/20
20%
27.5%
π0.5 [ 42 ]
13/20
12/20
12/20
11/20
10/20
6/20
11/20
10/20
10/20
19/20
8/20
16/20
16/20
52%
61.17%
GR00T N1.6 [ 43 ]
16/20
15/20
15/20
18/20
13/20
12/20
0/20
0/20
0/20
6/20
0/20
17/20
15/20
42%
48.67%
Fast-WAM [ 9 ]
0/20
0/20
0/20
8/20
3/20
2/20
0/20
0/20
0/20
4/20
0/20
9/20
4/20
6%
12.83%
TABLE II: Main comparison with 20 real-robot trials per task. WB-WAM achieves the highest mean SR and mean TP.
Fig. 6: Real-world experiments. WB-WAM enables the robot to perform a variety of whole-body manipulation tasks.
w/ locomotion
w/o locomotion
Training route
Robot [-1pt]Demos
Close the curtain
Push the cart
Wipe the table
Checkout
Mean SR
Mean TP
A
C
S
C
L
S
A
C
S
C
S
Direct post-training
30
13/20
12/20
12/20
13/20
11/20
11/20
12/20
11/20
11/20
6/20
5/20
48.75%
51.04%
Direct post-training
100
16/20
15/20
13/20
19/20
16/20
16/20
16/20
16/20
15/20
10/20
8/20
65%
70.42%
w/ PICO mid-training
30
19/20
17/20
16/20
16/20
17/20
16/20
19/20
18/20
18/20
12/20
9/20
73.75%
78.13%
TABLE III: Effect of PICO mid-training with 20 real-robot trials per condition.
Orange
Lemon
Apple
Method
C
S
C
S
C
S
WB-WAM
17/20
15/20
18/20
17/20
19/20
16/20
TABLE IV: Language-conditioned fruit manipulation.
Fig. 7: Video prediction and real-world execution under visual variation. The three central frames show a video rollout generated by the video expert under a modified visual condition. Real-robot cart pushing is shown in a familiar scene (left) and a novel scene absent from the dataset (right).
World Action Models (WAMs) offer a promising approach to general-purpose robot manipulation by jointly modeling visual dynamics and actions. However, most WAM studies focus on tabletop or arm-centric manipulation, while humanoid loco-manipulation remains less explored. To address this gap, we introduce WholeBodyWAM, which jointly predicts future visual dynamics, manipulation actions, and whole-body control intents for generalizable humanoid loco-manipulation. It preserves pre-trained world-action priors while grounding heterogeneous whole-body controller (WBC) semantics and coordinating whole-body behavior. Extensive experiments show that WholeBodyWAM achieves an overall simulation task success rate of 91.9%, with a 0.23 improvement in real-world out-of-distribution task progress and a 70% reduction in success-rate variance across WBCs relative to the respective baselines. These results suggest a path toward scalable humanoid whole-body intelligence by extending pre-trained world-action priors through structured WBC grounding and coordination, rather than relearning whole-body behavior from scratch. Project page: https://wholebodywam.github.io/.
Zhuo Li, Yiming Yao, Jim Tan +3
1The Chinese University of Hong Kong · 4Φ-Institute · 2The University of Hong Kong +1
World Action Models (WAMs) couple a video dynamics prior to the policy and have shown encouraging results on tabletop manipulation, but iterative denoising over high-dimensional video-action latents leaves them too slow for real-time humanoid loco-manipulation. The problem is compounded by the dominant hierarchical paradigm, in which a high-level manipulation policy controls only the upper body while a low-level controller tracks coarse base commands -- placing upper and lower body in inconsistent action spaces and reducing the legs to balance-preserving locomotion. We present MotionWAM, a real-time WAM that drives autonomous humanoid loco-manipulation from a single egocentric camera by conditioning the policy on the intermediate denoising features of a video world model. MotionWAM replaces the upper-lower split with a unified motion latent and predicts whole-body motion tokens that jointly cover locomotion, torso motion, height regulation, foot interaction, and hand manipulation in a single action space. A three-stage learning framework progressively adapts the video world model to egocentric visual dynamics and to the target humanoid embodiment. On nine real-world Unitree G1 tasks, MotionWAM runs in real time, substantially outperforms Vision-Language-Action (VLA) baselines fine-tuned on the same demonstrations by over 30% in overall success rate, and executes task-driven foot interaction that decoupled upper-lower policies cannot reach. Our results suggest that video-pretrained WAMs can be lifted from tabletop manipulation to coordinated, human-like whole-body humanoid control.
Jia Zheng, Teli Ma, Yudong Fan +3
Mondo Robotics 2 HKUST (GZ) · Mondo Robotics · HKUST (GZ) 3 HKUST
Humanoid whole-body manipulation requires coordinated whole-body dynamics, yet large-scale trajectories from a target robot are expensive to collect and difficult to scale. In contrast, whole-body motion from human and humanoid sources is abundantly available, although such data cannot be directly used as embodiment-specific robot actions. This work asks whether these scalable motion resources can instead provide a transferable predictive prior for humanoid world-action modeling. We introduce WholeBodyWAM, a humanoid world-action model that learns whole-body dynamics from large-scale heterogeneous motion before target-robot training. We curate UniMotion-4K, a motion corpus spanning more than 4K hours from human videos, native 3D motion datasets, and heterogeneous humanoid platforms, and canonicalize these diverse sources into a unified motion space. A language-conditioned Motion Expert is then pretrained to predict future whole-body motion without target-robot action supervision. During robot post-training, the pretrained Motion Expert is integrated with Video and Action Experts through asymmetric Mixture-of-Transformers (MoT) attention, enabling predictive scene dynamics and whole-body motion to jointly inform embodiment-specific action generation. Experiments show that WholeBodyWAM consistently benefits from increased motion-pretraining scale, improves future-motion prediction and downstream task performance, and transfers effectively to real-world humanoid manipulation. Moreover, the pretrained motion prior substantially improves data efficiency under limited target-robot demonstrations.
Bowei Zhang, Qiyao Zhang, Shuanghao Bai +8
Nankai University · Beijing Innovation Center of Humanoid Robotics · Beijing Institute of Technology +1