Organizations: IIIS, Tsinghua University, Beijing, China. · Xiong’an Institute of Artificial Intelligence, Xiong’an, China. · The University of Melbourne, Melbourne, Australia.
Humanoid loco-manipulation demands coordinated body and hand behavior, while conventional robot pre-training data provide limited coverage of such whole-body motion. We present WB-WAM, a World Action Model that incorporates explicit whole-body action supervision into generative video pre-training. A shared physical action space integrates body, root, and dexterous hand annotations from heterogeneous sources, enabling joint video and action learning from 1880.2 hours of partially annotated video and motion data. The resulting priors are refined through PICO mid-training and adapted to robot tasks with auxiliary forward kinematics supervision. We construct WB-Datasets to support these stages with retargeted egocentric human demonstrations and robot trajectories, allowing task-aligned human motion to supplement limited robot data. Evaluations in simulation demonstrate strong whole-body task performance with 81.9% in HumanoidArena, while real-world experiments further validate WB-WAM with 84.0% mean success across five tasks. Moreover, task-aligned PICO mid-training improves downstream task performance while reducing the need for real-robot demonstrations. These results support heterogeneous whole-body pre-training and human motion transfer as a practical route to data-efficient humanoid loco-manipulation.
Figures & tables
Fig. 2: WB-WAM architecture and three-stage training framework. Top: Progressive training integrates heterogeneous supervision, retargeted PICO motion, and real-robot demonstrations. Bottom: Video and action experts jointly model visual dynamics and whole-body actions conditioned on vision, language, and proprioception. Body and root references are executed through SONIC, while hand references directly control the finger joints.
Fig. 3: Composition of 9-source dataset used for pre-training.
Fig. 4: Pre-training data curation with pose validity, motion continuity, collision, and video quality checks.
Fig. 5: Robot hardware and system.
Method
Football
DoubleDesk
P&PBox
OpenDoor
SitSofa
Boxing
VisNavi
Mean SR
Best baseline (per task)
DP
π0.5
DP
DP
DP
DP
FM
FM 45.5 ± 25.2%
ACT 47.6 ± 24.2%
45.0 ± 10.8%
43.3 ± 6.2%
75.0 ± 4.1%
85.0 ± 10.8%
78.3 ± 14.3%
76.7 ± 2.4%
38.3 ± 14.3%
π0.5 51.2 ± 24.8%
DP 60.0 ± 24.2%
WB-WAM
70.0 ± 8.2%
65.0 ± 4.1%
86.7 ± 2.4%
98.3 ± 2.4%
95.0 ± 4.1%
81.7 ± 2.4%
76.7 ± 4.7%
81.9 ± 12.3%
TABLE I: Comparison on HumanoidArena [ 40 ] . For each task, we report the strongest baseline from the original benchmark.
w/ locomotion
w/o locomotion
Method
Wipe the table
Close the curtain
Make the bed
Move the pillow
Tidy the cloth
Mean SR
Mean TP
A
C
S
A
C
S
A
C
S
C
S
C
S
ACT [ 41 ]
0/20
0/20
0/20
0/20
0/20
0/20
0/20
0/20
0/20
16/20
8/20
19/20
12/20
20%
27.5%
π0.5 [ 42 ]
13/20
12/20
12/20
11/20
10/20
6/20
11/20
10/20
10/20
19/20
8/20
16/20
16/20
52%
61.17%
GR00T N1.6 [ 43 ]
16/20
15/20
15/20
18/20
13/20
12/20
0/20
0/20
0/20
6/20
0/20
17/20
15/20
42%
48.67%
Fast-WAM [ 9 ]
0/20
0/20
0/20
8/20
3/20
2/20
0/20
0/20
0/20
4/20
0/20
9/20
4/20
6%
12.83%
TABLE II: Main comparison with 20 real-robot trials per task. WB-WAM achieves the highest mean SR and mean TP.
Fig. 6: Real-world experiments. WB-WAM enables the robot to perform a variety of whole-body manipulation tasks.
w/ locomotion
w/o locomotion
Training route
Robot [-1pt]Demos
Close the curtain
Push the cart
Wipe the table
Checkout
Mean SR
Mean TP
A
C
S
C
L
S
A
C
S
C
S
Direct post-training
30
13/20
12/20
12/20
13/20
11/20
11/20
12/20
11/20
11/20
6/20
5/20
48.75%
51.04%
Direct post-training
100
16/20
15/20
13/20
19/20
16/20
16/20
16/20
16/20
15/20
10/20
8/20
65%
70.42%
w/ PICO mid-training
30
19/20
17/20
16/20
16/20
17/20
16/20
19/20
18/20
18/20
12/20
9/20
73.75%
78.13%
TABLE III: Effect of PICO mid-training with 20 real-robot trials per condition.
Orange
Lemon
Apple
Method
C
S
C
S
C
S
WB-WAM
17/20
15/20
18/20
17/20
19/20
16/20
TABLE IV: Language-conditioned fruit manipulation.
Fig. 7: Video prediction and real-world execution under visual variation. The three central frames show a video rollout generated by the video expert under a modified visual condition. Real-robot cart pushing is shown in a familiar scene (left) and a novel scene absent from the dataset (right).