Towards a General Humanoid Loco-Manipulation Model via Egocentric Whole-Body Human Data Pretraining
Authors: Chongyang Xu, Zhao Wu, Jin Chen, Yiming Jiang, Jinhui Ye, Yuming Jiang, Shifeng Zhang, Ziliang Feng, +4 more
Organizations: Alibaba Group · Sichuan University · Shanghai Innovation Institute · Beihang University · Hong Kong University of Science and Technology
Humanoid whole-body manipulation has advanced rapidly, enabling policies to coordinate locomotion, posture, bimanual interaction, and dexterous hand movements. Meanwhile, egocentric human videos provide diverse examples of everyday interactions across objects and scenes, offering scalable supervision without robot operation. However, existing supervision from these videos provides limited coverage of whole-body movement and coordination with hand-object interaction, while obtaining such supervision through humanoid teleoperation is also costly and difficult to scale. We therefore explore how human experience can support scalable learning of humanoid loco-manipulation. To support this study, we introduce HumanVerse-500, a 500-hour dataset of diverse human loco-manipulation behaviors in open-world environments, collected with a lightweight wearable system that synchronizes egocentric video with body and hand motion. Building on this dataset, we develop λ0, a whole-body humanoid vision-language-action policy, through three-stage training that first learns interaction from diverse egocentric datasets, then coordinates body and hand motion using HumanVerse-500, and finally adapts the policy to downstream tasks and robot embodiments. Across these stages, λ0 learns a shared representation space for human experience transfer, while domain-specific interfaces handle differences between human and robot states and actions. We evaluate λ0 on SIMPLE and 4 real-world loco-manipulation tasks, achieving state-of-the-art performance, and further analyze its scaling behavior, generalization, and training-stage contributions to understand how human data support downstream whole-body humanoid control. We will release our code, models, and data to support further research.
Figures & tables
Figure 1: Introducing λ0 with HumanVerse-500 . We introduce a robot-free egocentric human-data collection paradigm and use it to collect HumanVerse-500 , a 500-hour dataset of whole-body human loco-manipulation. Then, we develop λ0 , a unified whole-body humanoid vision–language–action policy, through a three-stage recipe combining large-scale Internet manipulation data, HumanVerse-500 , and embodiment-specific data. Finally, λ0 achieves state-of-the-art performance and strong generalization across simulated and real-world loco-manipulation tasks.
Figure 2: HumanVerse-500 . The wearable capture setup, task and motion distributions, and paired egocentric video and whole-body motions illustrate the dataset’s scale and diversity.
Figure 3: Egocentric task diversity in HumanVerse-500 . Sixty egocentric frames span carrying, cleaning, plant care, office tasks, and object handling.
Figure 4: Whole-body Human Data. Egocentric RGB frames, body motion, and hand motion.
Figure 5: Whole-body motion retargeting to G1 with dexterous hands.
Figure 6: Three-stage training of λ0 . The policy learns interaction from egocentric data, whole-body coordination from HumanVerse-500 , and downstream control from real-robot data.
Figure 7: Real-world data collection and evaluation. (a) Egocentric teleoperation collects whole-body demonstrations. (b) Four loco-manipulation tasks evaluate task success. (c) Robot-seen and robot-unseen objects are used for standard task and object-generalization evaluation, respectively.
Method
Bottle Disposal
Toy Storage
Box Transport
Chair Placement
Average
π0.5 ( Black et al., 2025 )
2/10
28.0
5/10
64.5
5/10
58.0
4/10
65.5
16/40
54.0
GR00T N1.6 ( GEAR Team et al., 2025 )
0/10
10.0
1/10
19.0
0/10
10.0
2/10
46.0
3/40
21.3
Ψ0 ( Wei and others, 2026 )
1/10
26.5
2/10
30.0
1/10
37.5
0/10
33.0
4/40
31.8
StarVLA ( StarVLA Community, 2026 )
0/10
10.0
5/10
55.0
3/10
71.0
3/10
82.5
11/40
54.6
λ0 (Ours)
5/10
79.0
7/10
81.0
7/10
84.5
6/10
77.0
25/40
80.4
Table 1: Real-world task performance. We report success rates over ten trials and weighted progress (%) for each task. Averages weight the four tasks equally.
Model
XMovePick
BendPick
Handover
Mobile P&P
Grasp
XMoveBendPick
Average
Overall
H-RDT ( Bi et al., 2026 )
0/0/20
0/0/10
0/10/0
0/0/0
0/0/0
0/0/0
0.0/1.7/5.0
2.2
InternVLA-M1 ( Chen et al., 2025 )
0/0/0
50/50/0
0/0/0
0/0/0
0/0/0
30/50/70
13.3/16.7/11.7
13.9
EgoVLA ( Yang et al., 2026 )
0/10/20
70/50/80
0/40/30
0/0/0
100/100/70
30/50/40
33.3/41.7/40.0
38.3
Diffusion Policy ( Chi et al., 2025 )
30/30/20
100/80/60
30/20/40
40/0/0
80/90/80
0/0/0
46.7/36.7/33.3
38.9
GR00T N1.6 ( GEAR Team et al., 2025 )
100/100/70
70/70/60
10/30/30
0/0/0
90/90/70
40/40/10
51.7/55.0/40.0
48.9
π0.5 ( Black et al., 2025 )
70/50/10
100/100/80
50/40/50
30/30/30
100/100/80
0/0/0
58.3/53.3/41.7
51.1
Table 2: SIMPLE task performance. Success rates (%) at Levels 0–2 with ten trials per task–level pair. Average is computed per level. Overall averages all 18 task–level pairs.
Figure 8: Human-data scaling of validation loss. (a) Training curves on representative and clean holdouts at five data fractions. (b) Minimum validation loss versus data scale with descriptive log-linear fits over the measured range. Minima are selected independently on each holdout.
Variant
S1
S2
Robot
SIMPLE (%)
Real-world (%)
L0
L1
L2
Overall
Success
Progress
w/o human-data pretraining
∘
∘
∙
83.3
83.3
71.7
79.4
15.0
54.0
w/o Stage II mid-training
∙
∘
∙
83.6
86.7
76.7
82.3
22.5
59.9
w/o Stage I pretraining
∘
∙
∙
88.0
91.7
83.7
87.8
50.0
75.9
λ0 (Full Recipe)
∙
∙
∙
90.0
93.3
86.7
90.0
62.5
80.4
Table 3: Training-stage ablations with fixed robot supervision. Filled/open circles indicate included/removed data. Metrics report SIMPLE success and real-world averages across four tasks.
Figure 9: Average real-world task progress versus Stage-II human-data fraction.
Whole-body humanoid loco-manipulation requires coordinating the robot's entire kinematic chain. However, most existing systems typically decouple the upper and lower bodies into separate controllers, limiting such coordination and yielding behaviors similar to those of a wheeled dual-arm platform. In this paper, we ask what it takes to build a whole-body native vision-language-action (VLA) model that maps language and pixels directly to all of the humanoid's degrees of freedom. We conduct a systematic empirical study organized as a roadmap of one-variable-at-a-time experiments across three phases: whole-body teleoperation, VLA model design, and heterogeneous co-training. Our study yields several intriguing findings: a joint-based whole-body teleoperation interface outperforms alternatives that only partially expose the humanoid's degrees of freedom; a VLA pretrained on static and wheeled dual-arm platforms transfers surprisingly well to a humanoid's full action space; and co-training with HuMI, the humanoid analog of UMI, extends the policy to new objects and instructions without additional whole-body teleoperation on those targets. Following this roadmap yields OpenHLM, an open-source recipe for whole-body humanoid loco-manipulation. In a challenging long-horizon task that spans a wide vertical range of the humanoid, OpenHLM outperforms two state-of-the-art humanoid VLA baselines (GR00T N1.6 and Ψ0) using less than half the total demonstration time. Our code, training data, and model checkpoints are available at [https://openhlm-project.github.io/].
Yingdong Hu, Haodong Zhu, Boyuan Zheng +6
Tsinghua University · Shanghai Qi Zhi Institute · Spirit AI
Human demonstrations capture diverse scenes and rich whole-body skills without requiring robot teleoperation. Prior work on egocentric transfer has emphasized scene generalization in loco-manipulation under decoupled control, leaving direct transfer of coordinated whole-body skills less explored. We present EgoHumanoid-V2, the first egocentric human-to-humanoid skill transfer framework for coordinated whole-body loco-manipulation. At its core, coarse-to-fine action alignment combines kinematic reference correction with dynamics-aware refinement. It improves end-effector pose accuracy while preserving whole-body coordination. We also use robot-arm rendering and training-time image augmentation to reduce the visual embodiment gap and improve viewpoint robustness. On four real-world tasks, vision-language-action (VLA) policies trained on aligned human data show zero-shot skill transfer without target-task robot demonstrations. Task scores are comparable to those of policies trained on teleoperation data at a lower collection cost. These results support human data as direct skill supervision.
Jin Chen, Yiming Jiang, Chongyang Xu +8
OpenDriveLab at The University of Hong Kong · Alibaba Group · Shanghai Innovation Institute +4
Humanoid loco-manipulation demands coordinated body and hand behavior, while conventional robot pre-training data provide limited coverage of such whole-body motion. We present WB-WAM, a World Action Model that incorporates explicit whole-body action supervision into generative video pre-training. A shared physical action space integrates body, root, and dexterous hand annotations from heterogeneous sources, enabling joint video and action learning from 1880.2 hours of partially annotated video and motion data. The resulting priors are refined through PICO mid-training and adapted to robot tasks with auxiliary forward kinematics supervision. We construct WB-Datasets to support these stages with retargeted egocentric human demonstrations and robot trajectories, allowing task-aligned human motion to supplement limited robot data. Evaluations in simulation demonstrate strong whole-body task performance with 81.9% in HumanoidArena, while real-world experiments further validate WB-WAM with 84.0% mean success across five tasks. Moreover, task-aligned PICO mid-training improves downstream task performance while reducing the need for real-robot demonstrations. These results support heterogeneous whole-body pre-training and human motion transfer as a practical route to data-efficient humanoid loco-manipulation.
Chuan Qin, Shaoting Zhu, Siyuan Luo +4
IIIS, Tsinghua University, Beijing, China. · Xiong’an Institute of Artificial Intelligence, Xiong’an, China. · The University of Melbourne, Melbourne, Australia.