Towards a General Humanoid Loco-Manipulation Model via Egocentric Whole-Body Human Data Pretraining
Authors: Chongyang Xu, Zhao Wu, Jin Chen, Yiming Jiang, Jinhui Ye, Yuming Jiang, Shifeng Zhang, Ziliang Feng, +4 more
Organizations: Alibaba Group · Sichuan University · Shanghai Innovation Institute · Beihang University · Hong Kong University of Science and Technology
Humanoid whole-body manipulation has advanced rapidly, enabling policies to coordinate locomotion, posture, bimanual interaction, and dexterous hand movements. Meanwhile, egocentric human videos provide diverse examples of everyday interactions across objects and scenes, offering scalable supervision without robot operation. However, existing supervision from these videos provides limited coverage of whole-body movement and coordination with hand-object interaction, while obtaining such supervision through humanoid teleoperation is also costly and difficult to scale. We therefore explore how human experience can support scalable learning of humanoid loco-manipulation. To support this study, we introduce HumanVerse-500, a 500-hour dataset of diverse human loco-manipulation behaviors in open-world environments, collected with a lightweight wearable system that synchronizes egocentric video with body and hand motion. Building on this dataset, we develop λ0, a whole-body humanoid vision-language-action policy, through three-stage training that first learns interaction from diverse egocentric datasets, then coordinates body and hand motion using HumanVerse-500, and finally adapts the policy to downstream tasks and robot embodiments. Across these stages, λ0 learns a shared representation space for human experience transfer, while domain-specific interfaces handle differences between human and robot states and actions. We evaluate λ0 on SIMPLE and 4 real-world loco-manipulation tasks, achieving state-of-the-art performance, and further analyze its scaling behavior, generalization, and training-stage contributions to understand how human data support downstream whole-body humanoid control. We will release our code, models, and data to support further research.
Figures & tables
Figure 1: Introducing λ0 with HumanVerse-500 . We introduce a robot-free egocentric human-data collection paradigm and use it to collect HumanVerse-500 , a 500-hour dataset of whole-body human loco-manipulation. Then, we develop λ0 , a unified whole-body humanoid vision–language–action policy, through a three-stage recipe combining large-scale Internet manipulation data, HumanVerse-500 , and embodiment-specific data. Finally, λ0 achieves state-of-the-art performance and strong generalization across simulated and real-world loco-manipulation tasks.
Figure 2: HumanVerse-500 . The wearable capture setup, task and motion distributions, and paired egocentric video and whole-body motions illustrate the dataset’s scale and diversity.
Figure 3: Egocentric task diversity in HumanVerse-500 . Sixty egocentric frames span carrying, cleaning, plant care, office tasks, and object handling.
Figure 4: Whole-body Human Data. Egocentric RGB frames, body motion, and hand motion.
Figure 5: Whole-body motion retargeting to G1 with dexterous hands.
Figure 6: Three-stage training of λ0 . The policy learns interaction from egocentric data, whole-body coordination from HumanVerse-500 , and downstream control from real-robot data.
Figure 7: Real-world data collection and evaluation. (a) Egocentric teleoperation collects whole-body demonstrations. (b) Four loco-manipulation tasks evaluate task success. (c) Robot-seen and robot-unseen objects are used for standard task and object-generalization evaluation, respectively.
Method
Bottle Disposal
Toy Storage
Box Transport
Chair Placement
Average
π0.5 ( Black et al., 2025 )
2/10
28.0
5/10
64.5
5/10
58.0
4/10
65.5
16/40
54.0
GR00T N1.6 ( GEAR Team et al., 2025 )
0/10
10.0
1/10
19.0
0/10
10.0
2/10
46.0
3/40
21.3
Ψ0 ( Wei and others, 2026 )
1/10
26.5
2/10
30.0
1/10
37.5
0/10
33.0
4/40
31.8
StarVLA ( StarVLA Community, 2026 )
0/10
10.0
5/10
55.0
3/10
71.0
3/10
82.5
11/40
54.6
λ0 (Ours)
5/10
79.0
7/10
81.0
7/10
84.5
6/10
77.0
25/40
80.4
Table 1: Real-world task performance. We report success rates over ten trials and weighted progress (%) for each task. Averages weight the four tasks equally.
Model
XMovePick
BendPick
Handover
Mobile P&P
Grasp
XMoveBendPick
Average
Overall
H-RDT ( Bi et al., 2026 )
0/0/20
0/0/10
0/10/0
0/0/0
0/0/0
0/0/0
0.0/1.7/5.0
2.2
InternVLA-M1 ( Chen et al., 2025 )
0/0/0
50/50/0
0/0/0
0/0/0
0/0/0
30/50/70
13.3/16.7/11.7
13.9
EgoVLA ( Yang et al., 2026 )
0/10/20
70/50/80
0/40/30
0/0/0
100/100/70
30/50/40
33.3/41.7/40.0
38.3
Diffusion Policy ( Chi et al., 2025 )
30/30/20
100/80/60
30/20/40
40/0/0
80/90/80
0/0/0
46.7/36.7/33.3
38.9
GR00T N1.6 ( GEAR Team et al., 2025 )
100/100/70
70/70/60
10/30/30
0/0/0
90/90/70
40/40/10
51.7/55.0/40.0
48.9
π0.5 ( Black et al., 2025 )
70/50/10
100/100/80
50/40/50
30/30/30
100/100/80
0/0/0
58.3/53.3/41.7
51.1
Table 2: SIMPLE task performance. Success rates (%) at Levels 0–2 with ten trials per task–level pair. Average is computed per level. Overall averages all 18 task–level pairs.
Figure 8: Human-data scaling of validation loss. (a) Training curves on representative and clean holdouts at five data fractions. (b) Minimum validation loss versus data scale with descriptive log-linear fits over the measured range. Minima are selected independently on each holdout.
Variant
S1
S2
Robot
SIMPLE (%)
Real-world (%)
L0
L1
L2
Overall
Success
Progress
w/o human-data pretraining
∘
∘
∙
83.3
83.3
71.7
79.4
15.0
54.0
w/o Stage II mid-training
∙
∘
∙
83.6
86.7
76.7
82.3
22.5
59.9
w/o Stage I pretraining
∘
∙
∙
88.0
91.7
83.7
87.8
50.0
75.9
λ0 (Full Recipe)
∙
∙
∙
90.0
93.3
86.7
90.0
62.5
80.4
Table 3: Training-stage ablations with fixed robot supervision. Filled/open circles indicate included/removed data. Metrics report SIMPLE success and real-world averages across four tasks.
Figure 9: Average real-world task progress versus Stage-II human-data fraction.
IIIS, Tsinghua University, Beijing, China. · Xiong’an Institute of Artificial Intelligence, Xiong’an, China. · The University of Melbourne, Melbourne, Australia.