Egocentric human demonstrations offer an accessible source of task experience, but differences in body scale and controller response, together with missing robot states, limit their value as humanoid training supervision. We present EgoAlign, a data-construction framework that converts these demonstrations into action and state supervision compatible with a general-purpose, continuous whole-body controller, without collecting physical-robot demonstrations. Using the target-robot model and simulator, EgoAlign guides demonstration collection through execution feedback. It preserves locomotion references for visually guided periodic stepping while adapting upper-body interaction geometry through scale alignment and controller-in-the-loop refinement. A final causal replay reconstructs the corresponding robot states and motion-token labels for training with the human observations. We assess the resulting supervision by fine-tuning a vision--language--action model solely on adapted human demonstrations and deploying it zero-shot on a physical humanoid. The resulting policies perform long-range object relocation, navigation to unseen goal positions, and independently evaluated foot interaction. Refinement improves simulated hand alignment and physical pickup success over kinematic alignment alone, while human collection reduces on-site acquisition time relative to teleoperation. https://lambdahumanoid.github.io/EgoAlign/
Fig. 1: EgoAlign pipeline. (1) A wearable setup captures synchronized dual-view images and human motion. (2) Real-time SONIC–MuJoCo feedback guides demonstrators to adjust their motions. (3) Scale alignment adapts upper-body motion to human-reference hand targets through optimization and replay. (4) State reconstruction pairs pre-action robot states and same-tick motion tokens. (5) Human-only task training produces a π0.5 policy for zero-shot robot deployment, predicting 64-D motion tokens and hand commands from live observations; SONIC decodes tokens with state history into body commands. Dashed arrows indicate collection-time and offline feedback.
Fig. 2: Human collection and robot hardware with matched camera heights. The downward view captures near-field interaction, while the level view provides distant navigation cues.
Fig. 3: Scale alignment. Constructing a robot-scale SMPL proxy (left), matching fixed human hand targets (middle), and optimizing bounded arm-joint updates (right). The bottom loop uses SONIC–MuJoCo replay residuals for iterative refinement.
Fig. 4: Physical-G1 tasks and endpoint protocols for object relocation and for navigation and foot-based interaction. Object relocation comprises pickup, transport, and placement. Navigation and foot interaction form a complete behavior but are evaluated independently to separate navigation and foot-operation performance.
Task / Subtask
Endpoint
Variant
Trials
Score
Full Suc.
S1
S2
S3
S4
Object Relocation
Direct
Aligned (ours)
20
82.5
65
100
90
75
65
w/o Nav. data
20
83.8
65
100
95
75
65
NoAlign
20
0.0
0
0
0
0
0
Kinematic (Round 0)
20
–
–
50
0
–
–
w/o State Recon.
20
0.0
0
0
0
0
0
Zero History
20
0.0
0
0
0
0
0
TABLE II: Main and ablation results on G1 (20 trials per row). Score: mean stage completion; Full Suc.: final-stage success. Round 0 evaluates pickup only; dashes indicate unreported entries.
Fig. 5: Stage score and full success for independently tested (a) navigation and (b) foot interaction, with 20 trials per setting.
Data
Demo
Reset
Total
Rate
Teleop
70s
60s
130s
1.0×
Human ego
15s
10s
25s
5.2×
TABLE III: On-site object relocation data-acquisition time per demonstration.
Fig. 6: Ablations of (a) navigation data, (b) alignment/state reconstruction, and (c) level view. (d) Refinement: simulated right-palm error (macro mean and 95% bootstrap CI) and physical S1/S2 success (20 Direct trials each for NoAlign and Rounds 0/2). GMR–SONIC fails at S1. Round 1 has simulation results only.
Human demonstrations capture diverse scenes and rich whole-body skills without requiring robot teleoperation. Prior work on egocentric transfer has emphasized scene generalization in loco-manipulation under decoupled control, leaving direct transfer of coordinated whole-body skills less explored. We present EgoHumanoid-V2, the first egocentric human-to-humanoid skill transfer framework for coordinated whole-body loco-manipulation. At its core, coarse-to-fine action alignment combines kinematic reference correction with dynamics-aware refinement. It improves end-effector pose accuracy while preserving whole-body coordination. We also use robot-arm rendering and training-time image augmentation to reduce the visual embodiment gap and improve viewpoint robustness. On four real-world tasks, vision-language-action (VLA) policies trained on aligned human data show zero-shot skill transfer without target-task robot demonstrations. Task scores are comparable to those of policies trained on teleoperation data at a lower collection cost. These results support human data as direct skill supervision.
Jin Chen, Yiming Jiang, Chongyang Xu +8
OpenDriveLab at The University of Hong Kong · Alibaba Group · Shanghai Innovation Institute +4
Vision-language-action (VLA) models across robot embodiments require high-quality observation--action supervision to learn deployable action distributions, yet scaling such robot data remains difficult, especially for high-DoF humanoids. Teleoperation provides controller-aligned supervision, while human egocentric videos capture diverse bimanual manipulation but do not directly provide executable robot actions. We introduce Human-as-Humanoid, a human-to-humanoid supervision framework that enables near-real-time human-centric action generation, making human demonstrations usable for high-DoF humanoid VLA training by jointly aligning the robot embodiment, the sensing setup, and the action-label interface. Built on PrimeU, a human-aligned 60-DoF upper-body humanoid, Human-as-Humanoid uses synchronized ego-exo videos to pair deployment-aligned egocentric observations with exocentric motion recovery, retargets the recovered human motion through staged Inverse Kinematics (IK) into controller-aligned 60-DoF action chunks, and trains the VLA model with Forward Kinematics (FK)-aware supervision to preserve wrist and fingertip task-space geometry. This converts large-scale human demonstrations from visual observations into executable observation--action supervision for the target humanoid. Experiments validate the conversion chain at the motion-recovery, robot-action-space, and real-robot deployment levels. Human-as-Humanoid yields a 4.8--7.2x raw demonstration-throughput gain over humanoid teleoperation in our data-collection analysis, and on several downstream tasks, policies post-trained only with the converted human labels generalize to real-robot deployment without target-task robot demonstrations. The official project website is available at https://zgc-embodyai.github.io/Human-as-Humanoid.
Xiaopeng Lin, Ruoqi Yang, Shijie Lian +14
The Hong Kong University of Science and Technology (Guangzhou) · DeepCybo · ZGCA +4
Human egocentric video captures rich manipulation demonstrations without any robot hardware, yet transferring these skills to robots remains challenging due to the embodiment gap between human and robot in both visual appearance and kinematics. We present HumanEgo, a framework that bridges the embodiment gap by lifting each human demonstration to an entity-level representation of hand-object interaction, and training a flow matching policy with dense auxiliary objectives that amplify supervision from every trajectory. HumanEgo is robot-data-free, hardware-agnostic, data-efficient, and zero-shot human-to-robot transferable. With only 30 minutes of human videos per task, HumanEgo achieves 92.5% average success across four real-world tasks (75% with just 15 minutes), outperforms matched-time robot teleoperation by 41%, and robustly transfers zero-shot across novel robots, cameras, and environments. We release HumanEgo as an easy-to-use, open-source framework for learning robot policies directly from human data: https://github.com/TX-Leo/HumanEgo