Egocentric human demonstrations offer an accessible source of task experience, but differences in body scale and controller response, together with missing robot states, limit their value as humanoid training supervision. We present EgoAlign, a data-construction framework that converts these demonstrations into action and state supervision compatible with a general-purpose, continuous whole-body controller, without collecting physical-robot demonstrations. Using the target-robot model and simulator, EgoAlign guides demonstration collection through execution feedback. It preserves locomotion references for visually guided periodic stepping while adapting upper-body interaction geometry through scale alignment and controller-in-the-loop refinement. A final causal replay reconstructs the corresponding robot states and motion-token labels for training with the human observations. We assess the resulting supervision by fine-tuning a vision--language--action model solely on adapted human demonstrations and deploying it zero-shot on a physical humanoid. The resulting policies perform long-range object relocation, navigation to unseen goal positions, and independently evaluated foot interaction. Refinement improves simulated hand alignment and physical pickup success over kinematic alignment alone, while human collection reduces on-site acquisition time relative to teleoperation. https://lambdahumanoid.github.io/EgoAlign/
Fig. 1: EgoAlign pipeline. (1) A wearable setup captures synchronized dual-view images and human motion. (2) Real-time SONIC–MuJoCo feedback guides demonstrators to adjust their motions. (3) Scale alignment adapts upper-body motion to human-reference hand targets through optimization and replay. (4) State reconstruction pairs pre-action robot states and same-tick motion tokens. (5) Human-only task training produces a π0.5 policy for zero-shot robot deployment, predicting 64-D motion tokens and hand commands from live observations; SONIC decodes tokens with state history into body commands. Dashed arrows indicate collection-time and offline feedback.
Fig. 2: Human collection and robot hardware with matched camera heights. The downward view captures near-field interaction, while the level view provides distant navigation cues.
Fig. 3: Scale alignment. Constructing a robot-scale SMPL proxy (left), matching fixed human hand targets (middle), and optimizing bounded arm-joint updates (right). The bottom loop uses SONIC–MuJoCo replay residuals for iterative refinement.
Fig. 4: Physical-G1 tasks and endpoint protocols for object relocation and for navigation and foot-based interaction. Object relocation comprises pickup, transport, and placement. Navigation and foot interaction form a complete behavior but are evaluated independently to separate navigation and foot-operation performance.
Task / Subtask
Endpoint
Variant
Trials
Score
Full Suc.
S1
S2
S3
S4
Object Relocation
Direct
Aligned (ours)
20
82.5
65
100
90
75
65
w/o Nav. data
20
83.8
65
100
95
75
65
NoAlign
20
0.0
0
0
0
0
0
Kinematic (Round 0)
20
–
–
50
0
–
–
w/o State Recon.
20
0.0
0
0
0
0
0
Zero History
20
0.0
0
0
0
0
0
TABLE II: Main and ablation results on G1 (20 trials per row). Score: mean stage completion; Full Suc.: final-stage success. Round 0 evaluates pickup only; dashes indicate unreported entries.
Fig. 5: Stage score and full success for independently tested (a) navigation and (b) foot interaction, with 20 trials per setting.
Data
Demo
Reset
Total
Rate
Teleop
70s
60s
130s
1.0×
Human ego
15s
10s
25s
5.2×
TABLE III: On-site object relocation data-acquisition time per demonstration.
Fig. 6: Ablations of (a) navigation data, (b) alignment/state reconstruction, and (c) level view. (d) Refinement: simulated right-palm error (macro mean and 95% bootstrap CI) and physical S1/S2 success (20 Direct trials each for NoAlign and Rounds 0/2). GMR–SONIC fails at S1. Round 1 has simulation results only.