Egocentric human data offer a path to scaling robot learning beyond costly robot demonstrations, yet the embodiment gap makes raw human trajectories a poor supervisory target for control. Our key insight is that, although low-level actions are embodiment-specific, their underlying motion intent can capture task-relevant structure that transfers across humans and robots. We introduce EgoLAP, a VLA pre-training framework that jointly learns from human and robot trajectories through a shared language-based action chain-of-thought. EgoLAP expresses motion intent as structured, temporally abstracted language actions and pairs them with motion-level reasoning grounded in scene geometry, physics, and object affordances. Across extensive real-world and simulated experiments, EgoLAP transfers human experience to robot control more effectively than alternative action representations and reaches 80.1% mean real-world task progress, a 2.3x performance gain over alternative action representations. Motion-level reasoning also outperforms a composite reasoning format that combines subtask, object-box, and visual-trace reasoning.
Figures & tables
Figure 1: EgoLAP turns egocentric human experience into transferable robot control. Human hands and robots differ in embodiment and low-level execution, yet share the same underlying motion intent when performing a task. EgoLAP captures this intent through motion-level reasoning and language actions, providing a shared supervision target across both data sources and leading to a 2.3× performance gain over alternative action representations.
Figure 2: EgoLAP places human and robot motion in a shared, physically grounded language space. By capturing the physical relationship between the current state and the motion needed to produce the intended effect, motion-level reasoning provides supervision that connects directly to low-level actions that transfers across embodiments.
Figure 3: Attention and gradient structure. Orange marks allowed attention, gray blocked attention, and blue attention whose gradients are stopped at the VLM boundary. Reasoning and language action tokens are causally masked and condition on the observation prefix; action-expert tokens never attend to them, so control requires no text generation.
Figure 4: Zero-shot evaluation on bimanual YAM robots. EgoLAP achieves 80.1% mean progress on real-world tasks—a 2.3× gain over alternative action representations—and a 38.3% success rate on simulated tasks, outperforming ABC-VLA [ 1 ] by 1.8× . When paired with language actions, motion-level reasoning yields a 35.6-point improvement, substantially greater than the gains achieved with other action representations.
Figure 5: Human data helps with EgoLAP. Adding egocentric human data more than doubles simulation success for the same language action policy.
Training data
No reasoning
Reasoning
Robot only
7.2
18.4
Robot + human
16.4
38.3
Table 1: Human data strengthens reasoning. Zero-shot success (%) on the 14-task simulation suite.
Figure 6: Motion-level reasoning is what carries the gain. Motion-level reasoning outperforms ECoT-style reasoning [ 54 ] , and augments it by a large margin.
Figure 7: GRPO over language actions over simulation tasks.
Retrieval
w/o
EgoLAP
Δ
R → H ⋅ 3w R@1
51
76
+25
H → R ⋅ 3w R@1
42
63
+21
R → H ⋅ 2w R@1
69
92
+23
R → H ⋅ 3w MRR
61
80
+19
Table 2: Motion-level reasoning improves cross-embodiment skill retrieval. All values are percentages.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Detailed EgoLAP architecture. A shared SigLIP encoder maps each camera view to visual tokens. The 2B VLM branch predicts motion-level reasoning and the language action; the 300M action expert maps a noisy action chunk to its flow velocity. The two branches hold separate parameters but join the same masked-attention operation at every transformer layer. Action tokens attend to the observation, instruction, and discretized state, never to the textual targets, and gradients through the VLM keys and values are stopped.
Component
Positive direction
Negative direction
Translation x
move forward n cm
move back n cm
Translation z
move up n cm
move down n cm
Translation y
move left n cm
move right n cm
Roll
tilt left θ degrees
tilt right θ degrees
Pitch
tilt back θ degrees
tilt forward θ degrees
Yaw
rotate counterclockwise θ degrees
rotate clockwise θ degrees
Appendix
Table 3: Canonical component vocabulary for language actions. Here n is an integer number of centimeters and θ is an integer multiple of 5 degrees. A component is omitted when its rounded magnitude is zero.
Figure 9: Training data mixture. Inner ring: group weights; outer ring: individual dataset contributions within each group. Only MECKA ( 20% ) and DROID ( 12% ) carry motion-level reasoning labels; the remaining samples are supervised with language actions and continuous actions alone.
Setting
Value
Observations per update
32
Completions per observation
64 (2,048 responses per update)
Gradient accumulation
16 microbatches of 128 responses, one step per rollout bank
linear warmup 5×10−8→10−6 over 20 updates, then constant
Gradient clipping
global norm 1.0, after accumulation
Appendix
Table 5: Shared language-action GRPO settings.
Figure 10: Masking the action-expert loss on human samples hurts. Simulation success over the 14-task suite for EgoLAP (“no human mask”, the main model) and for a matched variant that withholds the flow-matching loss on human samples (“use human mask”). Both variants receive identical language action and motion-level reasoning supervision; they differ only in whether the continuous action expert is trained on human motion. We compare the 25k step checkpoint for both models. Error bars are one binomial standard error.
Figure 11: Per-task zero-shot simulation success. Rows expand the same 14-task suite summarized in Fig. 4 (a), grouped as Pick, Place, Color, and Fold. Columns show ABC-VLA and the reasoning and no-reasoning variants of EgoLAP, EgoFAST, and EgoOAT. Each cell reports the empirical task success rate, with darker shading indicating higher success.
Representation
Reasoning
Robot only
Robot + human
Human-data gain
Language
No
7.2%
16.4%
+9.2 pp
Language
Yes
18.4%
38.3%
+19.9 pp
FAST
Yes
1.4%
8.7%
+7.3 pp
Appendix
Table 6: Language actions benefit more from human data under motion-level reasoning supervision. Zero-shot success on the 14-task simulation suite. Human-data gains are absolute percentage-point differences. Reasoning denotes motion-level reasoning supervision.
Vision-Language-Action (VLA) models benefit from large-scale and diverse embodied data, yet scaling robot trajectory collection is costly and labor-intensive. Recent advances show that large-scale egocentric human videos provide complementary real-world supervision in pretraining. However, joint training on human and robot data remains challenging due to divergences in action spaces, embodiment structures, temporal dynamics, and supervision quality. We introduce ACE-EGO-0, a unified VLA pretraining framework jointly leveraging heterogeneous data sources. To extract large-scale pretraining supervision from egocentric human videos, we build a scalable egocentric video-to-action pipeline that converts raw human videos into robot-format pseudo-action trajectories. To make these labels comparable with robot demonstrations, ACE-EGO-0 uses a unified action representation based on camera-space actions, morphology conditioning, and time-aligned action chunking. To robustly leverage noisy pseudo-action supervision from egocentric human videos, we formulate a reliability-aware training objective with a human auxiliary loss that concentrates supervision on reliable signals. We instantiate ACE-EGO-0 on 4.53K hours of robot and simulation data, together with 1.48K hours of pseudo-action-labeled egocentric human data. Experiments show that incorporating large-scale human supervision under reliability-aware weighting consistently improves both unified joint pretraining and supervised fine-tuning. ACE-EGO-0 achieves state-of-the-art performance on RoboCasa GR1 TableTop and RoboTwin 2.0, while demonstrating strong transfer to real-world bimanual manipulation.
Human egocentric video captures rich manipulation demonstrations without any robot hardware, yet transferring these skills to robots remains challenging due to the embodiment gap between human and robot in both visual appearance and kinematics. We present HumanEgo, a framework that bridges the embodiment gap by lifting each human demonstration to an entity-level representation of hand-object interaction, and training a flow matching policy with dense auxiliary objectives that amplify supervision from every trajectory. HumanEgo is robot-data-free, hardware-agnostic, data-efficient, and zero-shot human-to-robot transferable. With only 30 minutes of human videos per task, HumanEgo achieves 92.5% average success across four real-world tasks (75% with just 15 minutes), outperforms matched-time robot teleoperation by 41%, and robustly transfers zero-shot across novel robots, cameras, and environments. We release HumanEgo as an easy-to-use, open-source framework for learning robot policies directly from human data: https://github.com/TX-Leo/HumanEgo
Robotics faces a fundamental challenge of data scarcity. Unlike language or vision research, there is no internet-scale dataset for robotic manipulation. A promising path forward is to leverage egocentric human data, which can be collected more easily, with greater breadth, and at a larger scale. Towards this end, we investigate key design choices for learning across human and humanoid embodiments equipped with dexterous five-finger hands, using the π0.5 model as a foundation. Our results show that human data enables robots to learn new task semantics and compose existing skills into novel behaviors without corresponding robot data. The paper website is here: https://egopipaper.github.io/