Egocentric human data offer a path to scaling robot learning beyond costly robot demonstrations, yet the embodiment gap makes raw human trajectories a poor supervisory target for control. Our key insight is that, although low-level actions are embodiment-specific, their underlying motion intent can capture task-relevant structure that transfers across humans and robots. We introduce EgoLAP, a VLA pre-training framework that jointly learns from human and robot trajectories through a shared language-based action chain-of-thought. EgoLAP expresses motion intent as structured, temporally abstracted language actions and pairs them with motion-level reasoning grounded in scene geometry, physics, and object affordances. Across extensive real-world and simulated experiments, EgoLAP transfers human experience to robot control more effectively than alternative action representations and reaches 80.1% mean real-world task progress, a 2.3x performance gain over alternative action representations. Motion-level reasoning also outperforms a composite reasoning format that combines subtask, object-box, and visual-trace reasoning.
Figures & tables
Figure 1: EgoLAP turns egocentric human experience into transferable robot control. Human hands and robots differ in embodiment and low-level execution, yet share the same underlying motion intent when performing a task. EgoLAP captures this intent through motion-level reasoning and language actions, providing a shared supervision target across both data sources and leading to a 2.3× performance gain over alternative action representations.
Figure 2: EgoLAP places human and robot motion in a shared, physically grounded language space. By capturing the physical relationship between the current state and the motion needed to produce the intended effect, motion-level reasoning provides supervision that connects directly to low-level actions that transfers across embodiments.
Figure 3: Attention and gradient structure. Orange marks allowed attention, gray blocked attention, and blue attention whose gradients are stopped at the VLM boundary. Reasoning and language action tokens are causally masked and condition on the observation prefix; action-expert tokens never attend to them, so control requires no text generation.
Figure 4: Zero-shot evaluation on bimanual YAM robots. EgoLAP achieves 80.1% mean progress on real-world tasks—a 2.3× gain over alternative action representations—and a 38.3% success rate on simulated tasks, outperforming ABC-VLA [ 1 ] by 1.8× . When paired with language actions, motion-level reasoning yields a 35.6-point improvement, substantially greater than the gains achieved with other action representations.
Figure 5: Human data helps with EgoLAP. Adding egocentric human data more than doubles simulation success for the same language action policy.
Training data
No reasoning
Reasoning
Robot only
7.2
18.4
Robot + human
16.4
38.3
Table 1: Human data strengthens reasoning. Zero-shot success (%) on the 14-task simulation suite.
Figure 6: Motion-level reasoning is what carries the gain. Motion-level reasoning outperforms ECoT-style reasoning [ 54 ] , and augments it by a large margin.
Figure 7: GRPO over language actions over simulation tasks.
Retrieval
w/o
EgoLAP
Δ
R → H ⋅ 3w R@1
51
76
+25
H → R ⋅ 3w R@1
42
63
+21
R → H ⋅ 2w R@1
69
92
+23
R → H ⋅ 3w MRR
61
80
+19
Table 2: Motion-level reasoning improves cross-embodiment skill retrieval. All values are percentages.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Detailed EgoLAP architecture. A shared SigLIP encoder maps each camera view to visual tokens. The 2B VLM branch predicts motion-level reasoning and the language action; the 300M action expert maps a noisy action chunk to its flow velocity. The two branches hold separate parameters but join the same masked-attention operation at every transformer layer. Action tokens attend to the observation, instruction, and discretized state, never to the textual targets, and gradients through the VLM keys and values are stopped.
Component
Positive direction
Negative direction
Translation x
move forward n cm
move back n cm
Translation z
move up n cm
move down n cm
Translation y
move left n cm
move right n cm
Roll
tilt left θ degrees
tilt right θ degrees
Pitch
tilt back θ degrees
tilt forward θ degrees
Yaw
rotate counterclockwise θ degrees
rotate clockwise θ degrees
Appendix
Table 3: Canonical component vocabulary for language actions. Here n is an integer number of centimeters and θ is an integer multiple of 5 degrees. A component is omitted when its rounded magnitude is zero.
Figure 9: Training data mixture. Inner ring: group weights; outer ring: individual dataset contributions within each group. Only MECKA ( 20% ) and DROID ( 12% ) carry motion-level reasoning labels; the remaining samples are supervised with language actions and continuous actions alone.
Setting
Value
Observations per update
32
Completions per observation
64 (2,048 responses per update)
Gradient accumulation
16 microbatches of 128 responses, one step per rollout bank
linear warmup 5×10−8→10−6 over 20 updates, then constant
Gradient clipping
global norm 1.0, after accumulation
Appendix
Table 5: Shared language-action GRPO settings.
Figure 10: Masking the action-expert loss on human samples hurts. Simulation success over the 14-task suite for EgoLAP (“no human mask”, the main model) and for a matched variant that withholds the flow-matching loss on human samples (“use human mask”). Both variants receive identical language action and motion-level reasoning supervision; they differ only in whether the continuous action expert is trained on human motion. We compare the 25k step checkpoint for both models. Error bars are one binomial standard error.
Figure 11: Per-task zero-shot simulation success. Rows expand the same 14-task suite summarized in Fig. 4 (a), grouped as Pick, Place, Color, and Fold. Columns show ABC-VLA and the reasoning and no-reasoning variants of EgoLAP, EgoFAST, and EgoOAT. Each cell reports the empirical task success rate, with darker shading indicating higher success.
Representation
Reasoning
Robot only
Robot + human
Human-data gain
Language
No
7.2%
16.4%
+9.2 pp
Language
Yes
18.4%
38.3%
+19.9 pp
FAST
Yes
1.4%
8.7%
+7.3 pp
Appendix
Table 6: Language actions benefit more from human data under motion-level reasoning supervision. Zero-shot success on the 14-task simulation suite. Human-data gains are absolute percentage-point differences. Reasoning denotes motion-level reasoning supervision.