cs.ROOct 6, 2026

EgoLAP: Learning from Egocentric Human Data through Language-Action Reasoning

Authors: Lihan Zha, Shresth Grover, Tenny Yin, Samuel M. Bateman, Hengkai Pan, Mengchao Zhang, Aykut Onol, Allen Z. Ren, +2 more

Organizations: Princeton University · Toyota Research Institute · Physical Intelligence

Abstract

Egocentric human data offer a path to scaling robot learning beyond costly robot demonstrations, yet the embodiment gap makes raw human trajectories a poor supervisory target for control. Our key insight is that, although low-level actions are embodiment-specific, their underlying motion intent can capture task-relevant structure that transfers across humans and robots. We introduce EgoLAP, a VLA pre-training framework that jointly learns from human and robot trajectories through a shared language-based action chain-of-thought. EgoLAP expresses motion intent as structured, temporally abstracted language actions and pairs them with motion-level reasoning grounded in scene geometry, physics, and object affordances. Across extensive real-world and simulated experiments, EgoLAP transfers human experience to robot control more effectively than alternative action representations and reaches 80.1% mean real-world task progress, a 2.3x performance gain over alternative action representations. Motion-level reasoning also outperforms a composite reasoning format that combines subtask, object-box, and visual-trace reasoning.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining

    Jun 15, 2026Hao Li, Ganlong Zhao, Yufei Liu +8Robotic Data AcquisitionEgocentric Video

  2. Ego-Pi: VLA Fine-Tuning for Ego-Centric Human and Robot Data

    Jun 6, 2026Ji Woong Kim, Ke Wang, Zipeng Fu +4Egocentric DatasetHumanoid