Learning from human videos offers a promising route to acquiring diverse manipulation skills. Extending this capability beyond tabletop settings to long-horizon mobile manipulation requires adapting and composing demonstrated interactions across changing scenes and robot configurations. We present Skill Assembly and Kinematic Imitation (SAKI), a framework connecting human-video skill acquisition, cross-demonstration assembly and closed-loop whole-body execution. SAKI prepares reusable object-centric skills that preserve task-critical interactions while allowing transfer paths to adapt. Given a goal and supplied task dependencies, it selects and orders skills, binds their object roles to the current scene, and carries scene estimates and robot configuration between successive skills. Whole-body kinematic imitation generates coordinated base, arm and gripper motion. During execution, persistent object estimates maintain task references across viewpoint changes, while visual feedback updates remaining trajectories. Real-robot experiments demonstrate skill reuse across layouts and the composition of independently demonstrated interactions into continuous mobile tasks, including tidying and wiping. Ablation results show that task-conditioned reference preparation substantially improves long-horizon task completion with whole-body optimisation and visual feedback held fixed. Check https://aus.bot/research/saki/ for video demos!
Figures & tables
Fig. 1: SAKI transfers object motion and contact roles across articulation, pouring, placement and wiping. Separate task examples include bimanual placement, where container support accompanies manipulation.
Fig. 2: SAKI skill assembly and closed-loop execution. Object-role bindings ground reusable skills in the current scene, and whole-body optimisation generates coordinated robot motion. Accepted visual updates revise the remaining targets while replacement plans preserve committed motion. Paired base/arm completion and gripper events advance the sequence; retained scene estimates and the preceding terminal robot configuration initialise the next approach.
Fig. 3: Skill preparation separates interaction requirements from adjustable motion. Object roles, phases and contact requirements guide reference and grasp preparation for scene-dependent execution. Curves and contacts are schematic; the inset shows kinematic replay.
Fig. 4: Skill assembly across demonstration scenes. Object-role bindings place prepared interaction references in a shared target scene, while the selected order and successor approaches connect them into a mobile task.
Prepared skill
Retained interaction
L1 invocation
Can disposal
Lift, transport, final alignment and release
Can to floor bin; terminal release
Tape placement
Lift, align with opening and release
Tape to tabletop basket; terminal release
Board wiping
Tool grasp, surface contact and sweeping
Eraser to marked board region; terminal hold
TABLE I: Three independently prepared skills retain their interaction requirements while binding to objects in one tidying-and-wiping sequence.
Task
SAKI n/N (%)
Single tasks
S1 Door fully open
7/15 (46.7%)
S2 Long whiteboard wiping
11/15 (73.3%)
S3 Basket handling and tape placement
13/15 (86.7%)
S4 Can disposal
15/15 (100.0%)
S5 Pouring
15/15 (100.0%)
TABLE II: Complete physical outcomes across seven single tasks and three continuous sequences, with 15 initiated trials per condition. Sequence success requires every stage to complete.
Fig. 5: Door opening couples handle engagement to articulated motion. Contact is maintained as the door transitions to partial opening, illustrating an interaction that extends beyond the initial grasp.
Fig. 6: Autonomous suitcase packing requires lid support to overlap with toy placement. One arm maintains the opening while the other releases the toy and withdraws before closure, illustrating concurrent contact roles.
Fig. 7: Phase-specific pose requirements govern physical interaction: time-varying cup orientation during pouring (a) and sustained tool alignment during wiping (b).
Sequence
SAKI w/o ref. optimisation
SAKI
L1 Tidying + wiping
2/15 (13.3%)
9/15 (60.0%)
L2 Collection + pouring
4/15 (26.7%)
13/15 (86.7%)
L3 Three-location collection
1/15 (6.7%)
9/15 (60.0%)
TABLE III: Reference preparation improves continuous-task completion with scene grounding, whole-body solving and online updates retained.
Fig. 8: Visual feedback revises remaining motion while preserving committed execution. A 7.12 cm shift in the retained bin estimate (a) modifies reference targets (b) and the accepted base plan (c); execution reports confirm activation.
Shared subtask
DITTO-adapted
SAKI
Tape into stationary basket
13/15 (86.7%)
15/15 (100.0%)
Same-table cup pouring
10/15 (66.7%)
15/15 (100.0%)
Local whiteboard wiping
8/15 (53.3%)
15/15 (100.0%)
TABLE IV: Local manipulation under shared perception, grasping and execution modules, with 15 trials per operation within fixed-base reach.
Fig. 9: A frozen pouring skill accommodates changed cup, tray and robot starting configurations. Six of twenty trials pair initial arrangements with feedback-derived TCP paths at a common metric scale, registered to the fixed table.
Scaling imitation learning to diverse multi-task robot manipulation remains challenging due to suboptimal demonstrations, behavioral multi-modality, and destructive interference across tasks. While skill-based methods offer a promising direction by decomposing behaviors into reusable abstractions, existing approaches often learn skills that are either biased toward linguistic structure or lack semantic alignment across tasks, limiting generalization. In this work, we propose AtomSkill, a novel framework that learns a semantically aligned Atomic Skill Space from demonstrations and enables robust long-horizon execution through keypose imagination. Our method introduces: (1) semantic contrastive skill alignment, which partitions demonstrations into variable-length atomic skills and employs a contrastive objective to jointly enforce semantic consistency and temporal coherence, yielding a compact and reusable skill library; and (2) action decoding with keypose imagining, where the policy predicts both a skill's terminal keypose and immediate actions, thereby supporting progress-aware skill transitions. During inference, an atomic skill diffusion sampler generates plausible skill sequences, while predicted keyposes autonomously trigger smooth skill chaining. Extensive experiments in simulation and real-world settings show that AtomSkill consistently outperforms state-of-the-art imitation learning and skill-based baselines. Project page: https://atom-skill.github.io.
Yihang Zhu, Weiqing Wang, Shijie Wu +2
ShanghaiTech University, Shanghai, China · InstAdapt
Human manipulation videos are a convenient and intuitive source for robot learning. However, directly transferring human dexterity to robots remains challenging due to perception errors and embodiment gap. To address this, we introduce Video2Sim2Real, a full-stack framework for autonomous skill acquisition from a single human manipulation video. Our framework first uses off-the-shelf foundation models to reconstruct a simulator-ready digital twin and extract robot and object motion priors. Rather than treating the extracted robot motion as a reliable reference throughout execution, our key idea is to recover and leverage the most fundamental sources of supervision from the demonstrated skill: We identify object-centric keyframes to optimize the corresponding robot configurations using object information from the simulator, and use these configurations as anchors that refine the robot motion such that it ultimately has the desired impact on the environment. To bridge the remaining sim-to-real gap, we introduce a sim-to-real strategy that decouples robustness to noisy and incomplete perception from variations in hand-object interaction dynamics. Specifically, we learn to recalibrate robot configurations from noisy real-world point clouds via IL, and leverage residual RL to perform local finger-level adaptations to ensure for robust and effective interactions. Finally, a collision-aware motion planning module enables spatial generalization to novel object configurations. Across several everyday manipulation tasks, Video2Sim2Real improves simulated task success, safety, and trajectory coherence over numerous baselines, and achieves better sim-to-real transfer than existing techniques. These results demonstrate a promising path toward autonomous dexterous skill acquisition from human videos.
Yunhai Han, Jianuo Qiu, Linhao Bai +14
1Georgia Institute of Technology · University of Pennsylvania · 3Toyota Research Institute +1
The ability to acquire skills rapidly and effortlessly while retaining those already mastered is essential for robots. However, current methods still rely on a cumbersome training-time loop that is costly and slow, while eroding skills already mastered. In this paper, we introduce HOST (Human-to-robot One-Shot Skill AcquisiTion), a framework that enables a robot to acquire skills in seconds from a single human video while retaining previously mastered skills. HOST resolves skill acquisition through a cascade of self-grounded prediction. It first estimates the robot's progress within the demonstrated task, then translates the upcoming progression into the robot's own future observations, and finally derives actions from these predicted observations. This cascade is trained on targets coupled to the video demonstration, obtained by mapping the robot trajectory and the video demonstration onto a shared task progress manifold, then redefining each target to align with the future progression of the video. HOST thereby enables the robot to actively follow the demonstrated procedure and adapt it to the robot's embodiment. HOST acquires novel skills at inference time from a single human video in an average of 29 seconds and achieves a 62% average success rate. It exceeds the zero-shot baseline by 45% while retaining previously mastered skills. HOST even exceeds the baseline fine-tuned on 50 robot demonstrations per task while requiring 50 times fewer demonstrations and acquiring each skill 507 times faster. Additional information about HOST is available on the project website.
Guangyan Chen, Meiling Wang, Te Cui +9
Beijing Institute of Technology · X SQUARE ROBOT · Tsinghua University