Learning from human videos offers a promising route to acquiring diverse manipulation skills. Extending this capability beyond tabletop settings to long-horizon mobile manipulation requires adapting and composing demonstrated interactions across changing scenes and robot configurations. We present Skill Assembly and Kinematic Imitation (SAKI), a framework connecting human-video skill acquisition, cross-demonstration assembly and closed-loop whole-body execution. SAKI prepares reusable object-centric skills that preserve task-critical interactions while allowing transfer paths to adapt. Given a goal and supplied task dependencies, it selects and orders skills, binds their object roles to the current scene, and carries scene estimates and robot configuration between successive skills. Whole-body kinematic imitation generates coordinated base, arm and gripper motion. During execution, persistent object estimates maintain task references across viewpoint changes, while visual feedback updates remaining trajectories. Real-robot experiments demonstrate skill reuse across layouts and the composition of independently demonstrated interactions into continuous mobile tasks, including tidying and wiping. Ablation results show that task-conditioned reference preparation substantially improves long-horizon task completion with whole-body optimisation and visual feedback held fixed. Check https://aus.bot/research/saki/ for video demos!
Figures & tables
Fig. 1: SAKI transfers object motion and contact roles across articulation, pouring, placement and wiping. Separate task examples include bimanual placement, where container support accompanies manipulation.
Fig. 2: SAKI skill assembly and closed-loop execution. Object-role bindings ground reusable skills in the current scene, and whole-body optimisation generates coordinated robot motion. Accepted visual updates revise the remaining targets while replacement plans preserve committed motion. Paired base/arm completion and gripper events advance the sequence; retained scene estimates and the preceding terminal robot configuration initialise the next approach.
Fig. 3: Skill preparation separates interaction requirements from adjustable motion. Object roles, phases and contact requirements guide reference and grasp preparation for scene-dependent execution. Curves and contacts are schematic; the inset shows kinematic replay.
Fig. 4: Skill assembly across demonstration scenes. Object-role bindings place prepared interaction references in a shared target scene, while the selected order and successor approaches connect them into a mobile task.
Prepared skill
Retained interaction
L1 invocation
Can disposal
Lift, transport, final alignment and release
Can to floor bin; terminal release
Tape placement
Lift, align with opening and release
Tape to tabletop basket; terminal release
Board wiping
Tool grasp, surface contact and sweeping
Eraser to marked board region; terminal hold
TABLE I: Three independently prepared skills retain their interaction requirements while binding to objects in one tidying-and-wiping sequence.
Task
SAKI n/N (%)
Single tasks
S1 Door fully open
7/15 (46.7%)
S2 Long whiteboard wiping
11/15 (73.3%)
S3 Basket handling and tape placement
13/15 (86.7%)
S4 Can disposal
15/15 (100.0%)
S5 Pouring
15/15 (100.0%)
TABLE II: Complete physical outcomes across seven single tasks and three continuous sequences, with 15 initiated trials per condition. Sequence success requires every stage to complete.
Fig. 5: Door opening couples handle engagement to articulated motion. Contact is maintained as the door transitions to partial opening, illustrating an interaction that extends beyond the initial grasp.
Fig. 6: Autonomous suitcase packing requires lid support to overlap with toy placement. One arm maintains the opening while the other releases the toy and withdraws before closure, illustrating concurrent contact roles.
Fig. 7: Phase-specific pose requirements govern physical interaction: time-varying cup orientation during pouring (a) and sustained tool alignment during wiping (b).
Sequence
SAKI w/o ref. optimisation
SAKI
L1 Tidying + wiping
2/15 (13.3%)
9/15 (60.0%)
L2 Collection + pouring
4/15 (26.7%)
13/15 (86.7%)
L3 Three-location collection
1/15 (6.7%)
9/15 (60.0%)
TABLE III: Reference preparation improves continuous-task completion with scene grounding, whole-body solving and online updates retained.
Fig. 8: Visual feedback revises remaining motion while preserving committed execution. A 7.12 cm shift in the retained bin estimate (a) modifies reference targets (b) and the accepted base plan (c); execution reports confirm activation.
Shared subtask
DITTO-adapted
SAKI
Tape into stationary basket
13/15 (86.7%)
15/15 (100.0%)
Same-table cup pouring
10/15 (66.7%)
15/15 (100.0%)
Local whiteboard wiping
8/15 (53.3%)
15/15 (100.0%)
TABLE IV: Local manipulation under shared perception, grasping and execution modules, with 15 trials per operation within fixed-base reach.
Fig. 9: A frozen pouring skill accommodates changed cup, tray and robot starting configurations. Six of twenty trials pair initial arrangements with feedback-derived TCP paths at a common metric scale, registered to the fixed table.