Organizations: Robotics Institute, Carnegie Mellon University, USA. · Department of Computer Science, Keio University, Japan. · College of Engineering, University of California, Berkeley, USA. · Bosch Center for Artificial Intelligence, Pittsburgh, PA, USA.
Dexterous manipulation requires both large-scale task progression and precise contact-rich interaction, making it challenging to collect demonstrations that effectively support both regimes. We present SkillWeave, a heterogeneous demonstration framework for long-horizon dexterous manipulation that combines teleoperation for coarse reaching and transport with kinesthetic teaching for precise, contact-rich skills. To address the visual mismatch introduced by the demonstrator's presence during kinesthetic data collection, we propose an object-mask-conditioned diffusion policy that uses offline object segmentation for training supervision and a lightweight learned mask predictor at deployment, avoiding online segmentation and image inpainting. To mitigate distribution shift between independently trained sub-task policies, we introduce successor-aware terminal steering, which selects among actions sampled from the predecessor policy to guide the system toward states supported by the successor's demonstrated initial-state distribution. Across three real-world long-horizon tasks, SkillWeave achieves 27% average end-to-end success. Mask-conditioned kinesthetic policies improve dexterous sub-task success to an average of 65%, while successor-aware handoffs achieve an average composition efficiency of 87%. These results show that matching demonstration modality to interaction regime, explicitly addressing kinesthetic visual mismatch, and steering policy handoffs toward successor-supported states substantially improves long-horizon dexterous manipulation. Videos and code are available at skillweave-authors.github.io .
Figures & tables
\fnum@figure : Mask-conditioned diffusion policy architecture. External- and wrist-camera RGB observations are encoded by ResNet-18, and object masks guide cross-attention pooling over the resulting spatial feature maps. The pooled visual features are concatenated with end-effector and hand proprioception to condition a 1-D U-Net diffusion model that predicts robot action chunks. Training uses offline SAM 3 masks, while deployment uses learned mask predictors.
\fnum@figure : Overview of the policy composition and handoff framework. A predecessor policy π(i) executes its sub-task until entering the terminal phase. To facilitate a reliable transition to the independently trained successor policy π(i+1) , we repeatedly sample candidate action chunks from π(i) , evaluate their predicted terminal states according to proximity to the successor’s demonstrated start-state distribution Si+1start , and execute the best candidate. This closed-loop steering procedure moves the system toward a state supported by the successor’s training distribution, after which control is handed off to π(i+1) . The approach reduces distribution shift at policy boundaries without requiring additional transition demonstrations or a separately trained policy.
Task ( i )
j
Sub-task
Success Criterion
Cube Reorientation
1
Reach, grasp, and reorient cube
Cube securely grasped without dropping
2
Rotate cube in-hand
Cube rotated by at least 90∘
Nut Removal & Storage
1
Reach for nut
Hand positioned over the nut
2
Unscrew nut
Nut fully detached from screw
3
Place nut in container
Nut released inside the container
Kettle Preparation
1
Reach, grasp, and position kettle
Kettle grasped by handle and positioned under faucet
TABLE I: Long-horizon task decomposition and sub-task success criteria. Si(j) is determined by the corresponding criterion.
\fnum@figure : Example RGB observations for policies. Raw observations during teleoperation and kinesthetic teaching are directly fed to policies; KineDex inpaints the image to align with inference-time observation, but often blurs out robot hand and generates hallucinated parts.
Sub-task
Teleop.
Kin.
KineDex
Mask-Kin.
Reach, grasp, & reorient
70%
0%
0%
—
In-hand rotation
10%
30%
0%
70%
Reach for nut
85%
0%
40%
—
Unscrew nut
5%
0%
60%
75%
Place nut in container
30%
0%
5%
—
Reach, grasp, & position kettle
80%
0%
0%
—
TABLE II: Sub-task and long-horizon success rates. (a) Success rates of individual sub-task policies across demonstration and visual-processing strategies. (b) End-to-end success rates of long-horizon systems. All results are reported over 20 trials.
Teleop
KineDex
SkillWeave w/o Steering
SkillWeave
Task
S
S∗
C
S
S∗
C
S
S∗
C
S
S∗
C
Task 1
0%
7%
0%
0%
0%
–
10%
49%
20%
45%
49%
92%
Task 2
0%
1%
0%
0%
1%
0%
0%
19%
0%
15%
19%
79%
Task 3
0%
0%
–
0%
0%
–
5%
22%
23%
20%
22%
91%
Avg.
0%
3%
0%
0%
1%
0%
5%
30%
14%
27%
30%
87%
TABLE III: Long-horizon task performance and composition efficiency. Note Kinesthetic-only results are omitted because S∗=0 for all tasks, making C undefined.
\fnum@figure : SkillWeave long-horizon visualizations. From top to bottom: Cube Reorientation, Nut Removal and Storage, and Kettle Preparation. Blue shows frames for policies trained with teleoperation, while red shows frames for mask-conditioned kinesthetic teaching policies.
\fnum@figure : Comparison between using successor-aware terminal steering and not using terminal steering. Without terminal steering (left), predecessor policy ended at finger joint states significantly different from successor policy start joint states (middle); In comparison, with terminal steering (right), predecessor policy ended at finger joint states closer to successor policy start joint states.
Many dexterous manipulation tasks require the object to remain securely held throughout the interaction. From the perspective of hand-object relational motion, such manipulation comprises four canonical skills: grasping, relocation, in-hand rotation, and in-hand translation. Human hands flexibly compose these skills to accomplish complex tasks. Existing approaches, however, model these skills separately with skill-specific action constraints, objectives, or even dedicated hand morphologies, which breaks the compatibility and continuity required for long-horizon composition. In this work, we present a unified framework that models all four skills in a single formulation that shares the same state and action spaces and a common objective structure. This formulation enables distillation of a single cross-skill policy conditioned on the relational motion objectives, which achieves strong performance across all four skills, generalizes to unseen objects, remains robust to disturbances, and chains skills into long-horizon manipulation without switching policies. The framework also transfers effectively across different hand morphologies. Overall, our results suggest that different dexterous manipulation skills can be viewed as instantiations of a shared task formulation, revealing the intrinsic consistency. Project page: https://zdchan.github.io/UniCross/
Hui Zhang, Julian Ferchow, Jie Song +1
ETH Zürich, Switzerland · inspire AG, Zürich, Switzerland · HKUST (Guangzhou), China
Achieving human-level dexterity in complex, unstructured environments requires the seamless integration of whole-body scene interaction and dexterous object manipulation skills. While existing physics-based controllers generate physically plausible behaviors in each domain, they largely address these two capabilities independently. In this paper, we present Co2Skill that integrates scene interaction and dexterous manipulation through a unified policy formulation. Built on a pretrained motion prior, the policy uses task and phase dependent observation masks to select information relevant to the current interaction goals. We introduce a goal-conditioned loco-manipulation curriculum that combines partial reference guidance for precision with exploration from varied initial states while allowing goal-directed execution beyond the demonstrated trajectories. We further introduce a cross-task curriculum that jointly trains individual skills and selected task sequences, preserving physical states across task boundaries and maintaining grasps during subsequent scene interactions. Together, these support sequential task execution and simultaneous scene interaction with object manipulation. We evaluate sitting, standing, climbing, stair traversal, and goal-directed manipulation, together with sequential execution and with random different conditions. Additionally, we demonstrate skill compositions in indoor environments, illustrating their integration within the same control formulation.
Jeonghwan Kim, Hyeonwoo Kim, Hanbyul Joo
Department of Computer Science Seoul National University
Successfully automating dexterous, long-horizon robotic manipulation requires frameworks capable of both high-level reasoning and fine-grained execution. Traditional task and motion planning (TAMP), while excellent at symbolic planning, is often brittle in contact-rich operations. Simultaneously, imitation learning (IL), while effective in manipulation tasks with visual feedback, is limited by its low capability in spatial generalization and multi-stage operation. To reconcile their complementary strengths and limitations, we propose DR-LfD (Decomposed and Reorganized Skills Learned from Demonstrations), a framework that seamlessly integrates visuomotor policies into a TAMP-gated decision-making system. Based on contact relationships, DR-LfD decomposes human demonstrations into atomic skills, which are reproduced as visuomotor policies or object-centric primitives. The initiation, termination, and constraints of the visuomotor policies are carefully modeled and implemented in a TAMP-compatible form, enabling reorganization of skills learned from different sources. DR-LfD transforms the learning problem from one requiring exponential demonstration data over possible skill sequences to one whose demonstration burden scales with the number of distinct skill types, with limited data for each skill. Through comprehensive real-world and simulation benchmarking across diverse scenarios, we demonstrate the strong performance of DR-LfD on tasks involving multiple steps, unseen setups, and physical constraints. Project website: https://dr-lfd.github.io/DR-LfD-website.
Yizhou Chen, Hang Xu, Dongjie Yu +7
The University of Hong Kong, Hong Kong · JD.com (JingDong) · Nanyang Technological University (NTU), Singapore +3