Organizations: Robotics Institute, Carnegie Mellon University, USA. · Department of Computer Science, Keio University, Japan. · College of Engineering, University of California, Berkeley, USA. · Bosch Center for Artificial Intelligence, Pittsburgh, PA, USA.
Dexterous manipulation requires both large-scale task progression and precise contact-rich interaction, making it challenging to collect demonstrations that effectively support both regimes. We present SkillWeave, a heterogeneous demonstration framework for long-horizon dexterous manipulation that combines teleoperation for coarse reaching and transport with kinesthetic teaching for precise, contact-rich skills. To address the visual mismatch introduced by the demonstrator's presence during kinesthetic data collection, we propose an object-mask-conditioned diffusion policy that uses offline object segmentation for training supervision and a lightweight learned mask predictor at deployment, avoiding online segmentation and image inpainting. To mitigate distribution shift between independently trained sub-task policies, we introduce successor-aware terminal steering, which selects among actions sampled from the predecessor policy to guide the system toward states supported by the successor's demonstrated initial-state distribution. Across three real-world long-horizon tasks, SkillWeave achieves 27% average end-to-end success. Mask-conditioned kinesthetic policies improve dexterous sub-task success to an average of 65%, while successor-aware handoffs achieve an average composition efficiency of 87%. These results show that matching demonstration modality to interaction regime, explicitly addressing kinesthetic visual mismatch, and steering policy handoffs toward successor-supported states substantially improves long-horizon dexterous manipulation. Videos and code are available at skillweave-authors.github.io .
Figures & tables
\fnum@figure : Mask-conditioned diffusion policy architecture. External- and wrist-camera RGB observations are encoded by ResNet-18, and object masks guide cross-attention pooling over the resulting spatial feature maps. The pooled visual features are concatenated with end-effector and hand proprioception to condition a 1-D U-Net diffusion model that predicts robot action chunks. Training uses offline SAM 3 masks, while deployment uses learned mask predictors.
\fnum@figure : Overview of the policy composition and handoff framework. A predecessor policy π(i) executes its sub-task until entering the terminal phase. To facilitate a reliable transition to the independently trained successor policy π(i+1) , we repeatedly sample candidate action chunks from π(i) , evaluate their predicted terminal states according to proximity to the successor’s demonstrated start-state distribution Si+1start , and execute the best candidate. This closed-loop steering procedure moves the system toward a state supported by the successor’s training distribution, after which control is handed off to π(i+1) . The approach reduces distribution shift at policy boundaries without requiring additional transition demonstrations or a separately trained policy.
Task ( i )
j
Sub-task
Success Criterion
Cube Reorientation
1
Reach, grasp, and reorient cube
Cube securely grasped without dropping
2
Rotate cube in-hand
Cube rotated by at least 90∘
Nut Removal & Storage
1
Reach for nut
Hand positioned over the nut
2
Unscrew nut
Nut fully detached from screw
3
Place nut in container
Nut released inside the container
Kettle Preparation
1
Reach, grasp, and position kettle
Kettle grasped by handle and positioned under faucet
TABLE I: Long-horizon task decomposition and sub-task success criteria. Si(j) is determined by the corresponding criterion.
\fnum@figure : Example RGB observations for policies. Raw observations during teleoperation and kinesthetic teaching are directly fed to policies; KineDex inpaints the image to align with inference-time observation, but often blurs out robot hand and generates hallucinated parts.
Sub-task
Teleop.
Kin.
KineDex
Mask-Kin.
Reach, grasp, & reorient
70%
0%
0%
—
In-hand rotation
10%
30%
0%
70%
Reach for nut
85%
0%
40%
—
Unscrew nut
5%
0%
60%
75%
Place nut in container
30%
0%
5%
—
Reach, grasp, & position kettle
80%
0%
0%
—
TABLE II: Sub-task and long-horizon success rates. (a) Success rates of individual sub-task policies across demonstration and visual-processing strategies. (b) End-to-end success rates of long-horizon systems. All results are reported over 20 trials.
Teleop
KineDex
SkillWeave w/o Steering
SkillWeave
Task
S
S∗
C
S
S∗
C
S
S∗
C
S
S∗
C
Task 1
0%
7%
0%
0%
0%
–
10%
49%
20%
45%
49%
92%
Task 2
0%
1%
0%
0%
1%
0%
0%
19%
0%
15%
19%
79%
Task 3
0%
0%
–
0%
0%
–
5%
22%
23%
20%
22%
91%
Avg.
0%
3%
0%
0%
1%
0%
5%
30%
14%
27%
30%
87%
TABLE III: Long-horizon task performance and composition efficiency. Note Kinesthetic-only results are omitted because S∗=0 for all tasks, making C undefined.
\fnum@figure : SkillWeave long-horizon visualizations. From top to bottom: Cube Reorientation, Nut Removal and Storage, and Kettle Preparation. Blue shows frames for policies trained with teleoperation, while red shows frames for mask-conditioned kinesthetic teaching policies.
\fnum@figure : Comparison between using successor-aware terminal steering and not using terminal steering. Without terminal steering (left), predecessor policy ended at finger joint states significantly different from successor policy start joint states (middle); In comparison, with terminal steering (right), predecessor policy ended at finger joint states closer to successor policy start joint states.