Glove-based motion capture is emerging as a scalable approach to collecting dexterous-hand demonstration data. However, due to the kinematic gap between the human and robot hand, the recorded human motions cannot be executed directly on the robot, especially for contact-rich tool-use tasks involving in-hand reorientation. Prior work bridges this gap in simulation through reinforcement learning (RL) or trajectory optimization, but the human contact pattern is hard to preserve under such formulations, often producing unnatural manipulation and unstable functional grasps. These methods also train a separate policy or solve a separate optimization for each reference trajectory, which is inefficient. To solve these problems, we propose DexTaG, a tactile-guided RL framework for dexterous manipulation. During training, tactile signals captured by the glove guide policy search toward the measured human contact pattern, reducing reliance on precise reference geometry for contact supervision. To improve efficiency, we train a single generalizable retargeter jointly on all training trajectories of the same object. The retargeter is further distilled into a tactile-free student controller conditioned on the target object trajectory for real-world deployment. On marker-pen and hammer manipulation tasks, DexTaG learns natural, contact-rich behaviors that baselines with distance-based contact heuristics fail to learn, generalizes to held-out trajectories of the same object and task, and outperforms single-trajectory baselines on OakInk2.
Figures & tables
Fig. 2: Overview of DexTaG. The pipeline consists of three stages: (i) collecting hand–object demonstrations with tactile gloves, (ii) using measured tactile references to guide RL retargeting, and (iii) distilling the retargeter into a tactile-free controller for real-world deployment.
Figure 2
Fig. 5: Example collected trajectories with per-timestep tactile maps for marker-pen (top) and hammer (bottom) manipulation.
Fig. 6: Qualitative Comparison. First row: the reference glove tactile map vs. the simulated tactile map at each timestep. Second and third rows: policy rollouts of DexTaG and Geom-gated, respectively, on marker-pen and hammer manipulation.
Method
Comp.
Obj.
Hand
Marker pen
Ours
87.5±4.0
1.127±0.033
1.856±0.182
Prox-only
64.5±6.1
1.106±0.064
1.823±0.086
Sim-only
61.5±13.8
1.100±0.097
1.771±0.136
Geom-gated
48.6±24.0
1.161±0.138
1.983±0.315
w/o tactile
24.7±2.9
1.299±0.020
2.304±0.012
TABLE I: Tactile reward ablation. Trajectory Completion (%) and average per-step tracking rewards: mean ± standard deviation over five training seeds.