Organizations: Department of Computer Science, Stanford University, Stanford, CA 94305, USA. · Department of Mechanical Engineering, Columbia University, New York, NY 10027, USA.
Teleoperated demonstrations are a primary source of data for robot manipulation, and teleoperated interventions are a primary mechanism for correcting policies at deployment. Yet most teleoperation systems close the loop through vision alone and are built around parallel-jaw grippers, limiting both what the robot can execute and what the operator can express through it. This is most damaging in shared autonomy, where the operator sees the scene only through occluded cameras and must take over a dexterous hand mid-task, often with an object already grasped. We present DITTO-X, a hand-agnostic dexterous teleoperation interface that renders joint-level force and fingertip contact events from sensing already on the robot hand, and drives three commercial dexterous hands (Sharpa, Wuji, and Inspire) without per-hand redesign. Because the exoskeleton is actuated, DITTO-X also supports reverse teleoperation, in which the robot back-drives the operator's fingers into its own configuration before control is transferred, so the human enters the loop already matched to the state they inherit. Our results show that DITTO-X improves demonstration quality and throughput over a commercial hand-tracking glove, both in regular data collection and in human intervention during policy deployment for contact-rich manipulation tasks. More information can be found from our website: https://tml.stanford.edu/ditto-x/.
Figures & tables
Fig. 2: Design improvements from DITTO. (A) A third finger module instruments middle-finger MCP abduction/adduction, MCP flexion, and PIP flexion, and a simplified thumb reduces weight. (B) Fingertip haptics from 8 mm z-axis LRAs (JYLRA0825Z) driven by DRV2605L drivers.
Sharpa
Wuji2
Inspire
Actuated DoF
22
20
6
DITTO-X 1:1 Joint Mapping
Thumb
×
×
× (2 DOF only)
Index, Middle
✓
✓
× (1 DOF only)
Sensing
Joint Current
✓
✓
1D / finger
Tactile
✓
×
×
Feedback
Torque
✓
✓
1D / finger
TABLE I: Comparison of supported robotic hands. DITTO-X 1:1 Joint Mapping indicates whether a finger on the target hand can be paired joint-to-joint directly with the corresponding finger on DITTO-X.
Fig. 3: Evaluation tasks. (a) Battery, (b) tong, and (c) raspberry on the Sharpa hand span forceful insertion, tool use, and fragile grasping. (d) Drawer, on Wuji: one episode recruiting different parts of the hand in turn, hooking, pinching, then closing with a fist. (e) Bunny lifting, on Inspire, a hand with only six actuated degrees of freedom.
Fig. 4: Object size (T1) and compliance (T2) identification accuracy. Objects used for each task are shown on the right of the corresponding subfigure.
Fig. 5: User study T3 and T4 success rates (a) and throughput (b).
Fig. 6: Examples of human intervention with DITTO-X (Upper) and Manus (Bottom) . The bottom row shows corresponding wrist camera views and operator hand pose. Human intervention is annotated with blue border and policy execution is in green .
Fig. 7: Policy success rates of pretrain and two rounds of DAgger.
Fig. 8: Comparisons between DAgger and quantity match baselines.
Pretrain
DITTO-X DAgger
Manus
DITTO-X
r1
r2
Tong
Grasp
73.3
76.7
90.0
96.7
Toy
23.3
66.7
83.3
90.0
Place
23.3
63.3
80.0
90.0
Release
13.3
63.3
73.3
86.7
Rasp.
Grasp
60.0
90.0
96.7
93.3
TABLE II: Stage-wise completion rate (%) over 30 rollouts per run. A stage is credited only if all preceding stages succeeded.
Humans routinely wield tools, swap grasps, and reposition objects within a single hand, seamlessly orchestrating contact transitions that span translation, reorientation, and finger gaiting. Endowing robot dexterous hands with this level of in-hand dexterity through teleoperation requires precise control of object motion via dynamic hand-object contact, yet current teleoperation systems remain far from this capability. To bridge this gap, we take a major step towards human-level dexterous teleoperation by introducing TeleDexter, a hand-object co-tracking controller that maps operator intent into learned, low-level contact execution. The controller is trained on consecutive co-tracking subgoals derived from human reference motions, utilizing a hybrid reward that couples sparse subgoal objectives with dense tracking rewards to enable learning across diverse interaction modalities rather than frame-wise trajectory imitation. The entire pipeline requires only single-stage RL and, with random action masking and domain randomization, transfers zero-shot to the real robot. We evaluate TeleDexter on seven challenging dexterous teleoperation tasks spanning object reorientation and long-horizon tool use across two dexterous hands, achieving a 75% average success rate where all baselines consistently fail. Furthermore, the collected demonstrations successfully train autonomous policies via behavioral cloning, marking a concrete step towards human-level dexterous teleoperation.
Puhao Li, Zeyuan Chen, Yingying Wu +9
1Tsinghua University · 2State Key Lab of General Artificial Intelligence, BIGAI · 3Peking University
Dexterous teleoperation requires precise arm-hand coordination, low-latency feedback, and robust interaction in real-world contact-rich environments. This paper presents a modular bilateral teleoperation framework that integrates operator-side input interfaces with a robot-side dexterous hand and compliant robotic arm in a unified control architecture. The system supports position-based hand retargeting, differential arm control, multi-scale haptic feedback, and shared control for stable manipulation. We validate the framework through a real-world dexterous manipulation task, highlighting coordinated arm-hand control and contact-aware interaction. Beyond feasibility, we identify key design insights related to cross-embodiment mismatch, haptic feedback granularity, and shared control. The proposed platform provides a practical teleoperation system and a foundation for collecting high-quality demonstrations for future learning-from-demonstration research.
Stefano Dalla Gasperina, Dong Ho Kang, Haiyun Zhang +10
The University of Texas at Austin, Austin, TX, USA · Sony Group Corporation, Tokyo, Japan · Meta Reality Labs Research, Redmond, WA, USA
Teleoperation is a key interface for controlling dexterous robotic hands and collecting demonstrations for imitation learning. Its effectiveness largely depends on kinematic retargeting, which maps operator hand motions to feasible and intuitive robot hand motions. Existing methods often require hand-crafted objectives, precise calibration, or global shape matching between human and robot hand spaces, making them sensitive to hand-specific tuning and less reliable across different dexterous hands. We propose AnyDexRT, a calibration-free retargeting method for intuitive dexterous teleoperation across human-like dexterous hands. AnyDexRT combines self-supervised fingertip correspondence learning with few-shot human guidance to anchor the mapping in task-relevant regions, and further refines pinch-related poses using a contact classifier. Experiments on diverse dexterous hands and real-world teleoperation tasks show that AnyDexRT improves retargeting quality, reduces manual tuning, and provides more intuitive and efficient control than prior retargeting methods. Project website: https://chenxi-wang.github.io/projects/anydexrt
Chenxi Wang, Ying Feng, Hongjie Fang +4
1Noematrix · 2Shanghai Jiao Tong University · 3Shanghai Innovation Institute