Now You Feel It, Now You See Me: Digital-Twin-based Teleoperation Interface for Dexterous Manipulation
Authors: Youngchan Shim, Kyutae Lee, JooYun Kim, Jaeseong Hwang, Harim Ji, Yongseok Lee
Organizations: Department of Robotics and Mechatronics Engineering, DGIST, Daegu, Korea · Department of Mechanical Engineering, Seoul National University, Seoul, Korea
Teleoperation is becoming increasingly important for collecting high-quality demonstrations to teach robots dexterous manipulation skills. For dexterous manipulation, bare-hand tracking provides a practical way to control robotic hands and demonstrate coordinated finger movements without gloves or exoskeletons. However, this type of teleoperation faces two key feedback limitations: a lack of force feedback and visual feedback of occluded region. The absence of force feedback hinders precise and safe manipulation, as operators must infer contact force visually rather than feel them directly. Occlusion by objects or other robot parts impairs the assessment of hand positions, approach distances, and pre-grasp configurations suited to the object's shape. To address these limitations, we propose a digital-twin-based augmented reality (AR) teleoperation interface that integrates bare-hand tracking, bimanual robotic hands, and virtual representations of the robot and task objects. The interface renders tactile measurements on robot hand meshes as visuo-force feedback. Furthermore, it offers a user-controlled virtual view panel or occlusion-aware transparency rendering to provide occlusion-mitigating visual feedback. We evaluated the interface through user studies involving Task 1 and Task 2, examining how force visualization and visual assistance support demonstration collection in tasks requiring careful force regulation and manipulation under occlusion.
Figures & tables
Fig. 2: Visuo-force feedback representations for the left XHAND 1 (left column) and right Inspire Hand (right column). The top row presents force heatmap feedback, which encodes tactile signal magnitude through pad-level color. The bottom row presents force arrow feedback, which conveys force direction and magnitude through fingertip-anchored arrows.
Fig. 3: Hardware setup. Two RB5-850E arms were equipped with an XHAND 1 (left) and an Inspire Hand RH56DFTP (right). The operator wore a Meta Quest 3 and controlled the hands with bare-hand tracking.
Fig. 4: Experimental views across tasks and interface conditions. (a) Task 1 (Paper Cup Relocation): reference scene and grasping views with no visuo-force feedback, force heatmap feedback, and force arrow feedback. (b) Task 2 (Box Lifting): representative views under baseline, user-controlled virtual view panel, and occlusion-aware transparency rendering. Each row shows the corresponding experimental sequence or interface views from left to right.
Fig. 5: Distribution of task outcomes across feedback conditions. The reported p -values compare all three conditions through participant-level permutation tests.
TABLE IV: Task 2: Task Performance (14 participants; 28 trials per condition). VFF denotes visuo-force feedback. Insertion retries exclude the initial attempt and are averaged over successful trials only. Bold values indicate the best performance within each VFF condition.
Dexterous teleoperation requires precise arm-hand coordination, low-latency feedback, and robust interaction in real-world contact-rich environments. This paper presents a modular bilateral teleoperation framework that integrates operator-side input interfaces with a robot-side dexterous hand and compliant robotic arm in a unified control architecture. The system supports position-based hand retargeting, differential arm control, multi-scale haptic feedback, and shared control for stable manipulation. We validate the framework through a real-world dexterous manipulation task, highlighting coordinated arm-hand control and contact-aware interaction. Beyond feasibility, we identify key design insights related to cross-embodiment mismatch, haptic feedback granularity, and shared control. The proposed platform provides a practical teleoperation system and a foundation for collecting high-quality demonstrations for future learning-from-demonstration research.
Stefano Dalla Gasperina, Dong Ho Kang, Haiyun Zhang +10
The University of Texas at Austin, Austin, TX, USA · Sony Group Corporation, Tokyo, Japan · Meta Reality Labs Research, Redmond, WA, USA
Teleoperated demonstrations are a primary source of data for robot manipulation, and teleoperated interventions are a primary mechanism for correcting policies at deployment. Yet most teleoperation systems close the loop through vision alone and are built around parallel-jaw grippers, limiting both what the robot can execute and what the operator can express through it. This is most damaging in shared autonomy, where the operator sees the scene only through occluded cameras and must take over a dexterous hand mid-task, often with an object already grasped. We present DITTO-X, a hand-agnostic dexterous teleoperation interface that renders joint-level force and fingertip contact events from sensing already on the robot hand, and drives three commercial dexterous hands (Sharpa, Wuji, and Inspire) without per-hand redesign. Because the exoskeleton is actuated, DITTO-X also supports reverse teleoperation, in which the robot back-drives the operator's fingers into its own configuration before control is transferred, so the human enters the loop already matched to the state they inherit. Our results show that DITTO-X improves demonstration quality and throughput over a commercial hand-tracking glove, both in regular data collection and in human intervention during policy deployment for contact-rich manipulation tasks. More information can be found from our website: https://tml.stanford.edu/ditto-x/.
Zhanpeng He, Joaquin Palacios, Zhangyu Wang +5
Department of Computer Science, Stanford University, Stanford, CA 94305, USA. · Department of Mechanical Engineering, Columbia University, New York, NY 10027, USA.
Teleoperation for contact-rich manipulation remains challenging, especially when using low-cost, motion-only interfaces that provide no haptic feedback. Virtual reality controllers enable intuitive motion control but do not allow operators to directly perceive or regulate contact forces, limiting task performance. To address this, we propose an augmented reality (AR) visualization of the impedance controller's target pose and its displacement from each robot end effector. This visualization conveys the forces generated by the controller, providing operators with intuitive, real-time feedback without expensive haptic hardware. We evaluate the design in a dual-arm manipulation study with 17 participants who repeatedly reposition a box with and without the AR visualization. Results show that AR visualization reduces completion time by 24% for force-critical lifting tasks, with no significant effect on sliding tasks where precise force control is less critical. These findings indicate that making the impedance target visible through AR is a viable approach to improve human-robot interaction for contact-rich teleoperation.
Gijs van den Brandt, Femke van Beek, Elena Torta
Department of Mechanical Engineering, Eindhoven University of Technology, The Netherlands