Now You Feel It, Now You See Me: Digital-Twin-based Teleoperation Interface for Dexterous Manipulation
Authors: Youngchan Shim, Kyutae Lee, JooYun Kim, Jaeseong Hwang, Harim Ji, Yongseok Lee
Organizations: Department of Robotics and Mechatronics Engineering, DGIST, Daegu, Korea · Department of Mechanical Engineering, Seoul National University, Seoul, Korea
Teleoperation is becoming increasingly important for collecting high-quality demonstrations to teach robots dexterous manipulation skills. For dexterous manipulation, bare-hand tracking provides a practical way to control robotic hands and demonstrate coordinated finger movements without gloves or exoskeletons. However, this type of teleoperation faces two key feedback limitations: a lack of force feedback and visual feedback of occluded region. The absence of force feedback hinders precise and safe manipulation, as operators must infer contact force visually rather than feel them directly. Occlusion by objects or other robot parts impairs the assessment of hand positions, approach distances, and pre-grasp configurations suited to the object's shape. To address these limitations, we propose a digital-twin-based augmented reality (AR) teleoperation interface that integrates bare-hand tracking, bimanual robotic hands, and virtual representations of the robot and task objects. The interface renders tactile measurements on robot hand meshes as visuo-force feedback. Furthermore, it offers a user-controlled virtual view panel or occlusion-aware transparency rendering to provide occlusion-mitigating visual feedback. We evaluated the interface through user studies involving Task 1 and Task 2, examining how force visualization and visual assistance support demonstration collection in tasks requiring careful force regulation and manipulation under occlusion.
Figures & tables
Fig. 2: Visuo-force feedback representations for the left XHAND 1 (left column) and right Inspire Hand (right column). The top row presents force heatmap feedback, which encodes tactile signal magnitude through pad-level color. The bottom row presents force arrow feedback, which conveys force direction and magnitude through fingertip-anchored arrows.
Fig. 3: Hardware setup. Two RB5-850E arms were equipped with an XHAND 1 (left) and an Inspire Hand RH56DFTP (right). The operator wore a Meta Quest 3 and controlled the hands with bare-hand tracking.
Fig. 4: Experimental views across tasks and interface conditions. (a) Task 1 (Paper Cup Relocation): reference scene and grasping views with no visuo-force feedback, force heatmap feedback, and force arrow feedback. (b) Task 2 (Box Lifting): representative views under baseline, user-controlled virtual view panel, and occlusion-aware transparency rendering. Each row shows the corresponding experimental sequence or interface views from left to right.
Fig. 5: Distribution of task outcomes across feedback conditions. The reported p -values compare all three conditions through participant-level permutation tests.
TABLE IV: Task 2: Task Performance (14 participants; 28 trials per condition). VFF denotes visuo-force feedback. Insertion retries exclude the initial attempt and are averaged over successful trials only. Bold values indicate the best performance within each VFF condition.
Department of Computer Science, Stanford University, Stanford, CA 94305, USA. · Department of Mechanical Engineering, Columbia University, New York, NY 10027, USA.