Organizations: Shanghai Jiao Tong University · Sharpa Robotics · Beijing Institute of Technology · State Key Laboratory of General Artificial Intelligence, BIGAI
Dexterous manipulation requires tactile feedback. However, robot tactile demonstrations are difficult to scale,because dexterous-hand teleoperation provides limited tactile feedback to the operator. In contrast, human demonstrations offer a substantially more scalable source of diverse tactile interactions. Motivated by a simple premise: hands can change, but the underlying physics of interaction does not. We leverage human tactile data to improve dexterous manipulation policies. Specifically, we first build a tactile motion-capture system that synchronously records images, tactile signals, and hand motions. Using this system, we construct the UVTA dataset spanning five contact-rich tasks, with 1,000 human demonstrations covering diverse interaction patterns and 150 robot demonstrations per task. To transfer the underlying physics of human interaction to robot control, we propose a Unified Visual-Tactile-Action Model that maps both embodiments into aligned tactile and action representations and jointly predicts future action and tactile trajectories. The joint objective enables human demonstrations to supervise contact-aware representation learning, while only robot actions are executed during deployment. In real-robot evaluations across five tasks, our method achieves an average success rate of 70%, outperforming the strongest visual-tactile baseline, which achieves 29%, and an architecture ablation, which achieves 42%. Performance improves consistently with additional human demonstrations and exhibits no saturation at 1,000 demonstrations per task, validating the effectiveness of scalable human tactile data for dexterous manipulation. Project page is available at https://uni-vta.github.io/.
Figures & tables
Fig. 2: Overview of the Tactile Motion-Capture System. A passive tactile exoskeleton captures hand motion and tactile signals. The integrated system further includes a VIVE tracker and a wrist-mounted RGB camera for capturing the wrist pose and wrist-centric images, respectively. The camera and VIVE tracker are calibrated to the Sharpa Wave wrist frame.
Fig. 3: Dataset Overview. Each task contains 150 robot demonstrations collected under limited variation, and 1,000 human demonstrations with diverse scenarios and contact patterns.
Fig. 4: Overview of Our Unified Visual-Tactile-Action Model. VTPM converts the visual-tactile input into a 20-D piezoresistive taxel token, while a Transformer extracts a 384-D visual token. The two tokens are concatenated into a unified 404-D interaction representation and passed to a diffusion action head and a tactile regression head. The action and tactile losses supervise their respective prediction heads, while their joint objective optimizes the shared Transformer to learn a unified tactile-aware visual representation.
Configuration
Flip Page
Light Bulb
Toggle Switch
Ball Classification
Liquid Transfer
Average
Baseline Methods
ViTacFormer [ 28 ]
0
4
0
0
0
1
RDP [ 43 ]
10
20
20
18
0
14
T-Rex [ 4 ]
43
26
24
36
15
29
Ablation Methods
Human Data only
0
0
0
0
0
0
Table 1: Comparison of Our Method with Baseline and Ablation Methods Across Five Tasks. Success rates (%) are computed over 10 rollouts per task, then averaged across tasks. The last row is our full model; the blocks above compare it against representative baselines and against ablations of human data only, future-tactile prediction, tactile input, and proprioceptive conditioning.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Setting
Visual observation
224×224 wrist RGB
Visual encoder
ViT-S/8, trained from scratch
Visual / tactile token
384-D / 20-D
Fused representation
404-D
Prediction horizon
16 steps
Action per step
9-D wrist + 22-D hand
Appendix
Table 2: Implementation Details. Settings shared by all task-specific UVTA policies.
Fig. 6: Real-Robot Rollout Trajectories. Representative successful executions of UVTA on five contact-rich tasks. From top to bottom: Flip Page, Screw Light Bulb, Toggle Switch, Ball Classification, and Liquid Transfer with a Dropper. Each row shows an execution sequence, with time progressing from left to right.
Task
Milestones and scores
Flip Page
0.3 : separate and lift the page edge; 0.7 : turn multiple pages together; 1.0 : turn exactly one page.
Screw Light Bulb
0.2 : establish stable contact with the bulb; 0.5 : rotate the bulb through one full turn; 1.0 : complete multiple turns until the bulb remains lit.
Toggle Switch
0.2 : establish stable contact; 0.6 : partially actuate the switch; 1.0 : push the switch fully to its on position.
Ball Classification
0.2 : securely grasp the ball; 0.6 : place it in a container; 1.0 : place it in the correct container for its material.
Liquid Transfer
0.2 : establish stable contact with the dropper; 0.5 : draw liquid into the dropper; 0.7 : lift and transport the filled dropper; 1.0 : dispense the liquid into the target container.
Appendix
Table 3: Task-Specific Evaluation Criteria. A rollout receives the highest score associated with a milestone it achieves; scores are not summed across milestones. A rollout that achieves none of the listed milestones receives 0. A score of 1.0 denotes full task success.
Human videos provide demonstrations of dexterous manipulation but lack robot-executable actions and tactile measurements. We present UniDex-ViTac, a framework that uses human-video-guided simulation to generate robot demonstrations paired with fingertip contact observations for training a deployable visuo-tactile policy. Object-specific residual reinforcement learning specialists adapt annotated human-object interaction references to a robotic arm-hand system. Their successful rollouts pair final robot action targets with robot-side fingertip contact observations. From 50 human demonstrations across ten objects, we collect 10,000 simulated trajectories to train a single Action Chunking with Transformers (ACT) based generalist. The policy combines point clouds, proprioception, and four binary contact signals encoded through fingertip labels and a separate token, without requiring human references or privileged object identity and pose at deployment. The contact-augmented configuration achieves 68.3% macro-average success in simulation, compared with 55.5% for the point-cloud-only baseline. Without real-robot demonstrations or policy fine-tuning, it succeeds in 73/110 physical trials (66.4%) across six seen and five unseen objects, compared with 60/110 (54.5%) for the baseline, an increase of 11.8 percentage points. These results support the feasibility of learning a unified visuo-tactile dexterous manipulation policy from video-guided simulated interactions. Project page: https://unidex-vitac.github.io/
Hyesung Lee, Si-Hwan Heo, Sungwook Yang
Center for Humanoid Research, Korea Institute of Science and Technology, Seoul 02792, Republic of Korea · Kim Jaechul Graduate School of AI, Korea Advanced Institute of Science and Technology, Seoul 02455, Republic of Korea
Human videos are an abundant source of dexterous manipulation behaviors, but they lack tactile information that is crucial for contact-rich interaction. This raises a fundamental question: can robots learn deployable visual-tactile dexterous manipulation policies from human video demonstrations without robot-side data collection? We present DEX-X, a framework for learning visual-tactile dexterous manipulation from human videos through simulation. Our key insight is that simulation can serve as a tactile completion engine. Given monocular human demonstrations, DEX-X reconstructs hand-object interactions in simulation, where physically grounded contact dynamics provide tactile supervision unavailable in the original videos. Leveraging this recovered tactile information, we train visual-tactile dexterous manipulation policies and distill them into deployable policies operating on point-cloud observations and tactile sensing. We demonstrate zero-shot sim-to-real transfer on a dexterous hand-arm platform across diverse grasping and contact-rich tool-use tasks. The teacher policy achieves 65.9% average success across six task categories in simulation, while the distilled visual-tactile policy achieves 93% success on real-world cube picking and 53% on the challenging table-cleaning task. Zero-shot generalization to unseen object geometries is also observed on object-picking tasks. Our results suggest that simulated interaction is a key bridge between human videos and deployable dexterous manipulation policies, providing the missing physical supervision needed for scalable robot skill learning from Internet-scale human video data.
Ruoqu Chen, Feixiang Ruan, Liu Cao +9
Tsinghua University · Shanghai Qizhi Institute · Sharpa +2
Dexterous manipulation requires coordinated multi-finger control and effective tactile feedback, yet learning these capabilities remains challenging due to the lack of large-scale real-world data and the difficulty of extracting effective representations from sparse tactile signals. We build a robot platform and teleoperation system to collect a 200-hour bimanual dexterous manipulation dataset with synchronized visual, tactile, and language annotations, comprising 10,576 trajectories across 65 tasks, 69.5% of which involve dexterous multi-finger manipulation. We further propose STAR, an integrated training recipe for vision-tactile-language-action (VTLA) models that addresses the spatial, temporal, and informational sparsity of tactile signals through visual-tactile joint pre-training, sparse-global tactile token representation, and sparse future tactile prediction. Trained on this dataset, STAR achieves a 61% average success rate across four real-world tasks with 100 post-training trajectories per task, demonstrating dexterous performance under task-specific post-training.