Unified Visual-Tactile-Action Modeling from Human Demonstrations for Dexterous Manipulation
Organizations: Shanghai Jiao Tong University · Sharpa Robotics · Beijing Institute of Technology · State Key Laboratory of General Artificial Intelligence, BIGAI
Abstract
Dexterous manipulation requires tactile feedback. However, robot tactile demonstrations are difficult to scale,because dexterous-hand teleoperation provides limited tactile feedback to the operator. In contrast, human demonstrations offer a substantially more scalable source of diverse tactile interactions. Motivated by a simple premise: hands can change, but the underlying physics of interaction does not. We leverage human tactile data to improve dexterous manipulation policies. Specifically, we first build a tactile motion-capture system that synchronously records images, tactile signals, and hand motions. Using this system, we construct the UVTA dataset spanning five contact-rich tasks, with 1,000 human demonstrations covering diverse interaction patterns and 150 robot demonstrations per task. To transfer the underlying physics of human interaction to robot control, we propose a Unified Visual-Tactile-Action Model that maps both embodiments into aligned tactile and action representations and jointly predicts future action and tactile trajectories. The joint objective enables human demonstrations to supervise contact-aware representation learning, while only robot actions are executed during deployment. In real-robot evaluations across five tasks, our method achieves an average success rate of 70%, outperforming the strongest visual-tactile baseline, which achieves 29%, and an architecture ablation, which achieves 42%. Performance improves consistently with additional human demonstrations and exhibits no saturation at 1,000 demonstrations per task, validating the effectiveness of scalable human tactile data for dexterous manipulation. Project page is available at https://uni-vta.github.io/.
Figures & tables
| Configuration | Flip Page | Light Bulb | Toggle Switch | Ball Classification | Liquid Transfer | Average |
|---|---|---|---|---|---|---|
| Baseline Methods | ||||||
| ViTacFormer [ 28 ] | 0 | 4 | 0 | 0 | 0 | 1 |
| RDP [ 43 ] | 10 | 20 | 20 | 18 | 0 | 14 |
| T-Rex [ 4 ] | 43 | 26 | 24 | 36 | 15 | 29 |
| Ablation Methods | ||||||
| Human Data only | 0 | 0 | 0 | 0 | 0 | 0 |
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
| Component | Setting |
|---|---|
| Visual observation | wrist RGB |
| Visual encoder | ViT-S/8, trained from scratch |
| Visual / tactile token | 384-D / 20-D |
| Fused representation | 404-D |
| Prediction horizon | 16 steps |
| Action per step | 9-D wrist + 22-D hand |
| Task | Milestones and scores |
|---|---|
| Flip Page | 0.3 : separate and lift the page edge; 0.7 : turn multiple pages together; 1.0 : turn exactly one page. |
| Screw Light Bulb | 0.2 : establish stable contact with the bulb; 0.5 : rotate the bulb through one full turn; 1.0 : complete multiple turns until the bulb remains lit. |
| Toggle Switch | 0.2 : establish stable contact; 0.6 : partially actuate the switch; 1.0 : push the switch fully to its on position. |
| Ball Classification | 0.2 : securely grasp the ball; 0.6 : place it in a container; 1.0 : place it in the correct container for its material. |
| Liquid Transfer | 0.2 : establish stable contact with the dropper; 0.5 : draw liquid into the dropper; 0.7 : lift and transport the filled dropper; 1.0 : dispense the liquid into the target container. |