OpenViTac: Learning and Benchmarking Visuo-Tactile Policies in a Unified Sim-and-Real Framework
Organizations: Fudan University · Shanghai Key Laboratory of Multimodal Embodied AI · Hefei University of Technology · National University of Singapore · China Unicom · Neote AI
Abstract
Tactile feedback provides embodied agents with physical information beyond visual observations, enabling more reliable interaction with the real world. However, despite the rapid progress of vision-tactile-language-action (VTLA) policies, there remains a lack of unified benchmarks for evaluating tactile-enabled robot manipulation across simulation and the real world. To address this gap, we introduce OpenViTac, a visuo-tactile manipulation benchmark for evaluating robot policies across simulation and the real world. OpenViTac organizes contact-rich manipulation into four tactile-relevant capability dimensions and provides paired simulation-real-world settings for consistent evaluation of VLA, WAM, and VTLA policies. Building upon this benchmark, we investigate how different tactile representations and integration strategies affect the performance of pretrained VLA models. Correspondingly, we introduce OpenVTLA, a tactile augmentation framework that combines the best-performing representation and integration strategy. Furthermore, we leverage the paired benchmark setting to study sim-real co-training and analyze factors affecting cross-domain policy learning. Together, OpenViTac provides a unified platform for evaluating and advancing visuo-tactile robot manipulation.
Figures & tables
| Benchmark | Gripper Types | Tactile Sensing | Tactile Sensor Types | Property Perception | Fragility-/Force- Aware | Precision Manipulation | Scene Diversity |
| CALVIN [ 25 ] | 1 | ✓ | 1 | ✗ | ✗ | ✗ | ✗ |
| LIBERO [ 20 ] | 1 | ✗ | - | ✗ | ✗ | ✗ | ✗ |
| ManiFeel [ 22 ] | 1 | ✓ | 1 | ✗ | ✗ | ✗ | ✗ |
| Tabero [ 37 ] | 1 | ✓ | 1 | ✗ | ✓ | ✗ | ✗ |
| UniVTAC [ 3 ] | 1 | ✓ | 1 | ✓ | ✗ | ✓ | ✗ |
| NeoSim [ 31 ] | 1 | ✓ | 1 | ✗ | ✓ | ✓ | ✗ |
| Models | PP | FA | CR | PR | Avg. | ||||||||
| WC | HC | RC | RR | ECS | GC | GA | PD | In-USB | In-B v1 | In-B v2 | |||
| VLAs | 53 | 60 | 64 | 52 | 48 | 39 | 87 | 99 | 14 | 9 | 9.3 | 48.6 | |
| StarVLA (Qwen3) | 44 | 38 | 42 | 35 | 32 | 10 | 26 | 35 | 4 | 0 | 0 | 24.2 | |
| InternVLA-A1.5 | 53 | 76 | 65 | 77 | 47 | 53 | 79 | 98 | 3 | 6 | 9.3 | 51.5 | |
| XVLA | 28 | 30 | 56 | 47 | 45 | 42 | 84 | 76 | 2 | 3 | 0.7 | 37.6 | |
| Lingbot-VLA-2 | 52 | 58 | 58 | 53 | 52 | 41 | 78 | 92 | 8 | 5 | 2.7 | 45.4 | |
| Models | PP | FA | CR | PR | Avg. | |||||
| RC | ECS | GC | GA | PD | In-USB | In-B v1 | In-B v2 | |||
| VLAs | 10/20 | 10/20 | 16/20 | 7/20 | 12/20 | 3/20 | 5/20 | 4/30 | 41.0 | |
| InternVLA-A1.5 | 10/20 | 10/20 | 17/20 | 9/20 | 13/20 | 3/20 | 3/20 | 6/30 | 43.1 | |
| Lingbot-VLA-2 | 10/20 | 10/20 | 15/20 | 3/20 | 10/20 | 1/20 | 2/20 | 5/30 | 34.0 | |
| Xiaomi-Robotics | 10/20 | 10/20 | 16/20 | 5/20 | 8/20 | 1/20 | 3/20 | 2/30 | 34.0 | |
| WAMs | FastWAM | 8/20 | 9/20 | 12/20 | 5/20 | 10/20 | 0/20 | 1/20 | 2/30 | 29.0 |
| Method | Representation | Integration | PP | FA | CR | PR | Avg. |
| OpenVTLA | AnyTouch2 | ① Concat | 90.2 | 69.0 | 95.0 | 15.3 | 68.7 |
| Sparsh | ② Modulation | 82.2 | 49.0 | 73.0 | 5.0 | 56.5 | |
| AnyTouch2 | ③ Action | 70.8 | 43.0 | 62.0 | 4.7 | 48.6 | |
| FTP-1 | T3 | ④ Expert | 77.8 | 60.0 | 88.5 | 15.8 | 61.2 |
| AnyTouch2 | ⑤ Refine | 63.0 | 51.0 | 81.0 | 8.7 | 50.4 |
| Model | Sim. | Std. Lighting | Lighting Perturb. | ||
| In-USB | GA | In-USB | GA | ||
| 0 | 3/20 | 7/20 | 0/20 | 2/20 | |
| 250 | 5/20 | 9/20 | – | – | |
| 500 | 8/20 | 10/20 | 4/20 | 8/20 | |
| OpenVTLA | 0 | 4/20 | 10/20 | 2/20 | 10/20 |
| 250 | 6/20 | 10/20 | – | – | |
Appendix figures & tables27 assets
Supplementary material from the paper’s appendix.
Appendix
| Category | Abbr. | Task | Capability under evaluation |
| Physical Property Perception | WC | Weight Classify | Infer object weight through interaction and place the object at the corresponding target. |
| HC | Hardness Classify | Infer object hardness through contact and place the object at the corresponding target. | |
| RC | Roughness Classify | Infer surface roughness through contact and place the object at the corresponding target. | |
| RR | Roughness-Guided Regrasp | Infer relative roughness and keep or regrasp the appropriate object accordingly. | |
| ECS | Empty-Can Selection | Distinguish visually similar empty and filled cans from their physical properties and select the instructed one. | |
| Fragility-Aware Manipulation | GC | Grasp Chip | Grasp and place a fragile chip while avoiding damage throughout the manipulation. |
| Task | Task Objective | Success Condition |
| Weight Classify (WC) | Infer whether a block is light or heavy through interaction and place it on the corresponding target. | The block is correctly classified and placed on the corresponding target plate. |
| Hardness Classify (HC) | Infer whether a block is soft or hard through contact interaction and place it on the corresponding target. | The block is correctly classified and placed on the corresponding target plate. |
| Roughness Classify (RC) | Infer whether a block is smooth or rough through contact interaction and place it on the corresponding target. | The block is correctly classified and placed on the corresponding target plate. |
| Empty-Can Selection (ECS) | Identify the empty can among visually similar cans through physical interaction and place it into the basket. | The empty can is correctly identified and successfully placed into the basket. |
| Roughness-Guided Regrasp (RR) | Determine whether to retain the initially grasped block based on roughness and regrasp when necessary. | The rough block is correctly selected and placed on the target plate. |
| Grasp Chip (GC) | Grasp and transport a fragile chip while avoiding damage during manipulation. | The chip is successfully grasped, transported, and placed without damage while satisfying physical safety constraints. |
| Models | Steps | Global batch | |
| VLAs | 20k | 64 | |
| StarVLA (Qwen3) | 20k | 64 | |
| InternVLA-A1.5 | 40k | 32 | |
| XVLA | 30k | 64 | |
| Lingbot-VLA-2 | 20k | 64 | |
| Xiaomi-Robotics | 20k | 64 | |
| Representation | PP | FA | CR | PR | Avg. |
| AnyTouch2 (Random Init.) | 78.5 | 34.0 | 82.0 | 10.2 | 56.5 |
| AnyTouch2 (Pretrained) | 90.2 | 69.0 | 95.0 | 15.3 | 68.7 |
| Temporal Input | PP | FA | CR | PR | Avg. |
| Repeated Current Frame | 79.0 | 52.0 | 90.0 | 5.2 | 58.4 |
| Tactile History | 90.2 | 69.0 | 95.0 | 15.3 | 68.7 |