Disambiguate Gripper State in Grasp-Based Tasks: Pseudo-Tactile as Feedback Enables Pure Simulation Learning
Organizations: State Key Laboratory of Industrial Control and Technology, Zhejiang University, Hangzhou 310027, China · Huawei
Abstract
Grasp-based manipulation tasks are fundamental to robots interacting with their environments, yet gripper state ambiguity significantly reduces the robustness of imitation learning policies for these tasks. Data-driven solutions face the challenge of high real-world data costs, while simulation data, despite its low costs, is limited by the sim-to-real gap. We identify the root cause of gripper state ambiguity as the lack of tactile feedback. To address this, we propose a novel approach employing pseudo-tactile as feedback, inspired by the idea of using a force-controlled gripper as a tactile sensor. This method enhances policy robustness without additional data collection and hardware involvement, while providing a noise-free binary gripper state observation for the policy and thus facilitating pure simulation learning to unleash the power of simulation. Experimental results across three real-world grasp-based tasks demonstrate the necessity, effectiveness, and efficiency of our approach.
Figures & tables
| SR-ND (%) | SR-D (%) | SR (%) | AT (s) | SR-R (%) | GSR-ND (%) | GSR-D (%) | GSR (%) | |
|---|---|---|---|---|---|---|---|---|
| DP-Real | 20 | 0 | 10 | 18.5 | 10 | 20 | 0 | 10 |
| RDT-1B-Real | 30 | 10 | 20 | 27.0 | 70 | 40 | 10 | 25 |
| Ours | 100 | 80 | 90 | 16.3 | 100 | 100 | 80 | 90 |
| SR-ND (%) | SR-D (%) | SR (%) | AT (s) | SR-R (%) | GSR-ND (%) | GSR-D (%) | GSR (%) | |
|---|---|---|---|---|---|---|---|---|
| ACT-Real | 0 | 0 | 0 | - | 30 | 0 | 0 | 0 |
| DP-Real | 50 | 0 | 25 | 24.6 | 20 | 50 | 0 | 25 |
| RDT-1B-Real | 50 | 30 | 40 | 28.0 | 30 | 50 | 30 | 40 |
| Ours | 100 | 80 | 90 | 18.0 | 100 | 100 | 80 | 90 |
| SR-ND (%) | SR-D (%) | SR (%) | AT (s) | SR-R (%) | GSR-ND (%) | GSR-D (%) | GSR (%) | |
|---|---|---|---|---|---|---|---|---|
| ACT-Real | 50 | 20 | 35 | 34.7 | 30 | 50 | 30 | 40 |
| DP-Real | 60 | 20 | 40 | 68.5 | 30 | 60 | 20 | 40 |
| RDT-1B-Real | 60 | 30 | 45 | 44.4 | 60 | 70 | 30 | 50 |
| Ours | 70 | 90 | 80 | 26.7 | 100 | 70 | 90 | 80 |
| No. | policy architecture | sim data with gripper randomization | pseudo-tactile feedback | SR-ND(%) | SR-D(%) | SR(%) | AT(s) | SR-R(%) | GSR-ND(%) | GSR-D(%) | GSR(%) |
| 1 | DP | ✗ | ✗ | 50 | 10 | 30 | 19.7 | 30 | 50 | 10 | 30 |
| 2 | DP | ✓ | ✗ | 0 | 0 | 0 | - | 100 | 70 | 80 | 75 |
| 3 | DP | ✗ | ✓ | 70 | 90 | 80 | 26.7 | 100 | 70 | 90 | 80 |
| 4 | ACT | ✓ | ✗ | 40 | 40 | 40 | 20.1 | 100 | 40 | 70 | 55 |
| 5 | RDT-1B | ✓ | ✗ | 0 | 0 | 0 | - | 100 | 50 | 60 | 55 |