Contact-rich manipulation requires robots to sequence precise contacts, maintain stable grasps, and apply directed forces. Reinforcement learning (RL) can acquire such behaviors automatically, but its performance hinges on reward design: sparse rewards reduce the learning efficiency, while dense rewards are hard to specify. Visual reward learning addresses this by inferring rewards from action-free demonstrations. Because it conditions only on visual observations, it fails to capture rewards beyond visual goals. We propose Tactile Reward Learning (TaRL), a framework that learns rewards from tactile demonstrations. TaRL takes a sequence of tactile deformation maps as input, and regresses task-completion progress from both successful and failed demonstrations. Because TaRL captures local robot-object interaction, it provides informative feedback to learn firm grasps and correctly directed forces; meanwhile, it is robust to changes in scene layout such as object position. We evaluate TaRL on four manipulation tasks in simulation and two in the real world. Used as a shaping reward, it substantially improves both sample efficiency and final success rate, raising success on Nut threading from 34% to 56% in simulation and on cube pickup from 37% to 97% in the real world. Combining tactile with visual rewards improves performance further. TaRL also generalizes across object instances: trained on box placement and directly deployed to can placement, it significantly improves policy learning on the new task. Project page is available at https://embodiedai-ntu.github.io/tarl.
Figures & tables
Figure 2: Overview of TaRL. Left: It is trained with successful, rewound and failed demonstrations. Middle: It conditions on tactile deformation maps and outputs task-progress rewards. Right: It is applied to downstream RL, offering dense rewards.
Figure 3: Task suites in simulation and the real world. Four contact-rich manipulation tasks in simulation (a–d) and two real-world tasks (e, f) that mirror the simulated peg insertion and cube pickup.
Figure 4: Downstream RL efficiency. We evaluate RL policies trained with and without TaRL in simulation. TaRL substantially enhances the training efficiency and task success.
Figure 5: Generalization to unseen object positions. We train TaRL on a single position, and deploy it to seen (ID) and unseen (OOD) object positions. We report the predicted reward at each time step, averaged within each configuration. TaRL is robust to OOD object positions.
Figure 6: Comparison of tactile and visual features. We extract features with the CNN of ReWiND and TaRL, each taking visual and tactile observations as input, on peg insertion, gear assembly and nut threading tasks. With t-SNE, features are transformed into 2-dim vectors and plotted in 2D. Dashed ellipses denote covariances of their distributions. Visual features extracted from in-domain (ID) object positions follow distinct distributions from those of out-of-distribution (OOD) positions. In contrast, tactile features in both cases share similar distributions.
Figure 7: Evaluation of visual rewards. Top: Visual reward quality on ID and OOD object positions. Bottom: Downstream RL efficiency of visual rewards and the combination of tactile and visual rewards.
Figure 8: Rewards contributed at individual contact phases. We show the visual and tactile rewards predicted by ReWiND and TaRL on a trajectory of gear assembly task. Visual rewards plateau once the robot reach the target object’s position. Tactile rewards vary across different contact phases.
Figure 9: Zero-shot generalization to unseen object instances. We deploy TaRL trained on box placement to can placement. TaRL achieves high reward quality and improves downstream RL efficiency on the new task.
Figure 10: The visual and tactile reward model’s generalization to unseen object positions. We show predicted tactile and visual rewards along successful rollouts at the training position (left) and an unseen position (right). The single-position visual model collapses at the unseen position while TaRL does not.
Figure 11: Comparison of handcrafted tactile reward and TaRL.
Figure 12: Tactile-based RL efficiency. We train a tactile-based RL policy with actors and critics both taking tactile observations as inputs. For the tactile-based (state-based) approach, we evaluate its downstream RL efficiency with ( with ) and without ( without ) TaRL on gear assembly, peg insertion and nut threading tasks in simulation. Using TaRL enhances training efficiency and task success in both cases.
Figure 13: Real-world RL efficiency. We compare downstream RL efficiency with and without TaRL in the real world. From left to right, we show the success rate during inference on cube pickup, peg pickup and peg insertion tasks. Policies optimized with both basic sparse rewards and tactile rewards substantially outperform those with only sparse rewards.