Contact-rich manipulation requires robots to sequence precise contacts, maintain stable grasps, and apply directed forces. Reinforcement learning (RL) can acquire such behaviors automatically, but its performance hinges on reward design: sparse rewards reduce the learning efficiency, while dense rewards are hard to specify. Visual reward learning addresses this by inferring rewards from action-free demonstrations. Because it conditions only on visual observations, it fails to capture rewards beyond visual goals. We propose Tactile Reward Learning (TaRL), a framework that learns rewards from tactile demonstrations. TaRL takes a sequence of tactile deformation maps as input, and regresses task-completion progress from both successful and failed demonstrations. Because TaRL captures local robot-object interaction, it provides informative feedback to learn firm grasps and correctly directed forces; meanwhile, it is robust to changes in scene layout such as object position. We evaluate TaRL on four manipulation tasks in simulation and two in the real world. Used as a shaping reward, it substantially improves both sample efficiency and final success rate, raising success on Nut threading from 34% to 56% in simulation and on cube pickup from 37% to 97% in the real world. Combining tactile with visual rewards improves performance further. TaRL also generalizes across object instances: trained on box placement and directly deployed to can placement, it significantly improves policy learning on the new task. Project page is available at https://embodiedai-ntu.github.io/tarl.
Figures & tables
Figure 2: Overview of TaRL. Left: It is trained with successful, rewound and failed demonstrations. Middle: It conditions on tactile deformation maps and outputs task-progress rewards. Right: It is applied to downstream RL, offering dense rewards.
Figure 3: Task suites in simulation and the real world. Four contact-rich manipulation tasks in simulation (a–d) and two real-world tasks (e, f) that mirror the simulated peg insertion and cube pickup.
Figure 4: Downstream RL efficiency. We evaluate RL policies trained with and without TaRL in simulation. TaRL substantially enhances the training efficiency and task success.
Figure 5: Generalization to unseen object positions. We train TaRL on a single position, and deploy it to seen (ID) and unseen (OOD) object positions. We report the predicted reward at each time step, averaged within each configuration. TaRL is robust to OOD object positions.
Figure 6: Comparison of tactile and visual features. We extract features with the CNN of ReWiND and TaRL, each taking visual and tactile observations as input, on peg insertion, gear assembly and nut threading tasks. With t-SNE, features are transformed into 2-dim vectors and plotted in 2D. Dashed ellipses denote covariances of their distributions. Visual features extracted from in-domain (ID) object positions follow distinct distributions from those of out-of-distribution (OOD) positions. In contrast, tactile features in both cases share similar distributions.
Figure 7: Evaluation of visual rewards. Top: Visual reward quality on ID and OOD object positions. Bottom: Downstream RL efficiency of visual rewards and the combination of tactile and visual rewards.
Figure 8: Rewards contributed at individual contact phases. We show the visual and tactile rewards predicted by ReWiND and TaRL on a trajectory of gear assembly task. Visual rewards plateau once the robot reach the target object’s position. Tactile rewards vary across different contact phases.
Figure 9: Zero-shot generalization to unseen object instances. We deploy TaRL trained on box placement to can placement. TaRL achieves high reward quality and improves downstream RL efficiency on the new task.
Figure 10: The visual and tactile reward model’s generalization to unseen object positions. We show predicted tactile and visual rewards along successful rollouts at the training position (left) and an unseen position (right). The single-position visual model collapses at the unseen position while TaRL does not.
Figure 11: Comparison of handcrafted tactile reward and TaRL.
Figure 12: Tactile-based RL efficiency. We train a tactile-based RL policy with actors and critics both taking tactile observations as inputs. For the tactile-based (state-based) approach, we evaluate its downstream RL efficiency with ( with ) and without ( without ) TaRL on gear assembly, peg insertion and nut threading tasks in simulation. Using TaRL enhances training efficiency and task success in both cases.
Figure 13: Real-world RL efficiency. We compare downstream RL efficiency with and without TaRL in the real world. From left to right, we show the success rate during inference on cube pickup, peg pickup and peg insertion tasks. Policies optimized with both basic sparse rewards and tactile rewards substantially outperform those with only sparse rewards.
Vision-language-action (VLA) models provide strong visual, language, and action priors for robot manipulation, but visual observations alone often miss the local contact state required for contact-rich tasks. We present TacCoRL, a scalable framework that injects Tactile feedback into VLA policies and improves them through sim-real Co-training and simulation-based reinforcement learning (RL), without requiring large-scale tactile pretraining or extensive real-world contact exploration. The key idea is not only adding touch as an input, but learning how contact readings should modulate action responses in near-failure states that are rare in demonstrations and risky to collect on hardware. We use a real-aligned simulator as a closed-loop training environment for contact interaction. Mixed simulated and real trajectories first warm-start tactile-conditioned actions in the pretrained policy. Reinforcement learning with verifiable task rewards then optimizes the policy using simulated contact rollouts. It reinforces tactile-conditioned actions that lead to task completion, while a supervised objective on real trajectories keeps the refined policy anchored to deployment visual, tactile, and action distributions. The resulting policy transfers directly to the real robot without privileged simulation state or online real-world RL. Across four bimanual contact-rich tasks, the final visuo-tactile policy achieves an average success rate of 72.5%, compared to baseline of 50.0%. Result videos and more details are available at https://tac-corl.github.io/
Siyu Ma, Yuqi Liang, Chang Yu +5
University of California, Los Angeles · University of California, San Diego · 4Peking University +1
Mastering robot manipulation skills via reinforcement learning (RL) remains largely sample-inefficient. The most common RL algorithms rely on random action sampling to discover new strategies, resulting in agents that allocate most of their training budget to motions in free space, away from the contacts from which manipulation skills emerge. Existing intrinsic motivation methods based on model disagreement or epistemic uncertainty improve on isotropic noise, but they can also reward uncertainty in functionally irrelevant transitions, such as erratic motions in free space. In this work, we argue that tactile feedback provides a natural signal for exploration, and introduce TacEx, a framework that incorporates touch into epistemic uncertainty-driven exploration by decomposing model uncertainty across sensory modalities and directing curiosity toward the tactile channel. By anchoring curiosity to the sense of touch, TacEx drives the robot to discover complex contact dynamics, learning to manipulate and grasp objects without task rewards or expert demonstrations during exploration. The interaction-dense dataset collected through this tactile-driven curiosity supports offline learning of downstream pick-and-place policies without additional environment interaction. We further use tactile-driven exploration to post-train vision-language-action (VLA) models. Although the VLAs are initially pre-trained without tactile feedback, post-training with TacEx substantially improves downstream performance while remaining highly sample-efficient.
Klemens Iten, Alexander Proshkin, Bhavya Sukhija +4
University of California, Berkeley United States · The University of Texas at Austin United States
Vision-Language-Action (VLA) models have become a powerful framework for robotic manipulation, and recent studies have introduced tactile or force feedback into VLAs to address contact-rich tasks. However, these models are typically deployed as offline policies. When contact conditions shift from the training distribution, the policy cannot perform online adaptation, leading to problems such as inappropriate contact forces and inefficient retries. Therefore, we propose TORL-VLA, a tactile-guided online reinforcement learning framework that couples tactile feedback with policy refinement for contact-rich manipulation. Our method introduces a tactile-derived wrench-aware VLA to predict reference actions and future wrench sequences, while a lightweight online RL module is used to refine the reference actions. To stabilize learning from mixed exploratory policy-generated and human-intervention data, we introduce an intervention-censored critic that prevents post-intervention success from being wrongly credited to policy-generated actions preceding intervention. Real-robot experiments on long-horizon contact-rich tasks, including latch manipulation, coffee-cup placement, and egg handling, show that TORL-VLA improves success rates at both subtask and full-task levels, as well as time-bounded execution efficiency over strong baselines. Project page: https://torl-vla.github.io/
Huaihang Zheng, Yi Yang, Kai Ma +8
Meituan · State Key Lab of Multimodal Artificial Intelligence Systems, Institute of Automation, CAS · Beijing Institute of Technology +2