Mastering robot manipulation skills via reinforcement learning (RL) remains largely sample-inefficient. The most common RL algorithms rely on random action sampling to discover new strategies, resulting in agents that allocate most of their training budget to motions in free space, away from the contacts from which manipulation skills emerge. Existing intrinsic motivation methods based on model disagreement or epistemic uncertainty improve on isotropic noise, but they can also reward uncertainty in functionally irrelevant transitions, such as erratic motions in free space. In this work, we argue that tactile feedback provides a natural signal for exploration, and introduce TacEx, a framework that incorporates touch into epistemic uncertainty-driven exploration by decomposing model uncertainty across sensory modalities and directing curiosity toward the tactile channel. By anchoring curiosity to the sense of touch, TacEx drives the robot to discover complex contact dynamics, learning to manipulate and grasp objects without task rewards or expert demonstrations during exploration. The interaction-dense dataset collected through this tactile-driven curiosity supports offline learning of downstream pick-and-place policies without additional environment interaction. We further use tactile-driven exploration to post-train vision-language-action (VLA) models. Although the VLAs are initially pre-trained without tactile feedback, post-training with TacEx substantially improves downstream performance while remaining highly sample-efficient.
Figures & tables
Figure 1: Overview of TacEx . The state s concatenates the fused latent z , a visual embedding v , and a tactile embedding ι , where v and ι are produced by separate CNN encoders from RGB and tactile force maps. At each step, the policy π(a∣s) acts in the environment, and the resulting transitions train an ensemble environment model that predicts the next state s′ . The model’s epistemic uncertainty decomposes per modality into σnz , σnv , and σnι ; unlike standard MaxInfoRL [ 31 ] , which rewards the agent for total predictive uncertainty, TacEx explicitly weights the tactile component σnι(s,a) in the exploration bonus. This anchors curiosity to the sense of touch, driving the agent toward contact-rich interactions and away from uncertainty in functionally irrelevant regions and transitions.
Figure 2: Tactile-driven exploration with a single object. Reward-free exploration under six weighting configurations. (Left) object-interaction frequency during exploration; (right) offline return on the downstream pick-and-place task, with DrQ as an online RL baseline. All results are across 5 seeds; we report the mean and standard error.
Figure 3: Tactile-driven exploration with multiple objects. The setup of fig. 2 extended to four objects. (Left) object-interaction frequency; (center) offline pick-and-place performance; (right) interaction diversity, measured as the effective number of objects grasped 2H . All results are across 5 seeds; we report the mean and standard error.
Figure 4: Online diffusion-steering performance on contact-rich LIBERO-90 tasks. DSRL-SAC uses visual and proprioceptive state only and optimizes sparse task reward with SAC. DSRL-TacEx additionally receives tactile observations and uses ensemble disagreement weightings (ωv,ωι,ωz) over tactile (0,1,0) , vision+tactile (1,1,0) , or latent+tactile (0,1,1) predictions as an intrinsic exploration bonus. Shaded regions denote standard error across 10 random seeds.
Table 1: Main hyperparameters used for the LIBERO simulation experiments. Values are shared by the SAC baseline and MaxInfo-SAC ablations unless marked otherwise.
Figure 5: Ablations on LIBERO task 58. Black curves are SAC diffusion-steering baselines, colored curves add TacEx -style exploration. In the top row, solid curves use tactile observations in the learner state, dashed curves remove tactile observations from the learner state. In the bottom-left panel, dashed curves remove tactile modalities from the intrinsic reward. The dotted curve shows the explore-then-exploit schedule for comparison. Adding tactile observations improves diffusion steering, but the best-performing exploration variants are those that also include tactile prediction disagreement in the intrinsic bonus. We report the average evaluation return over 10 random seeds along with standard error bands
Figure 6: Additional diffusion-policy comparisons on five LIBERO tasks. Final return is the time- weighted mean success rate over the final 10% of the logged training horizon; normalized AUC is the integral over the full horizon divided by that horizon. Points and error bars report means and standard errors across ten seeds per task. Circles and squares denote bonuses with and without explicit tactile prediction targets. Summed-taxel controls use the aggregate signal as an intrinsic reward or as a prediction target for information gain. DSRL-PPO steers a frozen policy using task reward; the DPPO adaptation updates the action-output head using tactile disagreement.