Organizations: Amazon Fulfillment Technologies & Robotics, Westborough, MA, USA. · Department of Electrical and Computer Engineering, The University of Texas at Austin, Austin, TX, USA.
Tactile sim-to-real learning must bridge simulated contact and device-specific sensor responses while preserving information needed for control. We propose a factorized tactile representation and control framework that maps normal force and contact patch to an effective contact response recoverable from sensor readings. The response is separated into contact geometry, force distribution, and temporal contact change, with representation-specific encoding and randomization. A Tactile Gated Policy preserves these representations separately through control and operates over all mask configurations without retraining. We evaluate the approach through response reconstruction, spatial alignment, force regulation, and contact-rich adversarial peg insertion in simulation and the real world, enabling the utility and transfer reliability of different tactile representations to be assessed independently. The approach achieves <1 mm contact localization, 1.69 N force-tracking error on unseen geometries, and a 35% improvement in real-world adversarial peg insertion over the unfactorized response, with different tactile representations benefiting different interactions.
Figures & tables
Fig. 2 : The response pipeline in both directions on a single calibration contact. From physical deformation on a real sensor or simulated compliant contact with penetration, force and contact patch produce a spread pressure field via an anisotropic Gaussian kernel, which is nonlinearly aggregated into the per-taxel effective contact response r. The forward path multiplies by per-taxel gain β to predict raw counts; the inverse path divides measured counts by β to recover r, separating the shared nonlinear contact response from unit-specific electronic gain.
Fig. 3 : Tactile Gated Policy. Separate encoders produce geometry, force, and temporal-change latents. The mask m gates these latents before concatenation with the encoded robot state and mask itself as input to the LSTM actor.
Fig. 4 : Contact data collection and pad-contact validation. Localized pokes use a wrist F/T reference; edge and whole-pad grasps use the load-cell rig. Separate grasping tests with flat punches (diameters 4 and 6mm ) and spheres (radii 5 and 20mm ) validate the simulated compliant pad-contact model. The simulated rig matches the physical geometry and loading; no direct physical contact-patch measurement is assumed.
Model
Poke (L/R)
Edge (L/R)
Baseline comparisons
Spatial holdout
Proposed r
18.7 / 18.1
28.6 / 30.9
MLP encoder
13.3 / 13.9
83.5 / 63.6
CNN encoder
15.1 / 13.9
75.1 / 52.8
Kernel ridge
59.0 / 53.8
140.3 / 119.4
TABLE I : Median per-contact taxel RMSE in counts (left/right pad; lower is better). Spatial holdout fits on training pokes, selected edge lines, and whole-pad data. Family holdout excludes all edge data from fitting and model selection. Both edge evaluations use the same 20 held-out positions. Ablations use the spatial-holdout protocol. Methods within each block use matched fitting and evaluation splits.
Fig. 5 : The effective response bridges physical tactile measurement while retaining task-relevant contact information. (a) measured and modeled count patterns for a localized poke, edge, and distributed pad contacts. (b) held-out force error by true force bin, in total force and per activated taxel; error stays low inside the identified operating range and rises as the sensor approaches compression. (c) contact localization recovered from counts on a held-out fine scan, substantially finer than the taxel pitch.
Fig. 6 : Controlled tactile manipulation. (a) Tactile Alignment: the policy aligns a cylinder with an unknown in-hand pose. Hardware evaluation records the deepest level reached in a fixture with 6, 4, 2, and 1mm radial clearance. (b) Grasp-Force Regulation: the policy commands gripper increments during a scripted hold-and-lift motion. The simulated object’s two plates are coupled by a prismatic spring with randomized stiffness. An independent load cell supplies the hardware force reference. Scale bars: 2cm .
Sim. [mm/ ∘ ]
Real [success/10]
Input
Interp.
Out-of-range
1 mm ( r10/r18 )
2 mm ( r10/r18 )
Prop.
14.31/20.35
13.77/19.67
0/0
0/0
G
0.62/0.65
1.06/1.49
9/7
10/10
F
1.83/1.97
2.50/2.74
3/3
9/8
C
9.91/14.25
9.77/14.04
0/0
0/0
GF
0.69/0.67
0.98/1.23
5/5
10/10
TABLE II : Tactile Alignment. Simulation reports terminal position/orientation error [mm/ ∘ ] on unseen interpolation and out-of-range radii (100 held-out resets per radius). Real-world reports successful 1mm - and 2mm -clearance trials out of 10 for r10 / r18 , with r18 unseen during training. Taxel-line is an oracle-calibrated discrete taxel readout; its simulation entries are recovered in-hand pose RMSE rather than terminal policy error.
Sim. MAE [N]
Real MAE [N]
Input
≤15 N
15–30 N
≤15 N
Prop.
1.61
3.87
6.09±1.03
G
1.66
2.26
3.80±1.02
F
0.95
1.33
1.91±0.47
C
5.48
6.13
11.78±13.25
GF
1.15
1.50
1.69±0.25
TABLE III : Grasp-Force Regulation. Simulation reports error on 50 unseen test meshes; real-world evaluation reports pointwise settle MAE as mean ± SD across five printed test shapes. Reference uses privileged true force in simulation and the best fixed gripper setpoint selected with load-cell hindsight on hardware. Bold highlights the consistently strong F-containing gates.
Fig. 7 : Adversarial Peg Insertion. (a) The square peg must enter a tilted bore and pass two internal protrusions before becoming fully seated; grasp height, offset, and in-hand tilt vary between trials. (b) Cross-section of the tilted bore and internal protrusions. (c) The fixture is evaluated at two positions 5mm apart and three yaw angles ( −15∘ , 0∘ , +15∘ ), yielding six poses.
Representation Baselines
Input
Sim. [%]
Real [%]
Prop.
52.6
5/30 (16.7)
Binary
59.9
9/30 (30.0)
CoP
64.3
2/30 (6.7)
CoP-DR
42.7
1/30 (3.3)
r
70.1
14/30 (46.7)
TABLE IV : Adversarial Peg Insertion. Simulation reports success over 1024 held-out starts; hardware reports 30 trials per condition. Representation baselines compare proprioception only (Prop.; no tactile input), Binary Contact, CoP [ 5 ] , and the complete unfactorized Response r against GFC. CoP summarizes each pad by total response and its response-weighted centroid; r directly encodes the complete response field. For CoP-DR and r -DR, all tactile domain randomizations are applied jointly to the single representation, whereas GFC applies them representation-specifically. Factorization ablation compares all non-empty proper subsets of G/F/C with GFC. Bold indicates the highest result in each domain.
A primary bottleneck in contact-rich manipulation is the difficulty of collecting real-world data. Sim-to-real reinforcement learning offers a scalable alternative, but the simulation-reality gap prevents information-dense modalities like touch from being effectively used. Existing sim-to-real methods often mitigate this gap by simplifying tactile data into coarse low-dimensional features -- sacrificing the richness required for complex manipulation. In this work, we introduce Center-of-Pressure (CoP), an effective tactile representation grounded in physical principles that preserves dense contact information while maintaining robustness for sim-to-real transfer. To support this representation, we propose a sensor calibration scheme based on differentiable dynamics, enabling the estimation of taxel orientations without requiring ground-truth force measurements. We evaluate CoP on two blind, challenging contact-rich manipulation tasks: peg-in-hole insertion and ball balancing. Across both tasks, policies conditioned on CoP achieve zero-shot sim-to-real transfer on a multi-fingered hand, and outperform both coarse binary-contact and raw-taxel baselines. Analysis of learned policy states further suggests that CoP-conditioned policies encode task-relevant physical properties, such as object mass, as an emergent byproduct of control.
Tactile sensing provides direct measurements of contact interactions that are essential for robotic manipulation. However, current simulators lack the fidelity to faithfully model the complex deformation and transduction mechanics of tactile sensors, severely hindering sim-to-real transfer in robot learning pipelines. To address this challenge, we propose a multi-modal representation learning framework that aligns heterogeneous tactile modalities within a shared latent space, eliminating the need for accurate raw-signal simulation while preserving relevant contact information. Our approach employs modality-specific encoders to project diverse tactile observations, such as simulated penetration depth and real-world capacitance, into a common embedding space. The model is trained using self- and cross-reconstruction objectives alongside contrastive alignment, encouraging modality-invariant yet information-rich representations. We evaluate the learned embeddings on indenter shape identification, force prediction, and geometric reconstruction tasks, training exclusively in simulation and testing directly on real sensor measurements. Our results demonstrate zero-shot sim-to-real transfer across physically dissimilar representations. Furthermore, incorporating multi-physics simulation modalities yields more informative embeddings that transfer across diverse downstream tasks, demonstrating a 16.7% reduction in force prediction error and a 45.8% reduction in shape reconstruction error. Finally, we release an efficient Warp-based implementation of a penalty-based tactile simulation model for Isaac Lab, enabling scalable tactile data generation.
Arunim Joarder, Arjun Bhardwaj, René Zurbrügg +6
Robotic Systems Lab, ETH Zürich · ETH AI Center · NVIDIA +1
Contact-rich manipulation benefits from tactile feedback, yet physical tactile sensors introduce hardware, calibration, synchronization, and maintenance costs that complicate policy learning and deployment. We formulate predicted touch as an alternative to measured tactile input and present PredTac, a framework that learns to infer tactile states from causal visual observations and robot states and uses the predicted touch as an explicit interface for policy learning and execution. A tactile predictor is first trained with tactile supervision and then used to provide contact information without requiring measured tactile input during downstream policy training or execution. We evaluate PredTac across three contact-rich manipulation tasks in simulation and on a real robot, and further examine how policy performance depends on the predicted contact content. In simulation goal-offset evaluations, predicted-touch policies achieve 27.0%, 52.0%, and 44.7% success on USB, Barbed-spike, and Valve, respectively, improving over the visual baseline by 8.0-13.7 percentage points. On the real robot, predicted-touch ACT achieves 70.0%, 50.0%, and 90.0% success on USB insertion, Barbed extraction, and Valve rotation, respectively, with a three-task mean of 70.0%, approaching measured-touch ACT at 72.2% and substantially outperforming visual ACT at 21.1%. Fixed-policy interventions further show that performance is sensitive to the spatial structure of predicted contact, with spatial rearrangement at fixed value distributions reducing Valve success by 10.7 percentage points. These results demonstrate that predicted touch can provide useful contact information for contact-rich manipulation without requiring tactile sensing as a policy input.
Weijia Fan, Daqiang Guo
The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China.