Visible Touch: Rendering Contact for Visuomotor Policies
Organizations: University of California, Los Angeles
Abstract
Integrating contact information into visuomotor policies remains an open problem. Touch is essential to robust manipulation, yet most modern policies, including pretrained vision-language-action (VLA) models, operate from vision and proprioception alone. Existing approaches to closing this gap require specialized tactile hardware, add separate tactile encoders, or commit to non-image policy backbones, all incompatible with the modern paradigm of image-conditioned policies built on pretrained 2D visual representations. Our key insight is that the bottleneck is not the contact information itself, but how it is delivered: when contact signals are exposed in the same spatial frame as the scene the policy already attends to, they become directly usable by any image-conditioned policy without architectural changes. We operationalize this insight in Visible Touch, paired with a custom low-cost magnetic contact sensor that is open-sourced and fabricated from off-the-shelf parts via a parametric CAD-to-mold pipeline. Across the LIBERO benchmark, Visible Touch improves BC-Transformer success by 15.7 percentage points on average in the 2-view setting, with similar gains in the 1-view setting; controlled comparisons show that the contact-integration strategy strongly affects how effectively tactile information is used. The pattern holds when fine-tuning pretrained VLAs: miniVLA on LIBERO gains 25 percentage points on average, and on four real-world contact-rich tasks gains 30 percentage points with our custom sensor.
Figures & tables
| Pretraining | Fine-tuning | |||||
|---|---|---|---|---|---|---|
| LIBERO_90 | OBJECT | GOAL | SPATIAL | LONG | Avg. | |
| Baseline | 48.3% | 51.0% | 48.5% | 71.5% | 32.5% | 50.9% |
| Visible Touch | 63.2% | % | % | % | % | 76.2% |
| Visible Touch variant | OBJECT | GOAL | SPATIAL | LONG | Avg. |
|---|---|---|---|---|---|
| Multi-arrows | |||||
| Avg. Arrow | |||||
| Binbars |
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
| Component | Spec | Qty | Unit price | Subtotal |
|---|---|---|---|---|
| Sensor pad (consumable) | ||||
| XP-565 silicone elastomer | 10:1 base/activator | 10 g | $0.05/g | $0.50 |
| Neodymium magnets | 2 mm 1 mm | 9 | $0.07/ea | $0.63 |
| PLA filament (molds + spacer) | — | 20 g | $0.013/g | $0.26 |
| Loctite SF 770 | — | 1 ml | $0.50/ml | $0.50 |
| Loctite 406 | — | 0.5 ml | $1.65/ml | $0.83 |
| Component / Hyperparameter | Value |
|---|---|
| Architecture | |
| Visual backbone | ResNet-18 (random init, remove_layer_num=4 , no stride change) |
| Language fusion | FiLM conditioning on visual features |
| Language encoder | 1-layer MLP (input , hidden , output ) |
| Transformer layers | |
| Hidden / embed dimension | (token embed); MLP hidden |
| 1-view (agentview) | 2-view (agentview + wrist) | ||||
|---|---|---|---|---|---|
| Suite | Plain | Visible Touch | Plain | Proprio | Visible Touch |
| OBJECT | |||||
| GOAL | |||||
| SPATIAL | |||||
| LONG | |||||
| Average | |||||
| Suite | Task | Plain 2v | Visible Touch 2v | |
|---|---|---|---|---|
| OBJECT | butter basket | |||
| GOAL | wine top of cabinet | |||
| SPATIAL | bowl on ramekin plate | |||
| LONG | cheese + butter basket |
| Raw Pearson | Partial (controlling for headroom) | |||
|---|---|---|---|---|
| Feature | ||||
| Headroom | — | — | ||
| #grasps | ||||
| Trajectory length | ||||
| %contact | ||||
| #contact onsets | ||||
| Suite | 1-view Plain | 2-view Plain | (2v 1v) |
|---|---|---|---|
| OBJECT | |||
| GOAL | |||
| SPATIAL | |||
| LONG | |||
| Average |
| Suite | Seed 12345 | Seed 23456 | Seed 34567 | Mean | SD |
|---|---|---|---|---|---|
| OBJECT | |||||
| GOAL | |||||
| SPATIAL | |||||
| LONG |
| Suite | Seed 12345 | Seed 23456 | Seed 34567 | Mean | SD |
|---|---|---|---|---|---|
| OBJECT | |||||
| GOAL | |||||
| SPATIAL | |||||
| LONG |
| Suite | Seed 12345 | Seed 23456 | Seed 34567 | Mean | SD |
|---|---|---|---|---|---|
| OBJECT | |||||
| GOAL | |||||
| SPATIAL | |||||
| LONG |
| Suite | Recipe | Baseline | Visible Touch | |
|---|---|---|---|---|
| OBJECT | b8 1-epoch | pp | ||
| OBJECT | DDP-4 b16 cosine | pp | ||
| GOAL | b8 1-epoch | pp | ||
| GOAL | DDP-4 b16 cosine | pp | ||
| SPATIAL | 2-phase | pp | ||
| SPATIAL | DDP-4 b16 cosine | pp |
| Setting | Value |
|---|---|
| Base checkpoint | pi05_droid |
| Fine-tuning method | LoRA (backbone , action expert ) |
| LoRA targets | attention and FFN projections |
| Image resolution | , 3 cameras |
| Action horizon | 10 steps (1.0 s at 10 Hz) |
| Batch size | 8 |
| Task | Length (steps) | Duration (s) | Min/Max (steps) | |
|---|---|---|---|---|
| cube | 100 | 90 / 972 | ||
| tube | 97 | 166 / 238 | ||
| charger | 100 | 182 / 329 | ||
| dishwasher | 100 | 236 / 554 | ||
| All | 397 | 90 / 972 |
| Lift | Transfer Tube | Put Mug in Dishwasher | Plug Charger | ||||||
| Method | Pick | Pick | Insert | Pull | Pick | Place | Pick | Insert | |
| Baseline | 20/30 | 5/30 | 0/30 | 21/30 | 11/30 | 11/30 | 9/30 | 0/30 | |
| Tac-View | 16/30 | 9/30 | 0/30 | 22/30 | 18/30 | 16/30 | 17/30 | 0/30 | |
| Position-Only | 22/30 | 21/30 | 7/30 | 17/30 | 14/30 | 11/30 | 9/30 | 0/30 | |
| Visible Touch | Binbars | 23/30 | 10/30 | 1/30 | 24/30 | 15/30 | 15/30 | 9/30 | 0/30 |
| Binary Contact (1x) | 19/30 | 17/30 | 2/30 | 13/30 | 5/30 | 4/30 | 16/30 | 0/30 | |
| Task | Step budget | Action scale |
|---|---|---|
| Lift | 150 | 0.8 |
| Transfer Tube | 400 | 0.4 |
| Put Mug in Dishwasher | 600 | 0.8 |
| Plug Charger | 600 | 0.6 |