Integrating contact information into visuomotor policies remains an open problem. Touch is essential to robust manipulation, yet most modern policies, including pretrained vision-language-action (VLA) models, operate from vision and proprioception alone. Existing approaches to closing this gap require specialized tactile hardware, add separate tactile encoders, or commit to non-image policy backbones, all incompatible with the modern paradigm of image-conditioned policies built on pretrained 2D visual representations. Our key insight is that the bottleneck is not the contact information itself, but how it is delivered: when contact signals are exposed in the same spatial frame as the scene the policy already attends to, they become directly usable by any image-conditioned policy without architectural changes. We operationalize this insight in Visible Touch, paired with a custom low-cost magnetic contact sensor that is open-sourced and fabricated from off-the-shelf parts via a parametric CAD-to-mold pipeline. Across the LIBERO benchmark, Visible Touch improves BC-Transformer success by 15.7 percentage points on average in the 2-view setting, with similar gains in the 1-view setting; controlled comparisons show that the contact-integration strategy strongly affects how effectively tactile information is used. The pattern holds when fine-tuning pretrained VLAs: miniVLA on LIBERO gains 25 percentage points on average, and π0.5 on four real-world contact-rich tasks gains 30 percentage points with our custom sensor.
Figures & tables
Figure 1: Overview of Visible Touch . A low-cost magnetic contact sensor captures per-taxel contact, which we render as an image-space overlay on the policy’s RGB observations. The augmented images are drop-in inputs for fine-tuning pretrained VLAs, requiring no architectural changes. Contact overlays improve average task success by +25.3pp on LIBERO fine-tuning suites and +30.3pp on real-world contact-rich manipulation tasks.
Figure 2: Hardware system. (a) Visuo-tactile setup: UFactory xArm7 with two RealSense cameras and two tactile sensors mounted on the gripper fingers. (b) Custom tactile sensor with anti-slip skin. (c) Internal structure of the tactile sensor. (d) Our parametric CAD-to-mold pipeline: CAD-generated alternative casting molds (top) and the 3D-printed molds with the fabricated sensor pads (bottom). (e) Physical characteristics of the sensor.
Figure 3: BC-Transformer on LIBERO across camera configurations ( n=3 seeds, error bars are std). Orange bars show Visible Touch (ours); gray bars show plain baselines; blue bars show the Proprio baseline that receives an aggregated 12-D contact wrench through the proprioceptive state. Visible Touch consistently improves over plain images in both 1-view and 2-view, with the largest gain on LIBERO_LONG ( +33.7 pp 2-view). See details in Appendix C .
Pretraining
Fine-tuning
LIBERO_90
OBJECT
GOAL
SPATIAL
LONG
Avg.
Baseline
48.3%
51.0%
48.5%
71.5%
32.5%
50.9%
Visible Touch
63.2%
70.6±4.6 %
80.5±4.3 %
87.0±1.8 %
66.6±2.2 %
76.2%
Table 1: miniVLA success rates (%) on LIBERO across pretraining and fine-tuning. LIBERO_90 pretraining results are from single runs. For fine-tuning, baseline entries are the strongest single-run baselines available per suite, while Visible Touch results are mean ± s.e. over 5 training seeds using the suite-specific headline recipe.
Visible Touch variant
OBJECT
GOAL
SPATIAL
LONG
Avg.
Multi-arrows
94.3±5.6
91.8±0.5
90.2±1.0
76.2±2.7
88.1
Avg. Arrow
98.2±1.2
92.8±2.5
92.3±2.7
86.5±2.3
92.5
Binbars
98.0±0.9
92.7±1.7
94.5±1.5
78.8±2.6
91.0
Table 2: Success rates (%) for overlay variants on LIBERO suites (BC-Transformer, 2-view, n=3 seeds).
Figure 4: Success rates of Visible Touch (Avg. Arrow overlay) over plain baselines, bucketed by gripper-close events. Tasks with no grasps show no benefit; single-grasp tasks show modest gains; multi-grasp tasks show the largest gains. The pattern holds across both BC-Transformer ( 4(a) ) and a pretrained miniVLA on LIBERO_90 ( 4(b) ).
Figure 5: Real-world experimental setup. (a) Contact-rich manipulation tasks: Lift, Transfer Tube, Put Mug in Dishwasher , and Plug Charger . (b) Example contact overlays rendered onto external and wrist camera views.
Figure 6: Tac-View : contact arrows as a separate image.
Figure 7: Real-world evaluation of π0.5 fine-tuning across four contact-rich manipulation tasks (successes out of 30 trials per condition, sub-task scored). Final-stage success corresponds to complete end-to-end task success. Hatched bars denote the binary ablations of the matching solid variant: Binary Contact (1x) and (9x) keep only per-finger and per-taxel contact occurrence, respectively. Exact counts are listed in Table 15 ( Section F.3 ).
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Spec
Qty
Unit price
Subtotal
Sensor pad (consumable)
XP-565 silicone elastomer
10:1 base/activator
10 g
$0.05/g
$0.50
Neodymium magnets
∅ 2 mm × 1 mm
9
$0.07/ea
$0.63
PLA filament (molds + spacer)
—
20 g
$0.013/g
$0.26
Loctite SF 770
—
1 ml
$0.50/ml
$0.50
Loctite 406
—
0.5 ml
$1.65/ml
$0.83
Appendix
Table 3: Bill of materials for one tactile sensor ( 3×3 configuration). Sensor pad components are consumable and can be replaced without replacing the PCB. PCB costs are rough estimates and vary significantly with order quantity.
Figure 8: Fabrication steps of our custom tactile sensor.
Figure 9: Sensor characterization experiments. (a) Experiment setup: UFactory xArm7 equipped with BOTA Force-Torque sensor and our custom tactile sensor, used to collect ground-truth contact normal force data for characterizing the tactile sensor’s response. (b) Standard deviation of the 4 independently fabricated sensor readings under different contact normal forces.
Figure 10: Directional shear response and hysteresis. Left: in-plane flux change of the center taxel during ±1 mm tangential sweeps along X and Y at 0.6 mm indentation (3 repetitions): orthogonal, monotonic, sign-symmetric response trajectories. Right: aggregate response r versus normal force over six consecutive load–unload cycles at 0.5 mm/s; maximum loading–unloading separation is ≤ 14% of full scale and identical across cycles.
Figure 11: Baseline stability over a 30-minute stationary hold. Top: aggregate response over the full session; orange lines mark the reference presses bracketing the hold. Bottom: signed per-channel baseline deviation (60 s means) of all 27 channels during the hold; 26 of 27 channels stay within ±20μ T, and the largest excursion ( −55μ T, center-taxel normal axis) is a settling transient of 0.7% of a typical press signal.
Component / Hyperparameter
Value
Architecture
Visual backbone
ResNet-18 (random init, remove_layer_num=4 , no stride change)
Table 4: BC-Transformer training configuration for LIBERO multitask experiments.
1-view (agentview)
2-view (agentview + wrist)
Suite
Plain
Visible Touch
Plain
Proprio
Visible Touch
OBJECT
90.7±4.4
95.0±0.9
82.8±7.0
80.0±2.0
98.2±1.2
GOAL
79.3±3.2
91.7±0.3
86.8±4.1
81.5±4.3
92.8±2.5
SPATIAL
79.0±2.0
90.5±0.4
84.8±4.0
80.0±1.4
92.3±2.7
LONG
41.8±6.1
63.8±1.2
52.8±2.1
42.8±5.1
86.5±2.3
Average
72.7
85.2
76.8
71.1
92.5
Appendix
Table 5: BC-Transformer success rates (%) on LIBERO suites across camera configurations ( n=3 seeds).
Suite
Task
Plain 2v
Visible Touch 2v
Δ(pp)
OBJECT
butter → basket
48%
98%
+50
GOAL
wine → top of cabinet
83%
100%
+17
SPATIAL
bowl on ramekin → plate
62%
93%
+32
LONG
cheese + butter → basket
8%
95%
+87
Appendix
Table 6: Rescue outliers: tasks where Avg. Arrow produces the largest gain over plain 2-view within each suite. All four outliers involve grasp configurations that are visually ambiguous (slippery, small, occluded, or perched objects), where the contact arrow provides information not available from pixels alone.
Raw Pearson
Partial (controlling for headroom)
Feature
r
p
r
p
Headroom
+0.924
<10−3
—
—
#grasps
+0.633
<10−3
−0.094
0.562
Trajectory length
+0.458
0.003
−0.368
0.019
%contact
+0.360
0.022
+0.214
0.186
#contact onsets
+0.329
0.038
−0.571
<10−3
Appendix
Table 7: Per-task correlations between Δ= Visible Touch 2v − plain 2v and task properties, pooled across all 40 LIBERO tasks. Raw Pearson correlations are dominated by headroom (overlays help most where vision fails). After controlling for headroom via partial correlation, only #contact onsets and trajectory length retain independent predictive signal.
Suite
1-view Plain
2-view Plain
Δ (2v − 1v)
OBJECT
90.7
82.8
−7.9
GOAL
79.3
86.8
+7.5
SPATIAL
79.0
84.8
+5.8
LONG
41.8
52.8
+11.0
Average
72.7
76.8
+4.1
Appendix
Table 8: Effect of camera configuration on the plain (no-overlay) BC-Transformer baseline. Mean success rate (%) over 3 seeds per cell, 10 tasks per suite. Three of four suites benefit from the wrist camera; OBJECT regresses.
Suite
Seed 12345
Seed 23456
Seed 34567
Mean
SD
OBJECT
84.5
73.5
90.5
82.8
7.0
GOAL
90.0
89.5
81.0
86.8
4.1
SPATIAL
90.5
82.0
82.0
84.8
4.0
LONG
50.5
52.5
55.5
52.8
2.1
Appendix
Table 9: Per-seed success rates (%) for the Plain 2-view BC-Transformer baseline.
Suite
Seed 12345
Seed 23456
Seed 34567
Mean
SD
OBJECT
82.5
77.5
80.0
80.0
2.0
GOAL
84.0
75.5
85.0
81.5
4.3
SPATIAL
82.0
79.0
79.0
80.0
1.4
LONG
50.0
38.5
40.0
42.8
5.1
Appendix
Table 10: Per-seed success rates (%) for the Proprio 2-view baseline (Plain RGB + 12-D contact wrench concatenated to the proprioceptive state).
Figure 12: Per-task success rates on LIBERO_90: Visible Touch -miniVLA (contact-overlay pretraining) vs. the re-evaluated Stanford miniVLA [ 1 ] baseline.
Suite
Recipe
Baseline
Visible Touch
Δ
OBJECT
b8 1-epoch
47.0%
57.5%
+10.5 pp
OBJECT
DDP-4 b16 cosine
51.0%
70.6±4.6%
+19.6 pp
GOAL
b8 1-epoch
48.5%
74.0%
+25.5 pp
GOAL
DDP-4 b16 cosine
37.4%
80.5±4.3%
+43.1 pp
SPATIAL
2-phase
71.5%
87.0±1.8%
+15.5 pp
SPATIAL
DDP-4 b16 cosine
55.0%
87.5%
+32.5 pp
Appendix
Table 12: Per-recipe comparison for History-2 miniVLA on LIBERO suites. Baseline : clean non-contact fine-tuning with clean evaluation. Visible Touch : contact-overlay fine-tuning with overlay-augmented evaluation. For the suite-specific recipes used in the main results, Visible Touch performance is reported as mean ± standard error over 5 training seeds; all other entries are single runs. Baseline results are single runs.
Setting
Value
Base checkpoint
pi05_droid
Fine-tuning method
LoRA (backbone r=16 , action expert r=32 )
LoRA targets
attention and FFN projections
Image resolution
224×224×3 , 3 cameras
Action horizon
10 steps (1.0 s at 10 Hz)
Batch size
8
Appendix
Table 13: π0.5 LoRA fine-tuning hyperparameters. All 32 runs share these settings; only the input variant differs.
Task
n
Length (steps)
Duration (s)
Min/Max (steps)
cube
100
131±86
13.1±8.6
90 / 972
tube
97
190±14
19.0±1.4
166 / 238
charger
100
238±28
23.8±2.7
182 / 329
dishwasher
100
281±32
28.1±3.2
236 / 554
All
397
210±74
21.0±7.4
90 / 972
Appendix
Table 14: Per-task demonstration statistics (post-acceptance, post-idle-trim). n is the number of episodes used in training; length is reported as mean ± standard deviation in control steps and seconds. Min/max bracket the full distribution.
Lift
Transfer Tube
Put Mug in Dishwasher
Plug Charger
Method
Pick
Pick
Insert
Pull
Pick
Place
Pick
Insert
Baseline
20/30
5/30
0/30
21/30
11/30
11/30
9/30
0/30
Tac-View
16/30
9/30
0/30
22/30
18/30
16/30
17/30
0/30
Position-Only
22/30
21/30
7/30
17/30
14/30
11/30
9/30
0/30
Visible Touch
Binbars
23/30
10/30
1/30
24/30
15/30
15/30
9/30
0/30
Binary Contact (1x)
19/30
17/30
2/30
13/30
5/30
4/30
16/30
0/30
Appendix
Table 15: Real-world evaluation of π0.5 fine-tuning across four contact-rich manipulation tasks (30 trials per condition, sub-task scored). Final-stage success corresponds to complete end-to-end task success. Binary Contact (1x) and (9x) denote aggregate per-finger and per-taxel binary contact, respectively.
Task
Step budget
Action scale
Lift
150
0.8
Transfer Tube
400
0.4
Put Mug in Dishwasher
600
0.8
Plug Charger
600
0.6
Appendix
Table 16: Rollout budget and action scale per real-world task.