TACROSS: An Efficient and Low-Cost Scalable Human Touch System Across Heterogeneous Tactile Sensors for Dexterous Robot Learning
Authors: Bo Chen, Huanzhang Hu, Junyang Ma, Bo Yue, Fangdi Yu, Haijier Chen, Xianxin Lai, Shuyu Pan, +6 more
Organizations: The University of Hong Kong · INFIFORCE · The Chinese University of Hong Kong (Shenzhen) · Wuhan Textile University · Ocean University of China · Wuhan University
Collecting tactile demonstrations on robots is costly and slow, motivating the use of lower-cost human tactile gloves for scalable data collection. However, human capacitive/piezoresistive gloves and robotic tactile sensors differ fundamentally in transduction principle, sensor layout, spatial resolution, and dynamic response, making alignment of raw sensor channels ill-posed. To address this problem, we present TACROSS, a scalable system for learning from human touch and transferring it to robots that bridges this heterogeneity by aligning tactile streams at the level of contact events rather than raw sensor values. The hardware component of TACROSS integrates a piezoresistive glove with five layers and a cost of USD 10.86 with 285 sensing points. To align contact semantics, we design canonicalizers and residual adapters that map heterogeneous signals into a shared tactile latent with 256 dimensions via a temporal Transformer with attention across fingers. We further introduce a robot-grounded policy learning scheme in which robot demonstrations provide the sole source of ground-truth action supervision, while human demonstrations support tactile representation learning and provide confidence-weighted auxiliary supervision through valid retargeted hand targets. We evaluate our system on four contact-rich manipulation tasks. Compared to conventional teleoperation, our proposed system achieves a 3.5-fold efficiency improvement while reducing demonstration acquisition equipment cost by 95.7%. We will open-source the TACROSS hardware and software system and publicly release a tactile dataset comprising over 150 hours of recordings. Project page: https://tacross-touch-project.github.io/.
TABLE I: Hardware specifications and characterization of the TACROSS glove. F.S. denotes full scale.
Fig. 3: TACROSS human demonstration acquisition pipeline. (a) Tactile glove and on-wrist readout integrated with motion capture (Rokoko Ego) or vision-based inverse kinematics (Vision-IK Ego). (b) Sensor calibration and causal synchronization using native timestamps, selecting the latest available sample at or before each reference time. (c) Aligned RGB images, hand poses, and five-finger tactile observations form multimodal episodes for subsequent tactile representation and policy learning.
Fig. 4: Contact canonicalization and adaptation, shared spatiotemporal encoding, and ACT learning. Each window contains 80 tokens (16 frames × five fingers), processed by shared Transformers with two layers, four attention heads, and features of dimension 128. H is a learned hand token, F denotes finger features, and repeated labels identify the same tensor. Training uses masked canonical targets in both domains and contrastive attraction/repulsion with phase/force hard negatives. P projects latents from 256 dimensions to tokens with 512 dimensions through Linear, GELU, and LayerNorm layers. The illustrated state has eight dimensions and uses one degree of freedom for hand control, Sec. III-D defines the articulated action interface with 13 dimensions. Deployment uses robot observations.
Fig. 5: Representative human and robot demonstrations for T1–T4 (top to bottom). Each task shows human keyframes (left), tactile maps (middle), robot keyframes (right), and index-finger (blue) and thumb (orange) force traces (far right). Timestamps indicate elapsed time in each demonstration.
Fig. 6: Evaluation setup on the robot with the arm, Revo2 tactile hand, RGB cameras, task fixture, and receiving tray.
Metric
Robot-only
Ego-R
Ego-V (Ours)
Total recording duration (h)
24.86
89.49
69.60
Acquisition system cost (USD) ↓
12,000
2,150
520
Recorded demonstrations
1,236
12,784
10,893
Valid demonstration rate (%)
86.89
91.28
89.44
Valid demonstrations
1,074
11,669
9,743
Collection speedup ↑
1.00×
2.80×
3.50×
TABLE II: Data Collection Efficiency, System Cost, and Task Performance.
TABLE III: Comparison of Full-Hand Tactile Gloves
Characteristic
A
B
C
Ours
Robot:Ego ratio
1:0
1:5
1:10
1:5
1:10
1:5
1:10
Robot episodes
30
30
16
30
16
30
16
Human episodes
0
150
160
150
160
150
160
Fixed-Point Dispensing
52.5
77.5
77.5
47.5
45.0
95.0
92.5
Beaker-to-Beaker Liquid Transfer
47.5
72.5
70.0
50.0
50.0
92.5
87.5
Precision Pipetting
50.0
67.5
65.0
47.5
45.0
92.5
87.5
TABLE IV: Ablation Study of Human Tactile Input and Contact-Semantic Alignment
Characteristic
A
B
C
D
E
F
Robot episodes
15
15
15
15
15
105
Human episodes
0
20
40
60
90
0
Human ratio (%)
0.0
57.1
72.7
80.0
85.7
0.0
Collection cost (USD)
2.00
2.03
2.06
2.09
2.14
14.00
T1 success (%)
40.0
42.5
65.0
80.0
90.0
95.0
T2 success (%)
45.0
47.5
65.0
77.5
85.0
87.5
TABLE V: Scaling human data with 15 fixed robot episodes.
Fig. 7: Success rates and acquisition equipment costs for different robot–Ego data mixtures. (a) Task success under fixed collection time budgets of 50, 150, 200, and 250 min for T1–T4, respectively. (b) Equipment cost for each acquisition system. A uses robot-only teleoperation, B combines Rokoko motion capture and a tactile glove, and Ours combines vision-based IK and a tactile glove. Both Ego systems use robot demonstrations as anchors. Ratios denote nominal robot:Ego episode counts.