TACROSS: An Efficient and Low-Cost Scalable Human Touch System Across Heterogeneous Tactile Sensors for Dexterous Robot Learning
Authors: Bo Chen, Huanzhang Hu, Junyang Ma, Bo Yue, Fangdi Yu, Haijier Chen, Xianxin Lai, Shuyu Pan, +6 more
Organizations: The University of Hong Kong · INFIFORCE · The Chinese University of Hong Kong (Shenzhen) · Wuhan Textile University · Ocean University of China · Wuhan University
Collecting tactile demonstrations on robots is costly and slow, motivating the use of lower-cost human tactile gloves for scalable data collection. However, human capacitive/piezoresistive gloves and robotic tactile sensors differ fundamentally in transduction principle, sensor layout, spatial resolution, and dynamic response, making alignment of raw sensor channels ill-posed. To address this problem, we present TACROSS, a scalable system for learning from human touch and transferring it to robots that bridges this heterogeneity by aligning tactile streams at the level of contact events rather than raw sensor values. The hardware component of TACROSS integrates a piezoresistive glove with five layers and a cost of USD 10.86 with 285 sensing points. To align contact semantics, we design canonicalizers and residual adapters that map heterogeneous signals into a shared tactile latent with 256 dimensions via a temporal Transformer with attention across fingers. We further introduce a robot-grounded policy learning scheme in which robot demonstrations provide the sole source of ground-truth action supervision, while human demonstrations support tactile representation learning and provide confidence-weighted auxiliary supervision through valid retargeted hand targets. We evaluate our system on four contact-rich manipulation tasks. Compared to conventional teleoperation, our proposed system achieves a 3.5-fold efficiency improvement while reducing demonstration acquisition equipment cost by 95.7%. We will open-source the TACROSS hardware and software system and publicly release a tactile dataset comprising over 150 hours of recordings. Project page: https://tacross-touch-project.github.io/.
TABLE I: Hardware specifications and characterization of the TACROSS glove. F.S. denotes full scale.
Fig. 3: TACROSS human demonstration acquisition pipeline. (a) Tactile glove and on-wrist readout integrated with motion capture (Rokoko Ego) or vision-based inverse kinematics (Vision-IK Ego). (b) Sensor calibration and causal synchronization using native timestamps, selecting the latest available sample at or before each reference time. (c) Aligned RGB images, hand poses, and five-finger tactile observations form multimodal episodes for subsequent tactile representation and policy learning.
Fig. 4: Contact canonicalization and adaptation, shared spatiotemporal encoding, and ACT learning. Each window contains 80 tokens (16 frames × five fingers), processed by shared Transformers with two layers, four attention heads, and features of dimension 128. H is a learned hand token, F denotes finger features, and repeated labels identify the same tensor. Training uses masked canonical targets in both domains and contrastive attraction/repulsion with phase/force hard negatives. P projects latents from 256 dimensions to tokens with 512 dimensions through Linear, GELU, and LayerNorm layers. The illustrated state has eight dimensions and uses one degree of freedom for hand control, Sec. III-D defines the articulated action interface with 13 dimensions. Deployment uses robot observations.
Fig. 5: Representative human and robot demonstrations for T1–T4 (top to bottom). Each task shows human keyframes (left), tactile maps (middle), robot keyframes (right), and index-finger (blue) and thumb (orange) force traces (far right). Timestamps indicate elapsed time in each demonstration.
Fig. 6: Evaluation setup on the robot with the arm, Revo2 tactile hand, RGB cameras, task fixture, and receiving tray.
Metric
Robot-only
Ego-R
Ego-V (Ours)
Total recording duration (h)
24.86
89.49
69.60
Acquisition system cost (USD) ↓
12,000
2,150
520
Recorded demonstrations
1,236
12,784
10,893
Valid demonstration rate (%)
86.89
91.28
89.44
Valid demonstrations
1,074
11,669
9,743
Collection speedup ↑
1.00×
2.80×
3.50×
TABLE II: Data Collection Efficiency, System Cost, and Task Performance.
TABLE III: Comparison of Full-Hand Tactile Gloves
Characteristic
A
B
C
Ours
Robot:Ego ratio
1:0
1:5
1:10
1:5
1:10
1:5
1:10
Robot episodes
30
30
16
30
16
30
16
Human episodes
0
150
160
150
160
150
160
Fixed-Point Dispensing
52.5
77.5
77.5
47.5
45.0
95.0
92.5
Beaker-to-Beaker Liquid Transfer
47.5
72.5
70.0
50.0
50.0
92.5
87.5
Precision Pipetting
50.0
67.5
65.0
47.5
45.0
92.5
87.5
TABLE IV: Ablation Study of Human Tactile Input and Contact-Semantic Alignment
Characteristic
A
B
C
D
E
F
Robot episodes
15
15
15
15
15
105
Human episodes
0
20
40
60
90
0
Human ratio (%)
0.0
57.1
72.7
80.0
85.7
0.0
Collection cost (USD)
2.00
2.03
2.06
2.09
2.14
14.00
T1 success (%)
40.0
42.5
65.0
80.0
90.0
95.0
T2 success (%)
45.0
47.5
65.0
77.5
85.0
87.5
TABLE V: Scaling human data with 15 fixed robot episodes.
Fig. 7: Success rates and acquisition equipment costs for different robot–Ego data mixtures. (a) Task success under fixed collection time budgets of 50, 150, 200, and 250 min for T1–T4, respectively. (b) Equipment cost for each acquisition system. A uses robot-only teleoperation, B combines Rokoko motion capture and a tactile glove, and Ours combines vision-based IK and a tactile glove. Both Ego systems use robot demonstrations as anchors. Ratios denote nominal robot:Ego episode counts.
Tactile sensors provide critical information for contact-rich manipulation, yet tactile representations and policies remain tightly coupled to each specific sensor, limiting transferability across robots and hardware platforms. We propose TactX, a framework for learning a transferable tactile representation across sensors spanning three fundamentally different transduction modalities: resistive, magnetic, and vision-based. TactX maps heterogeneous tactile observations into a shared latent space through modality-specific encoders trained on paired contact data. Such paired interactions provide a natural alignment signal across modalities, and the encoders are jointly trained across all sensor pairs, inducing a consistent latent space for all sensor types. Our experiments show that TactX aligns tactile representations across sensors while preserving object-level contact information, as evidenced by sensor-identity prediction and object classification in the learned latent space. We evaluate TactX on four contact-rich manipulation tasks: pick-and-place, plug insertion, board wiping, and object reorientation, and show that policies trained with one sensor transfer zero-shot to physically distinct sensors through the shared latent. This improves the average success rate from 27.5% for vision-only policy to 45.9%, providing a step toward sensor-agnostic tactile manipulation.
Junsung Park, Sachin Bhadang, Carmelo Sferrazza +2
UC San Diego · Seoul National University · Amazon FAR
As an essential modality for dexterous and contact-rich tasks, tactile sensing provides precise force feedback that cannot be reliably inferred from vision. However, limited by hardware and data collection systems, existing datasets with tactility remain small in scale and narrow in contact coverage. Meanwhile, Vision-Language-Action (VLA) models with tactile modality are constrained on dynamics-agnostic post-training, which limits the performance ceiling on downstream tasks. In this paper, we present H-Tac, a large-scale tactile-action dataset with 160-hour egocentric human videos containing more than 300 tasks and 135k episodes. Building upon this, we propose Transferable Tactile Pre-Training (TTP), a system of tactile-based pre-training on human data for fine-grained robotic tasks. To bridge the gap between humans and robots, we use unified tactile and action spaces throughout the pre-training and post-training phases, preserving prior knowledge during human-to-robot transfer. By leveraging a tactile expert for future tactile prediction, our framework explicitly models the contact dynamics and precise physical interactions. Extensive experiments in simulation and on real robots demonstrate that our model achieves superior performance, exhibiting robust generalization and fine-grained manipulation capabilities. TTP paves the way for scalable tactile pre-training via human-to-robot transfer.
Chi Zhang, Penglin Cai, Ziheng Xi +6
1Peking University · 2BeingBeyond · 3Tsinghua University
We present FlexiTac, a low-cost, open-source, and scalable piezoresistive tactile sensing solution designed for robotic end-effectors. FlexiTac is a practical "plug-in" module consisting of (i) thin, flexible tactile sensor pads that provide dense tactile signals and (ii) a compact multi-channel readout board that streams synchronized measurements for real-time control and large-scale data collection. FlexiTac pads adopt a sealed three-layer laminate stack (FPC-Velostat-FPC) with electrode patterns directly integrated into flexible printed circuits, substantially improving fabrication throughput and repeatability while maintaining mechanical compliance for deployment on both rigid and soft grippers. The readout electronics use widely available, low-cost components and stream tactile signals to a host computer at 100 Hz via serial communication. Across multiple configurations, including fingertip pads and larger tactile mats, FlexiTac can be mounted on diverse platforms without major mechanical redesign. We further show that FlexiTac supports modern tactile learning pipelines, including 3D visuo-tactile fusion for contact-aware decision making, cross-embodiment skill transfer, and real-to-sim-to-real fine-tuning with GPU-parallel tactile simulation. Our project page is available at https://flexitac.github.io/.