PolyUMI: Accessible Visual-Tactile-Audio Data Collection for Object Inference and Manipulation
Authors: Conor W. Hayes, Rickmer Krohn, Aravind Ramaswami, Anunth Ramaswami, Nils Dengler, Kevin M. Lynch, J. Edward Colgate, Georgia Chalvatzaki, +1 more
Organizations: Center for Robotics and Biosystems, Northwestern Univ., Evanston, IL, USA · Interactive Robot Perception & Learning (PEARL) Lab, TU Darmstadt, Germany · Robotics Institute Germany (RIG)
Humans typically rely on vision, touch, hearing, and proprioception to perceive contact and adapt their actions during manipulation. Providing robots with comparable responsiveness therefore requires hardware that can retain and use these complementary sensory signals. Most imitation-learning systems, however, observe demonstrations primarily through vision and proprioception, limiting access to contact information that is difficult to infer visually. We present PolyUMI, an open-source platform for scalable visual--tactile--audio demonstration collection and robot deployment. Its lightweight, wireless handheld gripper records synchronized wrist-camera, optical tactile, contact-audio, and proprioceptive observations without requiring a tethered workstation. The same sensing finger can be transferred to the robot end effector, preserving the sensing geometry between demonstration collection and policy execution. To effectively use these heterogeneous observations, we further introduce VisTA, a token-level multimodal policy that integrates information across sensors and time to predict contact-aware robot actions. Experiments spanning object inference, slip control, and contact-rich manipulation show that touch and audio reveal task-relevant information beyond vision and that VisTA is competitive with or outperforms existing multimodal policies. Together, PolyUMI and VisTA provide an accessible pipeline for collecting multimodal demonstrations and learning policies that perceive physical interaction beyond vision. Project Page: https://polyumi-vista.github.io
Figures & tables
Fig. 1 : PolyUMI connects wireless demonstration collection (right) to robot policy execution (left) using a shared multimodal sensing finger. Wrist vision, optical tactile images, and contact audio provide complementary observations for learning contact-rich manipulation.
Fig. 2 : Left : The relationship between different sensor views of the same scene. The tactile camera “sees” the screwdriver through the sensing surface, as well as the lightbulb through peripheral vision. The wrist camera sees the lightbulb directly, while the screwdriver is mostly occluded by the fingers holding it. As the device interacts with its environment, contact events are detected by the piezo microphone, which can be seen as spikes in the audio spectrogram, as well as heard through headphones during gripper operation. Center, Right : The PolyUMI system is composed of a gripper and end-effector, carefully designed to share the same tactile finger component and task-oriented geometry.
Sensor modalities
Method
Vision
Tactile
Audio
Open source
Cable free
Live feedback
UMI [ 7 ]
✓
×
×
✓
✓
×
FastUMI [ 13 ]
✓
×
×
✓
×
×
ViTaMin [ 12 ]
✓
✓
×
×
✓
×
Actuated UMI [ 11 ]
✓
✓
×
✓
✓
×
ManiWAV [ 3 ]
✓
×
✓
✓
✓
×
TABLE I : Comparison of multimodal data collection systems. PolyUMI is the only interface able to collect synchronous vision, tactile and audio data. It also supplies a novel modality of operator feedback through live audio monitoring. The system’s hardware and software are open-sourced for community use.
Modality
Sensor
Output
Tactile
Optical finger
20 fps, 1152×648 MJPEG
Audio
Contact mic
16 kHz mono PCM
Vision
GoPro Hero 12 + fisheye
60 fps, 1920×1080 MP4
Proprioception
SLAM / ArUco
6-DoF pose + gripper width
TABLE II : PolyUMI sensing modalities and their native output specifications.
Robot End-Effector [ 7 ]
Handheld UMI Gripper [ 7 ]
Multimodal Sensing Finger
Hardware cost
$742.48
$700.27
$235.96
Assembly time
30 min
2 h
4 h
TABLE III : Estimated hardware cost and manual assembly time for the main PolyUMI components. The multimodal sensing finger adds approximately $240 and four hours of assembly relative to the original UMI platform.
Fig. 3 : The VisTA model consists of sensor-specific CNN encoders and joint self-attention fusion. The high number of fused conditioning tokens C , then conditions a flow matching transformer.
Fig. 4 : Tactile images of the shape recognition objects. From left to right: Flat , Square , Circle , Triangle , and the letter N .
Fig. 6 : Audio signals (left) generated by shaking gears, thin screws, and thick screws (middle) with the robot arm (right) enable classification of visually occluded objects, substantially outperforming vision and touch.
Fig. 7 : Prediction distributions for object-in-box classification, aggregated over three random seeds. Each panel corresponds to one ground-truth object class, and each row represents a sensor configuration. The green outlined cells indicate correct predictions.
Fig. 8 : Sensor ablation success rates (SR) across 10 trials (left) and side view of the Slip Control Task setup (right). The combination of tactile and audio sensing enables closed-loop slippage control, where other sensor modalities fail.
Fig. 10 : Manipulation results of Board Wiping and Lightbulb. VisTA outperforms state of the art visual tactile audio models in the board wiping task. All models are able to solve Lightbulb turning. while multisensor models fail to detect task stages, rather than contact modes.
Touch sensing is beneficial for solving a wide variety of manipulation tasks. While there exists a wide range of tactile sensors with different properties, exploiting the fusion of multiple heterogeneous tactile sensors to improve manipulation learning remains underexplored. We present Multi-Resolution Tactile Sensing (MiTaS), a representation framework that leverages multiple tactile sensors operating at different temporal resolutions in order to solve complex contact-rich manipulation tasks. We propose a novel architecture using modality-specific convolutional stems and transformer-based fusion that effectively fuses information from an RGB camera stream, a vision-based GelSight Mini sensor and a high-frequency event-based Evetac sensor. This multi-sensor representation then conditions a flow-matching policy for solving downstream tasks. Experimental results across five contact-rich manipulation tasks demonstrate the effectiveness of multi-resolution tactile features in imitation learning. MiTaS achieves an average success rate of 80 %, while vision-only (31 %) and visual-tactile (54 %) baselines cannot solve the task reliably. Co-training a visuo-tactile model with multi-tactile data boosts performance by over 10 % in certain tasks, without having access to the Evetac sensor during policy evaluation. A detailed sensor-reading and attention analysis reveals the importance of different sensors throughout task execution, validating our multi-resolution tactile sensing approach. Project Page: http://mitas-touch.github.io.
Rickmer Krohn, Erik Helmut, Niklas Funk +3
Interactive Robot Perception & Learning, TU Darmstadt · Hessian AI · Robotics Institute Germany +1
Whole-body humanoid manipulation of bulky, deformable, and shared-load objects requires distributed contact sensing and explicit force regulation, yet most imitation policies treat contact force only implicitly. On the other hand, different demonstration sources provide complementary modalities with inherent trade-offs: human demonstrations capture natural contact forces but not robot-executable actions, while teleoperation directly records robot actions but with less natural force regulation. This paper presents \textbf{WT-UMI}, a wearable whole-body tactile interface worn by human operators or mounted on humanoids, providing accurate observations of tactile images, contact forces, and end-effector poses across both human demonstration and humanoid teleoperation modes. We introduce a force-conditioned target-pose correction module that converts measured human poses into contact-aware robot targets by learning corrections from teleoperation data. To leverage the natural force interaction in human data, we propose a force-supervised planner that predicts end-effector pose chunks and contact-force trajectories. The predicted contact force serves as the reference for a tactile-based admittance controller. Across five contact-rich tasks spanning deformable objects, bulky rigid objects, and human--humanoid collaboration, WT-UMI improves success rate and reduces contact-position tracking error over four policy baselines. Our project page is available at https://wt-umi.github.io/WTUMI/.
Jaehwi Jang, Zhaoyuan Gu, Alfred Cueva +15
The Institute for Robotics and Intelligent Machines, Georgia Institute of Technology
Tactile sensing can substantially improve contact-rich robotic manipulation, yet its practical deployment remains limited by the fragility, calibration requirements, and maintenance burden of tactile hardware. This raises a fundamental question: can robots benefit from tactile knowledge without requiring tactile sensors at deployment? We present TacImag, a tactile imagination framework that predicts tactile observations from vision and proprioception and uses the generated signals to guide manipulation policies. Trained from paired visuotactile demonstrations, TacImag enables touch-informed manipulation using only visual observations at test time. We evaluate TacImag on six simulated and four real-world manipulation tasks. Across simulation and real-world experiments, imagined tactile observations consistently improve manipulation performance without requiring tactile hardware. In real-world experiments, imagined force fields improve contact-sensitive tasks by 44.4% on average, whereas imagined tactile images improve texture-sensitive tasks by 23.3%, revealing that the effectiveness of tactile imagination depends strongly on the relationship between tactile representation and task requirements. Our results further suggest that tactile imagination does not simply recover missing tactile measurements. Instead, it acts as a form of contact-aware supervision that transforms subtle visual interaction cues into representations that are easier for manipulation policies to exploit.
Zhiyuan Zhang, Adeesh Desai, Jyun-Chi Hu +7
Department of Industrial Engineering, Purdue University, West Lafayette, IN 47907, USA · Department of Mechanical Engineering, Texas A&M University, College Station, TX 77843, USA