PolyUMI: Accessible Visual-Tactile-Audio Data Collection for Object Inference and Manipulation
Authors: Conor W. Hayes, Rickmer Krohn, Aravind Ramaswami, Anunth Ramaswami, Nils Dengler, Kevin M. Lynch, J. Edward Colgate, Georgia Chalvatzaki, +1 more
Organizations: Center for Robotics and Biosystems, Northwestern Univ., Evanston, IL, USA · Interactive Robot Perception & Learning (PEARL) Lab, TU Darmstadt, Germany · Robotics Institute Germany (RIG)
Humans typically rely on vision, touch, hearing, and proprioception to perceive contact and adapt their actions during manipulation. Providing robots with comparable responsiveness therefore requires hardware that can retain and use these complementary sensory signals. Most imitation-learning systems, however, observe demonstrations primarily through vision and proprioception, limiting access to contact information that is difficult to infer visually. We present PolyUMI, an open-source platform for scalable visual--tactile--audio demonstration collection and robot deployment. Its lightweight, wireless handheld gripper records synchronized wrist-camera, optical tactile, contact-audio, and proprioceptive observations without requiring a tethered workstation. The same sensing finger can be transferred to the robot end effector, preserving the sensing geometry between demonstration collection and policy execution. To effectively use these heterogeneous observations, we further introduce VisTA, a token-level multimodal policy that integrates information across sensors and time to predict contact-aware robot actions. Experiments spanning object inference, slip control, and contact-rich manipulation show that touch and audio reveal task-relevant information beyond vision and that VisTA is competitive with or outperforms existing multimodal policies. Together, PolyUMI and VisTA provide an accessible pipeline for collecting multimodal demonstrations and learning policies that perceive physical interaction beyond vision. Project Page: https://polyumi-vista.github.io
Figures & tables
Fig. 1 : PolyUMI connects wireless demonstration collection (right) to robot policy execution (left) using a shared multimodal sensing finger. Wrist vision, optical tactile images, and contact audio provide complementary observations for learning contact-rich manipulation.
Fig. 2 : Left : The relationship between different sensor views of the same scene. The tactile camera “sees” the screwdriver through the sensing surface, as well as the lightbulb through peripheral vision. The wrist camera sees the lightbulb directly, while the screwdriver is mostly occluded by the fingers holding it. As the device interacts with its environment, contact events are detected by the piezo microphone, which can be seen as spikes in the audio spectrogram, as well as heard through headphones during gripper operation. Center, Right : The PolyUMI system is composed of a gripper and end-effector, carefully designed to share the same tactile finger component and task-oriented geometry.
Sensor modalities
Method
Vision
Tactile
Audio
Open source
Cable free
Live feedback
UMI [ 7 ]
✓
×
×
✓
✓
×
FastUMI [ 13 ]
✓
×
×
✓
×
×
ViTaMin [ 12 ]
✓
✓
×
×
✓
×
Actuated UMI [ 11 ]
✓
✓
×
✓
✓
×
ManiWAV [ 3 ]
✓
×
✓
✓
✓
×
TABLE I : Comparison of multimodal data collection systems. PolyUMI is the only interface able to collect synchronous vision, tactile and audio data. It also supplies a novel modality of operator feedback through live audio monitoring. The system’s hardware and software are open-sourced for community use.
Modality
Sensor
Output
Tactile
Optical finger
20 fps, 1152×648 MJPEG
Audio
Contact mic
16 kHz mono PCM
Vision
GoPro Hero 12 + fisheye
60 fps, 1920×1080 MP4
Proprioception
SLAM / ArUco
6-DoF pose + gripper width
TABLE II : PolyUMI sensing modalities and their native output specifications.
Robot End-Effector [ 7 ]
Handheld UMI Gripper [ 7 ]
Multimodal Sensing Finger
Hardware cost
$742.48
$700.27
$235.96
Assembly time
30 min
2 h
4 h
TABLE III : Estimated hardware cost and manual assembly time for the main PolyUMI components. The multimodal sensing finger adds approximately $240 and four hours of assembly relative to the original UMI platform.
Fig. 3 : The VisTA model consists of sensor-specific CNN encoders and joint self-attention fusion. The high number of fused conditioning tokens C , then conditions a flow matching transformer.
Fig. 4 : Tactile images of the shape recognition objects. From left to right: Flat , Square , Circle , Triangle , and the letter N .
Fig. 6 : Audio signals (left) generated by shaking gears, thin screws, and thick screws (middle) with the robot arm (right) enable classification of visually occluded objects, substantially outperforming vision and touch.
Fig. 7 : Prediction distributions for object-in-box classification, aggregated over three random seeds. Each panel corresponds to one ground-truth object class, and each row represents a sensor configuration. The green outlined cells indicate correct predictions.
Fig. 8 : Sensor ablation success rates (SR) across 10 trials (left) and side view of the Slip Control Task setup (right). The combination of tactile and audio sensing enables closed-loop slippage control, where other sensor modalities fail.
Fig. 10 : Manipulation results of Board Wiping and Lightbulb. VisTA outperforms state of the art visual tactile audio models in the board wiping task. All models are able to solve Lightbulb turning. while multisensor models fail to detect task stages, rather than contact modes.
Department of Industrial Engineering, Purdue University, West Lafayette, IN 47907, USA · Department of Mechanical Engineering, Texas A&M University, College Station, TX 77843, USA