cs.ROOct 6, 2026

TacZero: Training-Free Peg Insertion Using a General-Purpose Vision-Language Model with Tactile Feedback

Authors: Kazutoshi Tanaka

Organizations: OMRON SINIC X Corporation, Bunkyo-ku, Tokyo 113-0033, Japan

Abstract

Robots that autonomously determine their actions from language instructions and sensory observations could perform new contact-rich manipulation tasks without task-specific training or hand-designed rules. To perform these tasks, robots must infer how objects contact one another and move as a result, then select actions. For contact inference and action selection, prior approaches involve designing estimation models and tactile feedback control laws, or learning models for object-motion estimation, action-outcome prediction, and action selection from tactile data. Instead, we propose TacZero, which uses a pretrained general-purpose vision-language model (VLM) to interpret visual and tactile observations and select robot actions without additional tactile or manipulation training or task-specific rules for contact interpretation or action selection. TacZero provides the VLM with camera images, robot state, and three-axis tactile responses represented as numerical values or vectors overlaid on the images. From these observations and interaction history, the VLM generates commands specifying target end-effector positions and gripper opening or closing, which a low-level controller executes. In real-world cylindrical-peg insertion experiments, TacZero succeeded in 15 of 20 trials with numerical tactile input, compared with 10 of 20 without tactile input. This study provides a concrete starting point for further research on contact-rich manipulation using general-purpose VLMs and highlights challenges in pursuing this direction.

Figures & tables

Explore similar work

CardsList
  1. Tactile Curiosity Drives Robot Interaction

    Sep 30, 2026Klemens Iten, Alexander Proshkin, Bhavya Sukhija +4TactileRobot Systems

  2. TACO: TActile World Model as a Self-COrrector for Scalable Robot Policy Post-Training

    Jul 3, 2026Shengbang Liu, Yueru Jia, Yuyang Yan +7Tactile World ModelDiffusion-Based Vision-Language-Actions

  3. TacCoRL: Integrating Tactile Feedback into VLA via Simulation

    Jun 10, 2026Siyu Ma, Yuqi Liang, Chang Yu +5TactileSimulation-Based Reinforcement Learning