cs.ROOct 6, 2026

MIM-VLA: Learning Physical Interaction Representations from Gripper Motor Feedback

Authors: Jaeyoung Lee, Jiyeon Koo, Taehwa Kim, Yerin Cha, Andrew Jaeyong Choi

Organizations: School of Computing, iRASC Lab., Gachon University, Seongnam, Republic of Korea

Abstract

Vision-language-action (VLA) policies infer grasp actions primarily from visual observations and robot state, but do not explicitly represent the physical response observed after contact. We present MIM-VLA, a motor-feedback-based architecture that encodes recent gripper current, position, velocity, and signal validity as a 128-dimensional interaction token. A motor-only Motor Interaction Module (MIM) is pretrained with human-reviewed contact and interaction-phase labels and then conditions only the gripper-action pathway of SmolVLA; arm actions and the position-control interface remain unchanged. The same token supports the MEM selector VLM that compares candidate interactions and produces evidence-conditioned selections and explanations. We evaluate MIM-VLA in three real-world settings: comparing the interaction resistance of visually different objects, disambiguating visually similar real and replica objects through active probing, and gently grasping fragile objects, including held-out instances. Across 13 object pairs, MIM-VLA selects the higher-resistance object in 75.0% of trials, compared with 48.8% for the SmolVLA baseline. For the evaluated tasks, the approach uses motor feedback already available from the gripper and does not require an additional tactile array, force-torque sensor, calibrated force estimate, or direct current control.

Figures & tables

Explore similar work

CardsList
  1. TacCoRL: Integrating Tactile Feedback into VLA via Simulation

    Jun 10, 2026Siyu Ma, Yuqi Liang, Chang Yu +5TactileSimulation-Based Reinforcement Learning

  2. FM-VLA: Force-based Memory for Vision-Language-Action Models in Contact-Rich Manipulation

    Jul 20, 2026Ruicheng Li, Qixiu Li, Ruichun Ma +8Robotic ManipulationContact-Rich Manipulation

  3. VISTA: Vision-Grounded and Physics-Validated Adaptation of UMI data for VLA Training

    Jun 3, 2026Siyuan Yang, Linzheng Guo, Ouyang Lu +10Vision-Language-Action FrameworkRobotic Data Acquisition