Vision-language-action (VLA) policies infer grasp actions primarily from visual observations and robot state, but do not explicitly represent the physical response observed after contact. We present MIM-VLA, a motor-feedback-based architecture that encodes recent gripper current, position, velocity, and signal validity as a 128-dimensional interaction token. A motor-only Motor Interaction Module (MIM) is pretrained with human-reviewed contact and interaction-phase labels and then conditions only the gripper-action pathway of SmolVLA; arm actions and the position-control interface remain unchanged. The same token supports the MEM selector VLM that compares candidate interactions and produces evidence-conditioned selections and explanations. We evaluate MIM-VLA in three real-world settings: comparing the interaction resistance of visually different objects, disambiguating visually similar real and replica objects through active probing, and gently grasping fragile objects, including held-out instances. Across 13 object pairs, MIM-VLA selects the higher-resistance object in 75.0% of trials, compared with 48.8% for the SmolVLA baseline. For the evaluated tasks, the approach uses motor feedback already available from the gripper and does not require an additional tactile array, force-torque sensor, calibrated force estimate, or direct current control.
Figures & tables
Fig. 1: Interaction-grounded candidate selection. The robot probes both objects, compares their motor responses, selects the object satisfying the requested physical property, and explains its decision.
Fig. 2: MIM-VLA architecture. Multi-view observations and language are processed by the frozen SmolVLM backbone. The MIM interaction token conditions gripper-action prediction and supports evidence-conditioned comparison through the Material Explain Module (MEM).
Fig. 3: Motor-only MIM pretraining. Human-reviewed contact and phase targets supervise auxiliary heads during training; video and the commanded position shown for annotation context are visualization only and are excluded from the MIM encoder input.
Fig. 4: Qualitative motor-response validation on the MIM pretraining data. The 18 legend entries correspond to the objects used for MIM pretraining. Small points summarize individual episodes, large markers denote object-level medians, and dashed contours indicate recurring response regions (C1–C4) in gripper-position change and mean current during squeeze . Objects with similar interaction resistance exhibit nearby response distributions despite differences in appearance.
Fig. 5: Candidate-specific interaction histories are summarized independently and compared by the Material Explain Module (MEM) to produce an object choice and evidence-conditioned explanation.
Exp.
Task
Evaluation protocol
Split
1
Cross-Appearance Hardness Comparison
Sequentially probe two visually different objects, select the harder object, and explain the relative response.
7 seen + 3 unseen pairs
2
Appearance-Matched Physical Disambiguation
Probe a real object and its appearance-matched soft replica, then identify the harder or genuine candidate from interaction evidence.
3 matched pairs
3
Gentle Grasping of Fragile Objects
From a fixed arm pose, detect contact and establish a stable hold with minimal compression.
Fig. 6: Shared probe–compare–select protocol for Experiments 1 and 2. The robot observes both candidates, probes each object, compares their interaction memories, and manipulates the selected candidate.
TABLE II: Harder-target selection results for Experiments 1 and 2. Each object pair is evaluated over 20 trials per policy. Best results are shown in bold.
Fig. 7: Qualitative evaluation of gentle grasping. Top: damage-free grasps of an unseen raw egg (left) and a seen Oreo cookie (right), each shown with its SAM 2 evaluation mask. Bottom: marshmallow deformation for MIM-VLA (left, 18.0% mask-area reduction) and the SmolVLA-flow baseline (right, 41.2%). SAM 2 masks are used only for evaluation.
TABLE III: Gentle-grasp results. Deformation is mean ± SD over 20 trials per object and policy; lower SAM 2 mask-area reduction is better. Breakable-object results report successful trials and rates; higher is better.
K
Total tokens ↓
Selection acc. ↑
Position F1 ↑
Current F1 ↑
1
2
97.6±0.7
78.1±3.8
77.3±10.3
2
4
98.1±0.7
78.4±1.1
81.3±3.8
4
8
99.5±0.7
71.4±9.2
78.0±3.2
8
16
96.2±1.8
78.0±2.4
67.1±8.1
TABLE IV: Summary-token capacity ablation. K is per candidate and total tokens count both candidates. Results are mean ± SD over three seeds.