Vision-language-action (VLA) policies infer grasp actions primarily from visual observations and robot state, but do not explicitly represent the physical response observed after contact. We present MIM-VLA, a motor-feedback-based architecture that encodes recent gripper current, position, velocity, and signal validity as a 128-dimensional interaction token. A motor-only Motor Interaction Module (MIM) is pretrained with human-reviewed contact and interaction-phase labels and then conditions only the gripper-action pathway of SmolVLA; arm actions and the position-control interface remain unchanged. The same token supports the MEM selector VLM that compares candidate interactions and produces evidence-conditioned selections and explanations. We evaluate MIM-VLA in three real-world settings: comparing the interaction resistance of visually different objects, disambiguating visually similar real and replica objects through active probing, and gently grasping fragile objects, including held-out instances. Across 13 object pairs, MIM-VLA selects the higher-resistance object in 75.0% of trials, compared with 48.8% for the SmolVLA baseline. For the evaluated tasks, the approach uses motor feedback already available from the gripper and does not require an additional tactile array, force-torque sensor, calibrated force estimate, or direct current control.
Figures & tables
Fig. 1: Interaction-grounded candidate selection. The robot probes both objects, compares their motor responses, selects the object satisfying the requested physical property, and explains its decision.
Fig. 2: MIM-VLA architecture. Multi-view observations and language are processed by the frozen SmolVLM backbone. The MIM interaction token conditions gripper-action prediction and supports evidence-conditioned comparison through the Material Explain Module (MEM).
Fig. 3: Motor-only MIM pretraining. Human-reviewed contact and phase targets supervise auxiliary heads during training; video and the commanded position shown for annotation context are visualization only and are excluded from the MIM encoder input.
Fig. 4: Qualitative motor-response validation on the MIM pretraining data. The 18 legend entries correspond to the objects used for MIM pretraining. Small points summarize individual episodes, large markers denote object-level medians, and dashed contours indicate recurring response regions (C1–C4) in gripper-position change and mean current during squeeze . Objects with similar interaction resistance exhibit nearby response distributions despite differences in appearance.
Fig. 5: Candidate-specific interaction histories are summarized independently and compared by the Material Explain Module (MEM) to produce an object choice and evidence-conditioned explanation.
Exp.
Task
Evaluation protocol
Split
1
Cross-Appearance Hardness Comparison
Sequentially probe two visually different objects, select the harder object, and explain the relative response.
7 seen + 3 unseen pairs
2
Appearance-Matched Physical Disambiguation
Probe a real object and its appearance-matched soft replica, then identify the harder or genuine candidate from interaction evidence.
3 matched pairs
3
Gentle Grasping of Fragile Objects
From a fixed arm pose, detect contact and establish a stable hold with minimal compression.
Fig. 6: Shared probe–compare–select protocol for Experiments 1 and 2. The robot observes both candidates, probes each object, compares their interaction memories, and manipulates the selected candidate.
TABLE II: Harder-target selection results for Experiments 1 and 2. Each object pair is evaluated over 20 trials per policy. Best results are shown in bold.
Fig. 7: Qualitative evaluation of gentle grasping. Top: damage-free grasps of an unseen raw egg (left) and a seen Oreo cookie (right), each shown with its SAM 2 evaluation mask. Bottom: marshmallow deformation for MIM-VLA (left, 18.0% mask-area reduction) and the SmolVLA-flow baseline (right, 41.2%). SAM 2 masks are used only for evaluation.
TABLE III: Gentle-grasp results. Deformation is mean ± SD over 20 trials per object and policy; lower SAM 2 mask-area reduction is better. Breakable-object results report successful trials and rates; higher is better.
K
Total tokens ↓
Selection acc. ↑
Position F1 ↑
Current F1 ↑
1
2
97.6±0.7
78.1±3.8
77.3±10.3
2
4
98.1±0.7
78.4±1.1
81.3±3.8
4
8
99.5±0.7
71.4±9.2
78.0±3.2
8
16
96.2±1.8
78.0±2.4
67.1±8.1
TABLE IV: Summary-token capacity ablation. K is per candidate and total tokens count both candidates. Results are mean ± SD over three seeds.
Vision-language-action (VLA) models provide strong visual, language, and action priors for robot manipulation, but visual observations alone often miss the local contact state required for contact-rich tasks. We present TacCoRL, a scalable framework that injects Tactile feedback into VLA policies and improves them through sim-real Co-training and simulation-based reinforcement learning (RL), without requiring large-scale tactile pretraining or extensive real-world contact exploration. The key idea is not only adding touch as an input, but learning how contact readings should modulate action responses in near-failure states that are rare in demonstrations and risky to collect on hardware. We use a real-aligned simulator as a closed-loop training environment for contact interaction. Mixed simulated and real trajectories first warm-start tactile-conditioned actions in the pretrained policy. Reinforcement learning with verifiable task rewards then optimizes the policy using simulated contact rollouts. It reinforces tactile-conditioned actions that lead to task completion, while a supervised objective on real trajectories keeps the refined policy anchored to deployment visual, tactile, and action distributions. The resulting policy transfers directly to the real robot without privileged simulation state or online real-world RL. Across four bimanual contact-rich tasks, the final visuo-tactile policy achieves an average success rate of 72.5%, compared to baseline of 50.0%. Result videos and more details are available at https://tac-corl.github.io/
Siyu Ma, Yuqi Liang, Chang Yu +5
University of California, Los Angeles · University of California, San Diego · 4Peking University +1
Vision-language-action (VLA) models have achieved impressive generalization in robotic manipulation, and recent memory-augmented VLAs have relaxed the Markovian assumption by conditioning on past images or language summaries. Vision-based memory approaches address this by conditioning on sampled past image frames, but they are computationally expensive and fundamentally limited when temporal events are visually ambiguous, e.g., pushing a button multiple times with small movements. We propose FM-VLA, a VLA model with force-based memory, enabling temporal context reasoning for non-Markovian, contact-rich manipulation. We encode force histories into compact force memory tokens with a variational autoencoder (VAE) pretrained with force time series reconstruction. By projecting force latent representations and short state history as additional conditioning tokens to the action expert module, we enable VLAs to leverage accumulated contact event history to guide manipulation. We evaluate FM-VLA on three memory-dependent tasks, including finding a hidden block, pressing a button, and wiping a dish for a specific number of times. Our lightweight force memory achieves over 80% success rate with minimal inference overhead, significantly outperforming baseline approaches. Project page: https://qft-333.github.io/FM-VLA-Page/
Ruicheng Li, Qixiu Li, Ruichun Ma +8
Tsinghua University · Microsoft Research · Fudan University +1
Universal Manipulation Interface (UMI) enables scalable real-world robot data collection without hardware-specific teleoperation, yet leveraging UMI data to train large-scale Vision-Language-Action (VLA) models remains fundamentally challenging. We identify two critical mismatches: wrist-mounted fisheye views, with severe radial distortion and local gripper-centric perspectives, are out-of-distribution for pretrained VLMs; and human-collected trajectories frequently violate kinematic limits, incur collisions, or exceed controller bandwidth, teaching VLA policies physically infeasible actions. To address the challenges, we present VISTA, a framework that bridges this dual gap through three synergistic components. (i)~UMI-VQA, the first large-scale VQA dataset tailored to wrist-mounted fisheye observations, aligns VLM representations to the distorted visual regime via auxiliary vision-language supervision. (ii)~A systematic physical-validation pipeline performs a data-completeness pre-check and scores each valid trajectory for trajectory continuity, self-collision risk, and execution fidelity before it enters training. (iii)~A two-stage co-training recipe jointly learns vision-language grounding on UMI-VQA and action prediction on validated trajectories. Our experiments empirically show that incorporating UMI-VQA consistently improves downstream policy performance, and that physical-validation scores are strongly predictive of deployment success. On diverse simulation and real-world manipulation tasks, VISTA significantly outperforms strong baselines including π0.5, LingBot-VLA, and Wall-X. We release the physical-validation pipeline, UMI-VQA, validated trajectory data, and the pre-trained model for the community.
Siyuan Yang, Linzheng Guo, Ouyang Lu +10
Institute of AI (TeleAI), China Telecom · University of Science and Technology of China · 4Northwestern Polytechnical University +5