cs.ROOct 7, 2026

When Listening Becomes Easier: Scrubbing Visual Cues for Shortcut-Free VLAs

Authors: Jasper Gerigk, Kenzo Aspuru-Takata, Chin-Hsuan Wu, Mohammad Mohammadi, Shuhong Zheng, Igor Gilitschenski

Organizations: University of Toronto, Toronto, ON M5S 3H5, Canada · Vector Institute, Toronto, ON M5G 0C6, Canada

Abstract

Shortcut learning is a prevalent issue in robot learning. The limited diversity of robot demonstration datasets can mislead policies into exploiting spurious correlations between tasks and irrelevant features, such as viewpoint or background. Collecting sufficiently diverse robot demonstrations is costly and inefficient, motivating algorithmic alternatives. We focus on vision-language-action (VLA) models and discover that different vision-language model backbones exhibit substantially different levels of susceptibility to visual shortcut learning. We find that model behavior correlates with our proposed representation-level metric, action margin, which requires no policy rollouts. Visual shortcuts consistently enter action representations in early layers, with models differing in the extent to which later layers correct them by incorporating language information. To boost models' attention to language, we introduce task scrubbing, a new domain-adversarial training method that decreases models' likelihood of using visual shortcuts and improves VLAs' generalization. Experiments in both simulation and the real world across multiple VLAs and visual cues show that task scrubbing improves out-of-distribution robustness and often eliminates visual shortcut learning.

Explore similar work

CardsList
  1. LA4VLA: Learning to Act without Seeing via Language-Action Pretraining

    Jun 25, 2026Tao Lin, Yuxin Du, Yiran Mao +13Language Model PretrainingVisuomotor Policy Learning

  2. eRLT: Efficient VLA Reinforcement Learning via Action-Relevant Token Routing

    Oct 1, 2026Dehao Huang, Jianbang Liu, Jianpan Gao +7Efficient VLA ModelsRobotic Manipulation