Shortcut learning is a prevalent issue in robot learning. The limited diversity of robot demonstration datasets can mislead policies into exploiting spurious correlations between tasks and irrelevant features, such as viewpoint or background. Collecting sufficiently diverse robot demonstrations is costly and inefficient, motivating algorithmic alternatives. We focus on vision-language-action (VLA) models and discover that different vision-language model backbones exhibit substantially different levels of susceptibility to visual shortcut learning. We find that model behavior correlates with our proposed representation-level metric, action margin, which requires no policy rollouts. Visual shortcuts consistently enter action representations in early layers, with models differing in the extent to which later layers correct them by incorporating language information. To boost models' attention to language, we introduce task scrubbing, a new domain-adversarial training method that decreases models' likelihood of using visual shortcuts and improves VLAs' generalization. Experiments in both simulation and the real world across multiple VLAs and visual cues show that task scrubbing improves out-of-distribution robustness and often eliminates visual shortcut learning.
Figures & tables
Fig. 1: Visual shortcut-learning setup. During training (left), the model sees two different tasks with visual cues that match the language instruction. When the model is then evaluated (right) out-of-distribution where the visual cue does not match the language, the baseline model follows the visual cue rather than the language. Applying task scrubbing during training encourages the policy to follow the instruction.
Fig. 2: OOD success and SC across VLA variants and low, medium, and high task-correlated viewpoint disparity on LIBERO Spatial.
Fig. 3: Action margin shows a strong correlation with shortcut rate for the 21 model–disparity combinations. The representation-level metric strongly tracks shortcut reliance (Kendall’s τb=0.83 , p=1.35×10−7 ).
Fig. 4: Layer-wise attention-knockout effects for π0 , SmolVLA-V1, and SmolVLA. A positive value indicates that the removed attention promotes cue alignment, while negative values indicate it promotes task alignment. Removing visual attention generally shifts representations away from the cue, while language and action-token interactions can correct cue alignment in later layers.
Fig. 5: Task scrubbing makes task information less available in the visual features during training, so that the VLA backbone learns to pay more attention to the language information. Red denotes the auxiliary task-scrubbing head used during training, which makes the visual pathways less informative about task identity. Either option shown in blue allows the visual pathway to be trainable with little additional cost.
Fig. 6: Evaluation setup: The three different environments and task pairs used for evaluation.
Model
Condition
OOD Succ. (%)
SC (%)
ID Succ. (%)
Baseline
10±1
0±0
62±3
SmolVLA [ 27 ]
Scrubbed
73±3
0±0
74±4
Baseline
4±6
45±26
37±4
SmolVLA-V1
Scrubbed
64±9
0±0
66±3
Baseline
4±5
17±28
34±8
SmolVLA-VLM
Scrubbed
64±10
0±0
66±10
TABLE I: LIBERO Spatial results under high viewpoint disparity.
Fig. 7: Language share during training for SmolVLA, SmolVLA-VLM, SmolVLA-V1, π0 , and π0 without action pretraining. Task scrubbing increases the relative influence of language on the action representations. Action pretraining raises initial language share in both the SmolVLA and π0 comparisons, but pretrained models respond more slowly to scrubbing.
Model
Condition
OOD Succ. (%)
SC (%)
ID Succ. (%)
SmolVLA [ 27 ]
Baseline
5±2
0±0
98±3
Scrubbed
86±3
0±0
95±2
Baseline
0±0
100±0
90±11
SmolVLA-V1
Scrubbed
83±6
0±0
89±5
SmolVLA-VLM
Baseline
0±0
100±0
95±2
Scrubbed
45±15
0±0
50±11
TABLE II: Place-Block results with different backgrounds.
Cue
Model
Scrub.?
OOD Succ. (%)
SC (%)
ID Succ. (%)
SmolVLA [ 27 ]
No
77.5
0
95.0
Yes
95.0
2.5
95.0
No
57.5
17.5
95.0
Cup color
SmolVLA-V1
Yes
87.5
0
90.0
SmolVLA [ 27 ]
No
47.5
27.5
100.0
Yes
92.5
0
97.5
TABLE III: Real-world ALOHA performance.
Model
Method
OOD Succ. (%)
SC (%)
ID Succ. (%)
SmolVLA [ 27 ]
Scrubbed
73±3
0±0
74±4
Attention Reg.
67±4
0±0
67±3
Scrubbed
64±9
0±0
66±3
SmolVLA-V1
Attention Reg.
56±5
0±0
59±4
π0 [ 5 ]
Scrubbed
48±17
0±0
69±7
Attention Reg.
3±2
68±31
57±8
TABLE IV: Attention regularization versus task scrubbing.
Model
LIBERO
RoboTwin
Real robot
SmolVLA [ 27 ]
2
0.75
2
SmolVLA-V1
1
0.175
1
SmolVLA-VLM
2
0.25
–
MiniVLA [ 3 ]
12
–
–
π0 [ 5 ]
0.5
0.35
–
π0.5 [ 4 ]
0.8
0.55
–
TABLE V: Selected auxiliary-loss scales for task-scrubbing.
Vision-Language-Action (VLA) models are commonly pretrained on robot demonstrations by jointly mapping visual observations and language instructions to actions. However, dense visual-action supervision can dominate the comparatively sparse language-action signal. As a result, policies may rely on visual shortcuts rather than learn how language conditions action execution, making them sensitive to visual variations. To address this limitation, we propose LA4VLA, a language-action pretraining framework that enables policies to acquire language-conditioned action priors without visual observations. These priors capture reusable manipulation skills shared across tasks and scenes, reducing reliance on scene-specific visual cues. Specifically, LA4VLA decomposes expert demonstration trajectories into atomic action segments and pairs each segment with a corresponding low-level action description. This yields LA-33K, a dataset of 33K Language-Action (LA) episodes derived entirely from existing demonstrations without additional robot data collection. We further develop LA4VLA-1B, a lightweight 1B-parameter VLA model, and investigate three paradigms for incorporating language-action supervision into VLA learning: LA-only pretraining, sequential LA-to-VLA pretraining, and mixed LA-VLA pretraining. Across simulation and real-world tasks, LA-pretrained policies consistently outperform matched VLA-pretrained counterparts, while combining LA and VLA supervision leads to further gains. In particular, mixed LA-VLA pretraining improves the average success rate of LA4VLA-1B over the no-pretraining baseline by up to 17.8 and 45.0 percentage points in simulation and real-world tasks, respectively. These results establish LA4VLA as an effective and complementary pretraining strategy for building stronger and more robust VLA policies.
Tao Lin, Yuxin Du, Yiran Mao +13
School of AI, Shanghai Jiao Tong University · Alibaba Group · KAUST +1
Scaling robot data is crucial for building generalist Vision-Language-Action (VLA) models, yet robot trajectories are harder to scale than web-scale image-text data because embodied collection is costly and sparsely covers the physical world. This makes representation quality a central bottleneck: under a fixed robot-data budget, continued pre-training must turn limited trajectories into transferable visual-action knowledge rather than merely fit actions. We propose VLAct, a VLA-oriented VLM backbone trained on broad, heterogeneous, multi-embodiment robot data before task-specific fine-tuning. VLAct preserves the broad VLM prior and encourages shared action semantics across embodiments through VLM-prior preservation, multi-head continuous action co-supervision, and a partially unified cross-embodiment action layout, while allowing task-specific action heads during fine-tuning. Across simulation, real-world, and unseen-embodiment transfer, VLAct consistently improves downstream performance under fixed fine-tuning protocols. On LIBERO-Plus and RoboTwin 2.0, VLAct surpasses industrial VLA systems including ABot-M0 and LingBot-VLA, achieving success rates of 82.6% and 92.5%. On RoboDojo, VLAct ranks sixth among all policies by success rate and outperforms all explicitly designated world-action model (WAM) entries on both metrics. Most notably, on RoboCasa-GR1, an unseen humanoid embodiment, VLAct using only 20% of downstream trajectories outperforms the full-data GR00T-N1.6 baseline. These results are obtained using fully open-source data and only a 16-GPU training setup, showing that representation-centric continued pre-training can deliver highly competitive performance under a modest compute budget and is an important independent axis of VLA progress beyond data scaling.
Vision-Language-Action (VLA) models provide strong behavioral priors for robotic manipulation, yet efficiently adapting them to downstream tasks remains challenging. Recent work addresses this challenge by adapting frozen VLAs through online reinforcement learning (RL), whose sample efficiency depends on the quality of the state representation used by the actor and critic. Existing methods construct such representations either with VLA-independent visual encoders or through fixed compression of internal VLA representations. Neither design explicitly extracts the task-specific action-relevant VLA features most useful for downstream action refinement and action-value estimation, therefore limiting sample efficiency. To address this limitation, we introduce eRLT, which constructs an effective state representation by routing task-specific action-relevant information across both tokens and layers of the frozen VLA. Specifically, learned routing tokens dynamically aggregate visual-language features at multiple depths, while a lightweight layer router combines these summaries into a fixed-dimensional RL token. The routing module is initialized using expert demonstrations to capture features predictive of expert actions and then refined using critic feedback from online interactions for action-value estimation. Across seven LIBERO and RoboTwin tasks, eRLT improves mean normalized learning-curve AUC by up to 23.7% over representative baselines. Real-robot experiments on USB connector insertion and motherboard ribbon-cable insertion further show AUC improvements of 108.9% and 46.7%, respectively, over the strongest baseline.
Dehao Huang, Jianbang Liu, Jianpan Gao +7
Southern University of Science and Technology, Shenzhen, China. · Beijing Zhongguancun Academy, Beijing, China. · Samsung Robotics eXperience. +2