cs.ROOct 7, 2026

Do Vision-Language-Action Models Understand Instructions? A Mechanistic Interpretability Study on Language Grounding

Authors: Theodor Wulff, Angelo Cangelosi

Organizations: Department of Computer Science The University of Manchester Manchester, United Kingdom

Abstract

Vision-Language-Action models are designed to generalise across environments and task descriptions, raising the question of whether their action generation actually depends on the language instruction, or whether they largely rely on visual cues and superficial correlations. Robustness to variance in the visual and linguistic observation space is critical for real-world deployment, yet VLAs lack explicit grounding modules and instead rely on the intrinsic language grounding capabilities of their Vision-Language model backbones. For this reason, we conduct a controlled mechanistic interpretability study on the language grounding capabilities of two state-of-the-art Vision-Language-Action models, π0.5π_{0.5} and GR00T N1.7, by applying activation and attribution patching to the residual stream of the action generation modules. We systematically corrupt the task instruction of input samples of the LIBERO benchmark following five strategies: synonym replacement, semantic scaling, directional corruption, random object substitution, and empty string. Our experiments find that both models are comparatively insensitive to abstract rephrasing and to referencing non-existent objects, but react strongly to empty task descriptions and, especially, to directional language. During action generation, this sensitivity is concentrated in different loci for each model: mainly in the early, periodic cross-attention layers for GR00T N1.7, versus distributed across the earliest and selected later layers for π0.5π_{0.5}. For GR00T N1.7, directional perturbations drive some of the largest causal effects while leaving the internal representational geometry comparatively unchanged, a dissociation we do not observe clearly for π0.5π_{0.5}. Finally, the reliability of attribution patching is model-dependent: it closely tracks activation patching for GR00T N1.7 but not for π0.5π_{0.5}.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Grounded Semantic Re-Binding for Robust Instruction Generalization in Vision-Language-Action Models

    Aug 3, 2026Zhaokai Yin, Zhipeng ZhangRobotic ManipulationAction Expert

  2. Restoring Linguistic Grounding in VLA Models via Train-Free Attention Recalibration

    Mar 6, 2026Ninghao Zhang, Bin Zhu, Shijie Zhou +1Robot PoliciesLinguistics