cs.ROSep 29, 2026

Correcting WHERE, Preserving HOW: Compositional Generalization for Vision-Language-Action Models via Referential Guidance

Authors: Yanyan Zhang, Disheng Liu, Xinpeng Li, Chaoda Song, Mohsen Hariri, Debargha Ganguly, Wang Yang, Kai Ye, +3 more

Organizations: Case Western Reserve University Cleveland, OH, USA

Abstract

While Vision-Language-Action (VLA) models enable flexible action generation, their generalization across diverse environmental elements, including manipulated objects, destinations, and backgrounds, is limited by the lack of diversity in robotic training data. Trained end-to-end on such data, VLAs tend to exploit visual shortcuts, associating actions with task-irrelevant visual features rather than the intended task semantics. These shortcuts block recomposition of elements already seen by the policy, that is, compositional generalization. Existing approaches mitigate such entanglement through task-relevant perception or targeted data diversification, but offer no explicit mechanism for unseen recomposition and require backbone-specific modifications with retraining. We observe that under such recomposition, VLAs often fail at global grounding while retaining local manipulation skills that recover near the correct target in familiar configurations. Therefore, we propose Referential Guidance (ReGuide), a training-free wrapper that, given object poses from a grounding module, combines semantic and geometric rebinding to guide the end-effector into demonstration-supported configurations of the instructed referent, where the frozen policy can resume execution. Experiments in simulation across multiple VLA backbones as well as on a real robot show that ReGuide improves success rates under compositional shifts by up to 56.8 and 75.0 percentage points, respectively, while preserving standard-task performance.

Figures & tables

Appendix figures & tables31 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Grounded Semantic Re-Binding for Robust Instruction Generalization in Vision-Language-Action Models

    Aug 3, 2026Zhaokai Yin, Zhipeng ZhangRobotic ManipulationAction Expert

  2. GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization

    May 12, 2026Xiaosong Jia, Bowen Yang, Zuhao Ge +17Recent Vision-Language ModelsGuidance

  3. AC-VLA: Robust Out-of-Distribution Action Execution via Compositional Learning

    Jul 17, 2026Xiaojiang Peng, Kai Peng, Jie Lu +3Robotic ManipulationSpatial Grounding