cs.ROAug 26, 2026

RA-VLA: Retrieval-Augmented VLA for Test-Time Adaptation

Authors: Sanghwan Jang, Minjin Jeon, Minsoo Kim, Seongjin Choi, Dongha Kim, Hwanjo Yu

Abstract

Vision-Language-Action (VLA) models provide a versatile foundation for general robotic manipulation, yet they exhibit significant brittleness when confronted with novel task distributions. While In-Context Imitation Learning (ICIL) offers a training-free alternative, existing frameworks suffer from an adaptation bottleneck that hinders the effective translation of expert context to executable actions. This failure originates from superficial retrieval mechanisms and an inherent behavioral inertia that anchors the policy to its pre-trained priors. To address these limitations, we present RA-VLA, a retrieval-augmented VLA framework that integrates behavior-aligned context retrieval with a grounded execution pipeline. By enforcing faithful adherence to functional cues within a scalable architecture, RA-VLA facilitates seamless task adaptation while preserving inference efficiency. Our empirical evaluations across the LIBERO benchmark and a real-world UR5e environment demonstrate that RA-VLA achieves superior success rates and computational efficiency, establishing a robust framework for training-free robotic adaptation.

Explore similar work

CardsList
  1. ICI-VLA: In-Context Imitation with Spatiotemporally Aligned Demonstrations for Vision-Language-Action Models

    Sep 7, 2026Songhua Yang, Ziyu Liu, Xuetao Li +3Imitation LearningRobotic Manipulation

  2. Inline Memory Meets Reusable Skills: Memory-centric Framework for Vision-Language-Action Model

    Sep 30, 2026Zaijing Li, Rui Shao, Bing Hu +3Memory-Augmented VLMsRobot Skill Learning