RA-VLA: Retrieval-Augmented VLA for Test-Time Adaptation
Authors: Sanghwan Jang, Minjin Jeon, Minsoo Kim, Seongjin Choi, Dongha Kim, Hwanjo Yu
Organizations: Department of Computer Science and Engineering, POSTECH, Pohang, South Korea · Graduate School of Artificial Intelligence, POSTECH, Pohang, South Korea
Vision-Language-Action (VLA) models provide a versatile foundation for general robotic manipulation, yet they exhibit significant brittleness when confronted with novel task distributions. While In-Context Imitation Learning (ICIL) offers a training-free alternative, existing frameworks suffer from an adaptation bottleneck that hinders the effective translation of expert context to executable actions. This failure originates from superficial retrieval mechanisms and an inherent behavioral inertia that anchors the policy to its pre-trained priors. To address these limitations, we present RA-VLA, a retrieval-augmented VLA framework that integrates behavior-aligned context retrieval with a grounded execution pipeline. By enforcing faithful adherence to functional cues within a scalable architecture, RA-VLA facilitates seamless task adaptation while preserving inference efficiency. Our empirical evaluations across the LIBERO benchmark and a real-world UR5e environment demonstrate that RA-VLA achieves superior success rates and computational efficiency, establishing a robust framework for training-free robotic adaptation.
Figures & tables
Figure 1 : Adaptation bottleneck in existing ICIL frameworks. Existing ICIL methods (a) prioritize superficial visual similarity over functional intent in context retrieval, (b) fail to effectively leverage contextual guidance, and (c) exhibit significant computational overhead for larger contexts. While this overhead is plotted for RICL, other ICIL baselines follow a nearly identical trend.
Figure 2 : Overview of RA-VLA. Our framework retrieves task-relevant expert segments to serve as in-context guidance, facilitating action generation for unseen task adaptation. The architecture integrates two key components: a retriever trained to identify behaviorally similar segments (Section 4.2 ) and a VLA policy finetuned to effectively ground action generation on the retrieved context (Section 4.3 ).
Retriever
Method
LIBERO-Spatial
LIBERO-Object
LIBERO-Goal
LIBERO-Long
Average
–
Vanilla VLA
0.000
0.066
0.002
0.000
0.0170
Off-the-shelf (SigLIP 2)
RAEA
0.060
0.106
0.008
0.006
0.0450
RICL G
0.092
0.156
0.106
0.000
0.0885
RICL R
0.048
0.052
0.086
0.000
0.0465
Action-aware (Ours)
RAEA
0.164
0.182
0.104
0.044
0.1235
RICL G
0.156
0.198
0.116
0.048
0.1295
Table 1 : Novel task adaptation performance on the LIBERO benchmark. All methods are evaluated on an entirely unseen task suite, utilizing three expert demonstrations per each held-out task as in-context guidance. Baselines are evaluated with both off-the-shelf SigLIP 2 and our SigLIP 2-based action-aware retriever. We report the mean success rate (%) for each task suite, and Average denotes the average success rate across all task suites. Bold and underline indicate the best and second-best performances, respectively.
Stack Box
Throw Trash
Close Drawer
Press Pedal
Method
Pick Box
Success
Pick Trash
Success
Touch Drawer
Success
Touch Pedal
Success
Average
Vanilla VLA
0.000
0.000
0.083
0.000
0.500
0.333
0.000
0.000
0.0833
RAEA
0.000
0.000
0.250
0.000
0.667
0.500
0.000
0.000
0.1250
RICL R
0.583
0.250
0.833
0.250
0.667
0.583
0.667
0.333
0.3542
RA-VLA
0.750
0.417
0.917
0.750
0.750
0.667
0.583
0.417
0.5625
Table 2 : Novel task adaptation performance on the real-world UR5e environment. All methods are evaluated on an entirely unseen task suite, utilizing four expert demonstrations per held-out task as in-context guidance. Baselines are evaluated with our SigLIP 2-based action-aware retriever. We report both sub-goal and overall success rates for each task.
Figure 3 : Qualitative comparison of policy rollouts on the real-robot Stack Box task. While all frameworks utilize our action-aware retriever to obtain expert context, only RA-VLA precisely executes unseen instructions by effectively leveraging this guidance. In contrast, baselines fail by reverting to pre-trained priors or lacking control precision. For further qualitative examples, see Figures 6 – 9 .
Method
Sctx
LIBERO-Goal
RAEA
0.0247
0.104
RICL R
0.0788
0.206
RA-VLA w/o Ladhere
0.0353
0.098
RA-VLA
0.3639
0.532
Table 3 : Contextual sensitivity. The Relative Contextual Sensitivity Sctx is averaged over 1,000 samples. We observe a positive trend where Sctx correlates with increased success rates.
Figure 4 : Visualization of segment-level trajectory alignment. For a given query segment, the retrieval model searches for the most relevant segment within the source trajectory. For instance, for the first segment of the query trajectory, the ViT-based retriever identifies the first image in the retrieval results as the most relevant match. However, while the query segment depicts the initial pick-up phase, the retrieved segment corresponds to a later transport phase, illustrating a failure in capturing behavioral relevance.
Figure 5 : Inference scalability. The latency is averaged over 1,000 inference runs for the generation of a single action chunk. We use 4 denoising steps for the action head. Unlike existing ICIL approaches, RA-VLA enables the seamless integration of multiple expert segments. RAEA shows a similar trend to RICL, as illustrated in Figure 11 .
1
2
3
4
N
0.482
0.518
0.532
0.552
K
0.532
0.540
0.548
0.546
Table 4 : Impact of buffer size N and retrieval size K . We report the average success rate (%) on the LIBERO-Goal.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
K
Segment Assignment
1
[1, 1, 1, 1, 1, 1, 1, 1]
2
[2, 2, 2, 2, 1, 1, 1, 1]
3
[3, 3, 2, 2, 2, 1, 1, 1]
4
[4, 4, 3, 3, 2, 2, 1, 1]
5
[5, 4, 3, 3, 2, 2, 1, 1]
6
[6, 5, 4, 3, 2, 2, 1, 1]
Appendix
Table 5 : Segment assignment patterns across different values of K .
Figure 6 : Qualitative comparison of policy rollouts for the ‘put the wine bottle on top of the cabinet’ task in LIBERO-Goal. While all frameworks utilize our action-aware retriever to obtain expert context, only RA-VLA precisely executes unseen instructions by effectively leveraging this guidance. In contrast, baselines fail by reverting to pre-trained priors or lacking control precision.
Figure 7 : Qualitative comparison of policy rollouts on the real-robot Throw Trash task. While all frameworks utilize our action-aware retriever to obtain expert context, only RA-VLA precisely executes unseen instructions by effectively leveraging this guidance. In contrast, baselines fail by reverting to pre-trained priors or lacking control precision.
Figure 8 : Qualitative comparison of policy rollouts on the real-robot Close Drawer task. For this specific task, the vanilla VLA exhibits a notable degree of adherence to the given instruction, suggesting that its pre-training distribution likely encompasses demonstrations relevant to drawer-closing maneuvers.
Figure 9 : Qualitative comparison of policy rollouts on the real-robot Press Pedal task. While all frameworks utilize our action-aware retriever to obtain expert context, only RA-VLA precisely executes unseen instructions by effectively leveraging this guidance. In contrast, baselines fail by reverting to pre-trained priors or lacking control precision.
Figure 10 : Qualitative comparison of policy rollouts with and without contextual adherence loss Ladhere . In the absence of the contextual adherence loss, the policy fails to align with the expert guidance and instead (a) reverts to its pre-trained priors or (b) produce erratic actions.
Figure 11 : Inference scalability. The latency is averaged over 1,000 inference runs for the generation of a single action chunk. We use 4 denoising steps for the action head. Unlike existing ICIL approaches, RA-VLA enables the seamless integration of multiple expert segments.
Vision-Language-Action (VLA) models have emerged as a promising paradigm for robotic manipulation by leveraging pre-trained vision-language representations. However, current VLA training methods suffer from two critical limitations: poor generalization to novel environments and low training efficiency requiring extensive demonstrations. We introduce Agentic-VLA, an agentic training framework that enables VLAs to efficiently adapt online through three key innovations: (1) Adaptive Reward Synthesis, which dynamically generates and adjusts reward functions based on the VLA's current capabilities and task complexity, decomposing complex tasks into learnable sub-goals for curriculum learning; (2) Language-Guided Exploration, where a critic model provides structured guidance for systematic exploration rather than random sampling; and (3) Experience Memory,which stores and retrieves task-relevant policy weights for warm-starting adaptation to similar tasks. We evaluate Agentic-VLA on the LIBERO benchmark, achieving substantial improvements: +12.3% on long-horizon tasks, +28.5% in 1-shot learning, and enabling cross-task transfer from 0% to 31.2% without task-specific demonstrations. Our framework also demonstrates 2.4x faster convergence compared to existing online adaptation methods. Beyond LIBERO, Agentic-VLA retains its advantage on the dual-arm RoboTwin 2.0 benchmark, including under its randomized Hard setting. These results establish Agentic-VLA as a significant step toward truly adaptive VLA systems capable of continuous learning in deployment.
Vision-Language-Action (VLA) policies are commonly adapted to new manipulation settings through additional gradient updates, which limits rapid deployment when task-specific data or compute is scarce. We present ICI-VLA, a training and retrieval framework that equips a text-action VLM with few-shot test-time adaptation through in-context demonstrations. Unlike mainstream VLA designs based on action-specific multimodal fusion, ICI-VLA retains the native text-generation interface. ICI-VLA updates its parameters only during offline training; at inference, the policy remains fixed and conditions action generation on retrieved micro-demonstrations. The framework decomposes long trajectories into short, semantically labeled examples and trains an RD-Encoder with positives mined by Dynamic Time Warping (DTW), aligning the retrieved context with the phase and geometry of the current subtask. We further introduce Target Action Masking, a context-corruption objective designed to reduce direct action copying and increase reliance on the current observation. ICI-VLA reaches average success rates of 97.7% on LIBERO and 60.4% on RoboTwin 2.0, exceeding the highest reported baseline average on RoboTwin 2.0 by 19.3 percentage points. It also achieves 83.2% across four physical tasks. These results indicate that a fixed VLA policy can benefit from conditioning on spatiotemporally aligned demonstrations at test time.
Songhua Yang, Ziyu Liu, Xuetao Li +3
Wuhan University · Wuhan University, Wuhan, China · Institute of Technological Sciences, Wuhan University +1
Vision-Language-Action (VLA) models have shown strong promise for general-purpose robotic manipulation, yet adapting them to new tasks and domains remains inefficient: existing methods often rely on parameter tuning, incurring substantial costs and risking catastrophic forgetting of previously learned tasks. To address this, we propose \textbf{Optimus-R}, a memory-centric VLA framework that formulates robotic adaptation as explicit query-skill memory tuning. Optimus-R introduces: (i) An \textbf{Inline Memory Interface for skill extraction}. It inserts learnable memory tokens into the VLA prefix stream, allowing the backbone to derive control-aware query and skill representations within the native action-conditioning pathway. (ii) A \textbf{Query-Skill Memory Bank for skill learning}. It externalizes skills into query prototypes for deciding \emph{what} to retrieve and skill values for specifying \emph{how} to act, supporting skill reuse and expansion with limited parameter updates. (iii) A lightweight \textbf{Bridge-and-Adapt mechanism for skill updating}. It aligns target-domain queries and skills with the existing memory space through a lightweight adapter and residual memory updates. Experiments on in-domain adaptation, cross-domain adaptation, and lifelong learning show that Optimus-R enables data-efficient skill learning while mitigating catastrophic forgetting.
Zaijing Li, Rui Shao, Bing Hu +3
Harbin Institute of Technology (Shenzhen) · Pengcheng Laboratory