When Does Retrieval Help? A Study of In-Context Adaptation in Vision-Language-Action Models
Organizations: Department of Computer Science, Tulane University · Department of Computer Science, Aalto University · Department of Industrial Engineering and Management, Aalto University
Abstract
Vision-language-action (VLA) models have shown strong potential as generalist robot policies, but adapting them to unseen tasks often requires costly parameter updates. Recent work such as RICL introduces in-context adaptability by retrieving expert demonstrations based on the current VLA observation and providing them as additional context at test time. The effectiveness of this adaptation therefore depends critically on the retrieval mechanism. In this work, we systematically study how different retrieval methods affect both retrieval quality and task performance within the RICL framework. Specifically, we compare four different methods: image-based retrieval, retrieval augmented with VLA's state, retrieval using features from the VLA backbone, and random retrieval. Our experiments yield three main findings. First, no retrieval method consistently dominates the others in task success, while surprisingly, random retrieval achieves a non-trivial success rate. Second, standard retrieval-quality diagnostics do not reliably reflect downstream VLA performance. Third, demonstrations from different but related tasks can provide useful transferable information. Together, these results provide an initial step toward understanding how retrieval mechanisms shape the in-context learning capability of VLA models and their downstream task performance, while highlighting the need for more careful design and evaluation of retrieval mechanisms for reliable test-time adaptation.
Figures & tables
| Method | Succ. | Task-hier. CI |
|---|---|---|
| RICL-Image | 40.0 | [27.2, 54.0] |
| RICL-ImageState | 36.8 | [24.0, 50.0] |
| RICL-VLAFeat | 40.4 | [28.4, 52.8] |
| RICL-Random | 26.0 | [15.6, 37.6] |
| Exact-task retrieval rate | NDCG | ||
|---|---|---|---|
| Method | Offline | On-Policy | |
| RICL-Image | 87.2% | 51.8% | 0.7461 |
| RICL-ImageState | 87.9% | 48.9% | 0.7897 |
| RICL-VLAFeat | 83.1% | 38.7% | 0.7280 |
| RICL-Random | 4.0% | 4.1% | 0.0199 |
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
| ID | Task Instruction | ID | Task Instruction |
|---|---|---|---|
| 20 | turn on the stove | 49 | pick up the tomato sauce and put it in the basket |
| 21 | turn on the stove and put the frying pan on it | 50 | pick up the alphabet soup and put it in the basket |
| 24 | put the black bowl in the bottom drawer of the cabinet | 51 | pick up the butter and put it in the basket |
| 25 | put the black bowl on top of the cabinet | 52 | pick up the milk and put it in the basket |
| 28 | close the top drawer of the cabinet | 55 | pick up the alphabet soup and put it in the tray |
| 29 | put the black bowl in the top drawer of the cabinet | 56 | pick up the butter and put it in the tray |