cs.ROOct 5, 2026

Recursive Video In-Context Learning for Agentic Robot

Authors: Wenrui Bao, Xinxin Liu, Bingxin Xu, Yuzhang Shang

Organizations: University of Central Florida · University of Southern California

Abstract

LLM agents that orchestrate frozen vision-language-action (VLA) policies improve across episodes through text memory, which records what the agent did but not how the task is done. A demonstration video shows it, but fits poorly into an agent's context. The full video slows every turn, fixed keyframes lose the contact detail that decides whether a grasp holds, and what the agent needs shifts from the task's structure while planning to the frames around each contact. We introduce Recursive Video In-Context Learning (RV-ICL), a training-free method that turns a demonstration into a hierarchy the agent navigates rather than a prompt it receives. The hierarchy is built from the sub-events of the demonstration, such as grasps and releases. Its levels grow finer, from keyframes of the whole task to phases, moments and short clips, and are exposed through read-only tools. The agent reads the coarse levels before planning. During execution it re-enters the hierarchy whenever a step needs more detail and loads only the clip of its current sub-goal. One demonstration per task is enough. Built on RPent, RV-ICL raises success from 92.6% to 96.5% on LIBERO-PRO and from 86.7% to 95.8% on LIBERO-Plus.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ICI-VLA: In-Context Imitation with Spatiotemporally Aligned Demonstrations for Vision-Language-Action Models

    Sep 7, 2026Songhua Yang, Ziyu Liu, Xuetao Li +3In-ContextMultimodal Fusion

  2. In-Context Robot Learning with VLM Agents

    Sep 16, 2026Dongzhou Cheng, Taoran Yi, Ye Fang +12Scalable Robot LearningRobot Systems

  3. Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents

    Jul 9, 2026Yixian Zhang, Huanming Zhang, Feng Gao +14Goal-Conditioned Dynamic ManipulationVisuomotor Control