LLM agents that orchestrate frozen vision-language-action (VLA) policies improve across episodes through text memory, which records what the agent did but not how the task is done. A demonstration video shows it, but fits poorly into an agent's context. The full video slows every turn, fixed keyframes lose the contact detail that decides whether a grasp holds, and what the agent needs shifts from the task's structure while planning to the frames around each contact. We introduce Recursive Video In-Context Learning (RV-ICL), a training-free method that turns a demonstration into a hierarchy the agent navigates rather than a prompt it receives. The hierarchy is built from the sub-events of the demonstration, such as grasps and releases. Its levels grow finer, from keyframes of the whole task to phases, moments and short clips, and are exposed through read-only tools. The agent reads the coarse levels before planning. During execution it re-enters the hierarchy whenever a step needs more detail and loads only the clip of its current sub-goal. One demonstration per task is enough. Built on RPent, RV-ICL raises success from 92.6% to 96.5% on LIBERO-PRO and from 86.7% to 95.8% on LIBERO-Plus.
Figures & tables
Figure 1: Overview of RV-ICL on a task from the LIBERO-10 suite of LIBERO-PRO. The agent receives one demonstration of the task as a hierarchy (a) and consults it while solving the same task in a new layout (b). (a) The hierarchy has four levels: L0 event-driven keyframes (thin lines mark their time), L1 phases between grasps and releases, L2 moments around each sub-event, and L3 20-frame clips, here of the first grasp. (b) The tool calls of one solved evaluation episode (consecutive calls grouped, robot actions in grey, re-entries in orange). The agent reads L0 – L2 while planning and loads clips only during execution. Each re-entry loads the clip of the current sub-goal: the grasp, the release or the closing of the door.
Figure 2: Architecture of RV-ICL on RPent. Each demonstration is segmented offline at its sub-events into a four-level hierarchy. During an episode the LLM agent acts through the harness. Robot tools execute on the environment, and read-only demonstration tools load a level of the hierarchy into the agent’s context. The agent reads the coarse levels ( L0 – L2 ) before planning and loads a clip whenever its current sub-goal needs more detail.
Algorithm 1 One RV-ICL episode. ⊕ appends to the context, and the budget B caps the views of each phase or moment.
Method
Spat-T
Spat-S
Obj-T
Obj-S
Goal-T
Goal-S
L10-T
L10-S
Overall
VLA
OpenVLA ( Kim et al., 2024 )
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
π0 ( Black et al., 2025 )
0.0
0.0
0.0
2.0
0.0
0.0
0.0
0.0
0.3
π0.5 ( Physical Intelligence et al., 2025 )
1.0
20.0
1.0
17.0
2.0
38.0
1.0
8.0
11.0
MolmoAct ( Lee et al., 2025 )
0.0
0.0
0.0
6.0
0.0
0.0
6.0
0.0
1.5
NORA ( Hung et al., 2025 )
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
X-VLA ( Zheng et al., 2026 )
0.0
0.0
8.0
2.0
9.0
1.0
10.0
0.0
3.8
Table 1: Success rate (%) on LIBERO-PRO under instruction redirection (T) and position swap (S). Baseline cells aggregate 100 episodes (10 tasks × 10 seeds) and our cells 50 episodes (10 tasks × 5 seeds). All rows except ours are reported by Zhang et al. (2026c) ; “/” marks cells the original works do not report, and the overall score of Cap-X and RATS averages their six reported cells. HarnessVLA uses π0.5 fine-tuned by RLinf as its VLA primitive; RV-ICL adds the hierarchy to its GPT-6 Astra configuration. Bold marks the best result in each column.
Figure 3: How the agent uses the demonstration in one LIBERO-10 episode from our LIBERO-PRO runs. Top: the policy closes the gripper on top of the moka pot (1). The agent does not replay the whole video. It opens only the grasp clip, where the demonstrator takes the pot by its handle (2), and then re-grasps the pot by the handle (3). Bottom: across the episode the agent reads the keyframes and the phase map before acting, then opens one short clip at each sub-goal. Robot views are crops of the episode video, demonstration views are frames of the demonstration, and the quote is the agent’s own message.
Arm
Demonstration
Re-entry
L10-T
L10-S
Text
none
–
85.0
72.0
Keyframes
L0 keyframes
–
86.0
74.0
Full video
whole video
–
84.0
66.0
RV-ICL
L0 – L3
✓
90.0
84.0
Table 2: Ablations on the two LIBERO-10 cells of LIBERO-PRO, the longest-horizon tasks (GPT-6 Astra, success %). L0 denotes the keyframes and L3 the clips of the demonstration hierarchy ( Section 3.2 ). Bold marks the best result in each column.
Suite
Episodes
Clip views / episode
With re-entry
Spatial
100
2.8
98%
Object
100
2.3
98%
Goal
100
3.9
92%
LIBERO-10
100
10.4
98%
All
400
4.8
97%
Table 3: How the agent uses the demonstration in the 400 LIBERO-PRO episodes of Table 1 (T and S cells pooled). Clip views are re-entries during execution; the last column is the share of episodes with at least one.
Dimension
Lay.
Cam.
Rob.
Lang.
Light
Back.
Noise
All
HarnessVLA
82.4
83.3
82.4
100.0
76.5
88.2
94.1
86.7
+ RV-ICL
88.2
94.4
94.1
100.0
94.1
100.0
100.0
95.8
Suite
Spatial
Object
Goal
L10
HarnessVLA
100.0
96.7
76.7
73.3
+ RV-ICL
96.7
100.0
90.0
96.7
Table 4: Success rate (%) on LIBERO-Plus by perturbation dimension (top: object layout, camera viewpoint, robot initial state, language, lighting, background and sensor noise) and by suite (bottom). Bold marks the best result in each column.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Suite
Length (s)
Phases
Moments
Map tokens
Spatial
6.4
3.0
2.0
2.4k
Object
7.8
3.2
2.2
2.3k
Goal
6.8
2.7
2.0
2.2k
LIBERO-10
14.6
4.6
3.9
2.6k
Appendix
Table 7: The demonstration hierarchy per LIBERO suite over its ten base tasks: mean demonstration length, phases and moments, and the median size of the phase map ( L1 and L2 ) in tokens.
Method
Time (min)
Model calls
Input tokens (M)
Cached (%)
Output tokens (k)
Final context (k)
HarnessVLA
8.3
18
1.19
88.0
3.6
91.4
+ RV-ICL
14.6
20
1.71
88.6
4.5
109.3
Appendix
Table 8: Median cost of one episode on the 120 LIBERO-Plus variants: wall-clock time, model calls, input tokens and the share served from the cache, output tokens, and the size of the final context.