General-purpose vision-language models offer a promising way to zero-shot robot control: \gptastra{} excels at open-ended and language- or image-conditioned manipulation but remains substantially weaker on high-precision and long-horizon tasks. We introduce \emph{RoboICL}, an in-context robot-control framework that narrows these gaps without robot-specific parameter updates or a learned VLA. RoboICL separates \emph{demonstration context}, which provides recorded examples when available, from \emph{interaction memory}, which accumulates the model's own actions and observed outcomes. Both use a shared observation--action--receipt--observation grammar. To preserve experience across task stages, RoboICL combines sampled demonstration blocks with bounded anchored memory. Fixed anchors keep earlier rollout interactions available for in-context learning, while the latest interaction supports immediate error correction. Across 30 RoboDojo tasks, using zero shot for Open and one demonstration elsewhere, RoboICL improves on official zero-shot \gptastra{} by 20--27 progress-score points in every category. It leads the leaderboard baselines on Memory and Open, achieves comparable performance to the strongest Precision baseline, and remains competitive on Long-Horizon. Its 30-task Overall score is 50.64, versus 33.68 for the strongest baseline. On a separate ten-task subset, RoboICL scores 60.60, within 2.00 points of the π0.5 + \gptastra{} hybrid approach. On three real-robot tasks, mean progress rises from 14.45 at zero shot to 63.33 at one shot and 78.89 at three shots. On two development tasks, optional Jev-gated action reuse reduces \gptastra{} calls by 33--48%. Code is available at \href{https://github.com/Mosi-AI/RoboICL}{https://github.com/Mosi-AI/RoboICL}.
Figures & tables
Figure 1: RoboDojo category-level comparison. Baseline scores are reproduced from the RoboDojo leaderboard, with GPT-6 Astra denoting the official zero-shot results ( Zhang et al., 2026 ) . As no training demonstrations are available for Open, RoboICL uses interaction memory alone (zero-shot) in this category; all other categories use one demonstration (one-shot). Integer labels on the official GPT-6 Astra bars show RoboICL’s absolute gains in progress-score points. Among the strongest methods currently listed on the leaderboard, RoboICL leads on Memory and Open, is comparable to the strongest method on Precision, and remains competitive on Long-Horizon, without robot-specific post-training. Table 2 reports the per-task evaluation counts.
Figure 2: A shared interaction grammar for demonstration context and interaction memory. Both sources use observation–action–receipt–observation records. Synthetic demonstration receipts describe recorded intervals, action counts, target tracking, and result states. Execution receipts report realized actions and controller feedback. The general action dimension is da , with da=14 for the illustrated RoboDojo interface. The lower panel shows a bottle-task request at step 687: four fixed anchors and the latest interaction, with explicit gaps between retained intervals.
Figure 3: Same frozen GPT-6 Astra under three zero-shot controller designs. All three systems receive the task, robot state, head and bilateral-wrist images, and online feedback, without expert demonstrations or a learned VLA. The official RoboDojo controller ( Zhang et al., 2026 ) retains recent images and older text, predicts absolute grasp-point poses and gripper commands, and motion-plans a joint trajectory lasting up to 10 s. GPT-as-Policy Direct ( Su et al., 2026 ) maintains a persistent tool-accessible conversation, predicts absolute dual-arm link-6 targets, and tracks them with local inverse kinematics for 1–5 steps at 25 Hz. RoboICL stitches the three views into one triptych, retains B=25 temporally distributed interactions in bounded anchored memory, and predicts an adaptive H×14 sequence of per-step arm increments and gripper targets for validated execution at 25 Hz. The comparison summarizes how context retention, image packaging, action parameterization, and execution horizon distinguish the three zero-shot controllers.
Figure 4: Scaling interaction memory and allocating demonstration context. Each point is a task’s mean progress score over five layouts. The horizontal axis gives J/B : retained demonstration blocks and interaction-memory blocks. Left (blue): without demonstrations ( J=0 ), increasing B retains more of the robot’s own experience. The plotted scores rise for Build Tower and Cover Blocks. Right (beige): one demonstration is sampled at different densities while J+B=24 . The intermediate allocations (12,12) and (16,8) achieve the highest four-task means; allocating 22 blocks to the demonstration and only two to interaction memory lowers scores on three tasks. Thus, denser demonstration coverage alone does not guarantee better control. Dashed links connect the zero- and one-shot configurations, where both demonstration content and memory capacity change.
VLA / WAM
Hybrid
0-shot
1-shot
Task
Galaxea G0.5
OpenWAM α
π0.5
GPT-6 Astra +π0.5
GPT-6 Astra Direct
GPT-6 Astra RoboICL
GPT-6 Astra RoboICL
Organize the table
46.33
62.50
23.33
60.00
30.00
45.00
50.00
Classify by language
1.07
1.33
0.60
38.00
60.00
44.00
44.00†
Imitate sorting sequence
1.67
2.90
1.60
53.00
0.00
90.00
90.00
Arrange largest number
4.11
4.36
2.29
50.00
57.00
71.00
58.00
Pack objects into a box
17.12
20.83
18.36
50.00
50.00
16.00
36.00
Table 1: Ten-task RoboDojo comparison. Mean progress score (0–100); Overall averages the ten task means. Direct denotes GPT-as-Policy’s Codex-based controller ( Su et al., 2026 ) ; Figure 3 compares it with the official baseline and RoboICL. RoboICL uses task-dependent horizons, with B=25 at zero shot and J=B=12 at one shot. External columns retain their published protocols ( Zhang et al., 2026 ; Su et al., 2026 ) . RoboICL targets five layouts per task and averages valid scores; † marks reused zero-shot results.
VLA / WAM
GPT-6 Astra
Galaxea G0.5
DM0.5
Liber-0 Preview
Simate-beta
RoboProbe
Task
( Liu et al., 2026 )
( Dexmal Team, 2026 )
( Zhang et al., 2026 )
( Zhang et al., 2026 )
( Zhang et al., 2026 )
RoboICL
Open
Align Blocks
0.00
0.00
0.00
0.00
50.00
90.00 50
Classify Objects by Language
1.07
0.47
3.87
6.00
46.00
70.80 50
General Pickup
12.67
14.00
34.00
49.33
84.00
90.00 50
Table 2: Four-category RoboDojo comparison. Mean progress score (0–100). Published VLA/WAM columns use three seeds and 50 episodes per task; official zero-shot GPT-6 Astra (RoboProbe) uses one seed and 50 episodes per task ( Chen et al., 2026b ; Zhang et al., 2026 ) . All RoboICL evaluations use seed 0. RoboICL uses zero shot on Open and one shot elsewhere; blue and orange cells mark these settings. Upper-right markers give the number of evaluated layouts; 0 marks a waived task assigned zero. Bold marks the highest score in each row.
Figure 5: The critical wrist-guided request in Deposit Coin. The figure shows the right-wrist observation from the request at step 205. GPT-6 Astra identifies the coin–slot offset and emits five 25 Hz commands with millimeter-scale translation and a small yaw correction. Subsequent requests complete alignment, seating, and release; the selected episode receives score 100.
Figure 6: Real-robot manipulation. Executions of Peg in Hole, Folding Towel, and Building Bridge using a Franka Research 3 with wrist and external RGB cameras.
Figure 7: Shot scaling in simulation and on a real robot. Left: five-layout simulator means at H=15 and J=B=5 . Right: five-trial real-robot means under staged progress scoring. Three demonstrations improve every shown task over zero shot.
Test condition
Progress score
Seen 1
100.00
Unseen 1
90.00
Unseen 2 (larger towel)
40.00
Mean across unseen conditions
65.00
Table 3: Folding Towel OOD progress scores. Three-shot mean progress score (0–100) over five trials per towel. Unseen towels do not appear in the demonstrations.
Task
Shot
Progress score
Tokens (M)
KV-cache hit (%)
Calls / chunks
API (min)
Wall (min)
Build Tower
0
16
5.040
91.66
77.4 / 72.4
48.8
58.7
1
100
3.265
92.28
40.2 / 39.6
31.5
37.9
Classify Objects
0
100
2.632
89.27
55.6 / 53.4
29.9
37.4
1
71
5.573
94.03
62.6 / 62.2
43.3
52.8
Put Bottles into Dustbin
0
70
3.246
90.93
64.4 / 61.2
31.4
37.9
1
73
5.091
93.78
61.8 / 61.0
38.8
46.3
Table 4: Inference use and latency on three five-layout task pairs. Each row averages five scored episodes. Tokens are provider-reported input plus output; KV-cache hit rate is the token-weighted cached share of input. Calls count completed API requests, chunks count executed tool calls, API time sums request latency, and wall time includes simulator and transport overhead.
Figure 8: KV-cache hit rate across model calls. The same three-task, 30-episode panel as Table 4 . Thin lines show each layout; thick lines give the pooled token-level rate 100∑eCe,t/∑eIe,t at each call index. Later points average the episodes that reach that call. Curves use completed responses with provider usage and include all input modalities.
Task
Pure GPT-6 Astra
GPT-6 Astra +Jev
Paired wall time (s)
Align Blocks
3/5
2/5
1128/865 ( n=1 )
General Pickup
4/5
5/5
3642/2327 ( n=4 )
Table 5: Complete-task success and paired completion time with Jev. Counts are over five layouts. The last column sums wall-clock seconds only over layouts successful in both conditions, in the order pure GPT-6 Astra / GPT-6 Astra +Jev; n gives the number of jointly successful layouts.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 9: Recovery after an interrupted action prefix. On Build Tower L0, the one-shot policy ( H=15 , J=B=12 ) proposes 15 commands at step 165. The controller executes five before a continuity rejection at step 170 and discards the ten-command suffix. The next execution note requests smaller rotations; construction continues and reaches score 100 at step 600. The same layout and seed under zero-shot RoboICL ( B=25 ) receive 10 at the 1050-step limit. All panels use native head-camera frames.
Figure 10: Complementary examples of demonstration-conditioned manipulation. Top: terminal views from Fasten Screws layout 2, where zero-shot RoboICL scores 20 and three-shot RoboICL scores 100. Middle: Fill Pen Holder demonstrations with opposite hand assignments. Bottom: the selected rollout before and after transferring the filled holder from the right hand to the left. Each observation concatenates the left-wrist, head, and right-wrist views in that order.
Figure 11: Selected three-shot Folding Towel executions with two unseen towels: Unseen 1 (left) and the larger Unseen 2 (right). Table 3 reports progress scores over five trials per condition.
Task
Layout
Pure GPT-6 Astra
GPT-6 Astra +Jev
Align Blocks
0
0/200/4489
0/200/1135
1
100/151/1128
100/148/865
2
100/167/1490
0/200/1120
3
100/138/1054
0/200/1199
4
0/200/2358
100/141/827
General Pickup
0
100/64/1335
100/70/440
Appendix
Table 6: Every Jev-panel layout. Each cell is progress score (0–100) / executed robot steps / wall-clock seconds (rounded). Scores in this panel are 0 or 100; Failed 200-step episodes are shown but excluded from successful-pair completion-time comparisons.
Task
Condition
Success
GPT-6 Astra calls
GPT-6 Astra API s
GPT-6 Astra tokens
Jev votes / reused steps
Align Blocks
Pure GPT-6 Astra
3/5
181
9407
9,534,190
0 / 0
GPT-6 Astra +Jev
2/5
121
4101
5,083,650
113 / 315
General Pickup
Pure GPT-6 Astra
4/5
116
4192
4,952,887
0 / 0
GPT-6 Astra +Jev
5/5
60
2095
1,391,331
53 / 140
Appendix
Table 7: Total model use across each five-layout condition. GPT-6 Astra calls and API seconds use completed requests; tokens are provider-reported response usage. Jev reuse counts executed steps approved by the gate. Totals include all five episodes.
Enabling robots to adapt to unfamiliar environments as readily as humans remains a moonshot goal of embodied AI. No finite collection of demonstrations can cover every task and situation a robot will encounter, making the ability to learn from context at deployment essential for generalization. Such in-context learning (ICL), however, remains largely beyond the reach of existing robotic policies. The broad agentic capabilities of commercial vision-language models (VLMs), such as GPT-6 Astra, raise a compelling question: can these models learn from demonstrations, examples, and interaction feedback, then translate that information into executable and verifiable robot behavior from a new initial state without gradient updates or persistent changes to task-specific parameters? We introduce GPT-Policy, a general-agent framework for in-context robot learning. GPT-Policy integrates a context compiler that preserves task-relevant visual transitions, a VLM that proposes robot-tool actions, and a constrained controller that verifies and executes each action and reports its outcome. We evaluate its reliability and limitations through task success and efficiency metrics, matched comparisons across models, and controlled context ablations. In real-robot trials, human video demonstrations improve task completion even without robot action labels, while aligned action references yield further gains on contact-sensitive tasks. These findings position GPT-Policy as a step toward robot adaptation through in-context learning, providing an empirical foundation for translating the general-purpose capabilities of VLMs into physical behavior and clarifying the challenges that must be overcome for reliable deployment.
We study robotic in-context learning (ICL), an emerging paradigm that enables robots to infer and execute tasks from visual demonstrations. Despite its growing promise, the problem itself remains under-defined: a visual demonstration simultaneously conveys action trajectories, object semantics, manipulation affordances, spatial relations, and task goals, making it unclear what information the robot is actually expected to follow. In this work, we first provide a clear problem definition of robot ICL that explicitly defines its learning target and resolves this fundamental prompt ambiguity. Building on this definition, we develop a minimalist and reproducible ICL framework (SimpleICL) with a visual prompt encoder and a low-cost data collection protocol. Without massive pre-training or specialized data infrastructure, our framework achieves strong performance in both simulation and real-world environments. Extensive experiments further reveal several key properties of robot ICL, including action, semantic, composition, and affordance discrimination. We will fully open-source our data and training pipeline to facilitate systematic and reproducible research on robot ICL. The project page can be found at https://simpleicl.github.io/simpleicl.
Embodied AI systems are often organized into System 1 and System 2. System 1 is typically a pretrained policy that generates actions at high frequency, whereas System 2 is often instantiated as a vision-enabled language model for high-level planning. We ask whether a large language model (LLM) can act as the policy for robot manipulation without task-specific finetuning. We call this setting LLM as policy. We evaluate three LLMs on all 42 RoboDojo tasks and compare their scores with 40 public policies. Astra and GPT-5.5 use the official 50-episode-per-task protocol; DeepSeek-Flash uses 10 episodes per task. GPT-6 Astra achieves 22.48% average success rate and 28.97 Score over 2,100 trials, ranking above every public entry. Yet GPT-5.5 and DeepSeek-Flash reach only 0.88% and 1.92% average success rate with the same post-processing. We find that Astra exhibits a sharply polarized capability profile. It generalizes well to tasks that require semantic understanding but not high-precision control. In contrast, it performs poorly on tasks that require precision, dynamic control, or complex bimanual coordination. In-context experiments show no aggregate benefit from one-shot demonstrations, while selected interaction traces show within-episode corrections under perturbations. Overall, the evaluated LLMs vary substantially in manipulation performance. Astra stands out and provides initial evidence for the potential of a general-purpose manipulation model, although reliable precision and dynamic control remain limitations in the evaluated setting.