RoboICL: Embodied In-Context Learning with GPT-6 Astra
Organizations: Samsung Research, Beijing, China · Samsung Robotics eXperience · Shanghai Jiao Tong University
Abstract
General-purpose vision-language models offer a promising way to zero-shot robot control: \gptastra{} excels at open-ended and language- or image-conditioned manipulation but remains substantially weaker on high-precision and long-horizon tasks. We introduce \emph{RoboICL}, an in-context robot-control framework that narrows these gaps without robot-specific parameter updates or a learned VLA. RoboICL separates \emph{demonstration context}, which provides recorded examples when available, from \emph{interaction memory}, which accumulates the model's own actions and observed outcomes. Both use a shared observation--action--receipt--observation grammar. To preserve experience across task stages, RoboICL combines sampled demonstration blocks with bounded anchored memory. Fixed anchors keep earlier rollout interactions available for in-context learning, while the latest interaction supports immediate error correction. Across 30 RoboDojo tasks, using zero shot for Open and one demonstration elsewhere, RoboICL improves on official zero-shot \gptastra{} by 20--27 progress-score points in every category. It leads the leaderboard baselines on Memory and Open, achieves comparable performance to the strongest Precision baseline, and remains competitive on Long-Horizon. Its 30-task Overall score is 50.64, versus 33.68 for the strongest baseline. On a separate ten-task subset, RoboICL scores 60.60, within 2.00 points of the + \gptastra{} hybrid approach. On three real-robot tasks, mean progress rises from 14.45 at zero shot to 63.33 at one shot and 78.89 at three shots. On two development tasks, optional Jev-gated action reuse reduces \gptastra{} calls by 33--48%. Code is available at \href{https://github.com/Mosi-AI/RoboICL}{https://github.com/Mosi-AI/RoboICL}.
Figures & tables
| VLA / WAM | Hybrid | 0-shot | 1-shot | ||||
|---|---|---|---|---|---|---|---|
| Task | Galaxea G0.5 | OpenWAM | GPT-6 Astra | GPT-6 Astra Direct | GPT-6 Astra RoboICL | GPT-6 Astra RoboICL | |
| Organize the table | 46.33 | 62.50 | 23.33 | 60.00 | 30.00 | 45.00 | 50.00 |
| Classify by language | 1.07 | 1.33 | 0.60 | 38.00 | 60.00 | 44.00 | |
| Imitate sorting sequence | 1.67 | 2.90 | 1.60 | 53.00 | 0.00 | 90.00 | 90.00 |
| Arrange largest number | 4.11 | 4.36 | 2.29 | 50.00 | 57.00 | 71.00 | 58.00 |
| Pack objects into a box | 17.12 | 20.83 | 18.36 | 50.00 | 50.00 | 16.00 | 36.00 |
| VLA / WAM | GPT-6 Astra | |||||
| Galaxea G0.5 | DM0.5 | Liber-0 Preview | Simate-beta | RoboProbe | ||
| Task | ( Liu et al., 2026 ) | ( Dexmal Team, 2026 ) | ( Zhang et al., 2026 ) | ( Zhang et al., 2026 ) | ( Zhang et al., 2026 ) | RoboICL |
| Open | ||||||
| Align Blocks | 0.00 | 0.00 | 0.00 | 0.00 | 50.00 | 90.00 |
| Classify Objects by Language | 1.07 | 0.47 | 3.87 | 6.00 | 46.00 | 70.80 |
| General Pickup | 12.67 | 14.00 | 34.00 | 49.33 | 84.00 | 90.00 |
| Test condition | Progress score |
|---|---|
| Seen 1 | 100.00 |
| Unseen 1 | 90.00 |
| Unseen 2 (larger towel) | 40.00 |
| Mean across unseen conditions | 65.00 |
| Task | Shot | Progress score | Tokens (M) | KV-cache hit (%) | Calls / chunks | API (min) | Wall (min) |
|---|---|---|---|---|---|---|---|
| Build Tower | 0 | 16 | 5.040 | 91.66 | 77.4 / 72.4 | 48.8 | 58.7 |
| 1 | 100 | 3.265 | 92.28 | 40.2 / 39.6 | 31.5 | 37.9 | |
| Classify Objects | 0 | 100 | 2.632 | 89.27 | 55.6 / 53.4 | 29.9 | 37.4 |
| 1 | 71 | 5.573 | 94.03 | 62.6 / 62.2 | 43.3 | 52.8 | |
| Put Bottles into Dustbin | 0 | 70 | 3.246 | 90.93 | 64.4 / 61.2 | 31.4 | 37.9 |
| 1 | 73 | 5.091 | 93.78 | 61.8 / 61.0 | 38.8 | 46.3 |
| Task | Pure GPT-6 Astra | GPT-6 Astra +Jev | Paired wall time (s) |
|---|---|---|---|
| Align Blocks | 3/5 | 2/5 | 1128/865 ( ) |
| General Pickup | 4/5 | 5/5 | 3642/2327 ( ) |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Task | Layout | Pure GPT-6 Astra | GPT-6 Astra +Jev |
|---|---|---|---|
| Align Blocks | 0 | 0/200/4489 | 0/200/1135 |
| 1 | 100/151/1128 | 100/148/865 | |
| 2 | 100/167/1490 | 0/200/1120 | |
| 3 | 100/138/1054 | 0/200/1199 | |
| 4 | 0/200/2358 | 100/141/827 | |
| General Pickup | 0 | 100/64/1335 | 100/70/440 |
| Task | Condition | Success | GPT-6 Astra calls | GPT-6 Astra API s | GPT-6 Astra tokens | Jev votes / reused steps |
|---|---|---|---|---|---|---|
| Align Blocks | Pure GPT-6 Astra | 3/5 | 181 | 9407 | 9,534,190 | 0 / 0 |
| GPT-6 Astra +Jev | 2/5 | 121 | 4101 | 5,083,650 | 113 / 315 | |
| General Pickup | Pure GPT-6 Astra | 4/5 | 116 | 4192 | 4,952,887 | 0 / 0 |
| GPT-6 Astra +Jev | 5/5 | 60 | 2095 | 1,391,331 | 53 / 140 |