We study robotic in-context learning (ICL), an emerging paradigm that enables robots to infer and execute tasks from visual demonstrations. Despite its growing promise, the problem itself remains under-defined: a visual demonstration simultaneously conveys action trajectories, object semantics, manipulation affordances, spatial relations, and task goals, making it unclear what information the robot is actually expected to follow. In this work, we first provide a clear problem definition of robot ICL that explicitly defines its learning target and resolves this fundamental prompt ambiguity. Building on this definition, we develop a minimalist and reproducible ICL framework (SimpleICL) with a visual prompt encoder and a low-cost data collection protocol. Without massive pre-training or specialized data infrastructure, our framework achieves strong performance in both simulation and real-world environments. Extensive experiments further reveal several key properties of robot ICL, including action, semantic, composition, and affordance discrimination. We will fully open-source our data and training pipeline to facilitate systematic and reproducible research on robot ICL. The project page can be found at https://simpleicl.github.io/simpleicl.
Figures & tables
Figure 1: Overview of SimpleICL. We identify the prompt ambiguity problem in robot ICL and clarify the intrinsic semantics that the model should learn. We further improve task understanding and generalization through efficient data engineering. Our method achieves the best performance across eight evaluation tasks.
Figure 2: Four data types constructed for ICL based on our problem definition.
Figure 3: Schematic illustration of the real-world ICL instance construction.
Model
Trained on
Average
Fast-WAM
Native
37.6
SimpleICL (Ours)
Native
40.3
ICL-oriented
72.1
w/o. video cond.
ICL-oriented
61.6
w/o. future pred.
ICL-oriented
20.2
Table 1: Simulation evaluation results on the visual ambiguity test of RoboTwin tasks.
Table 3: Real-world evaluation on eight unseen tasks. Success rates (%) are reported for each method. Task abbreviations: Wipe = wiping a plate; Tea = serving tea; Drawer = pulling a drawer; Soap = pressing a soap dispenser; Fruit = placing fruit; Chopsticks = organizing chopsticks and bowls. Shelf = organizing items on a shelf. Please refer to appendices for detailed task design.
Figure 5: Qualitative comparison between our method and existing baselines. At the Medium level (top), Fast-WAM is distracted by a nearby scale, while our method correctly locates the drawer. At the Hard level (bottom), Fast-WAM mistakenly places chopsticks onto the round bowl, whereas ours reaches the correct destination.
Setting
Composition
Affordance
Snack
Cup
Bag
Cup
Measure
Tape
w/. DC†
0.97
1.00
0.90
0.87
0.77
0.83
w/o. DC
0.93
0.97
0.93
0.60
0.40
0.47
Table 4: Analysis of composition and affordance under discriminative sampling (ours) versus without discriminative sampling. The metric is intent-following rate (%). †: DC means discriminative data collection.
Method
Orig.
Spatial
Human
Scene Variations
Avg.
App.
Sty.
Light
Bg.
Clut.
Ours
0.81
0.79
0.77
0.80
0.78
0.79
0.75
0.78
w/o. CP†
0.80
0.51
0.78
0.79
0.73
0.71
0.66
0.70
Table 5: Robustness analysis. Success rates (%) are reported under multiple disturbance categories. Spatial denotes that the target object position during robot execution is shifted relative to that in the human demonstration. Human denotes human action details, where App. changes the appearance of the demonstrator’s hand (e.g., wearing gloves), while Sty. changes the human subject and motion style. For Scene Variations , Light changes the illumination condition, Bg. introduces background color distractions, and Clut. adds irrelevant clutter objects to the scene. †: CP means cross-group pairing strategy for data augmentation.
Figure 7: The timeline of attention visualization heat map for the human conditioning videos on two real tasks of Place Cup on Tray and Move Plate Between Holders .
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Novel (OOD) tasks for simulation evaluation.
In-context imitation learning (ICIL) enables robots to learn new tasks from a small number of demonstrations by conditioning a pre-trained policy on task-specific examples, without retraining at test time. Despite this promise, training generalizable and scalable in-context imitation policies remains an open challenge. We present SynthICL, a scalable framework that trains ICIL policies entirely from RGB-only synthetic data. Specifically, we build a data generation pipeline to produce high-fidelity ICIL data and train a flow-matching transformer policy on the resulting dataset. SynthICL avoids the need for depth sensing, precise camera calibration, and real-world training data in prior approaches, offering a simpler and more scalable alternative. We further incorporate subgoal prediction by training the model to predict the next subgoal images, enabling more precise and visually grounded control. Evaluated on 16 unseen real-world manipulation tasks, SynthICL achieves an average success rate of 79% with only one demonstration provided at test time and outperforms prior methods. Project page: https://synth-icl.github.io
General-purpose robots must infer what a new task requires and translate that understanding into appropriate physical action. In-context learning (ICL) for robots supports this process by using demonstrations and interaction to direct existing competence with neural parameters held fixed during deployment. We organize this literature review around the interfaces connecting contextual evidence to execution, distinguishing four families: context-conditioned policies, geometric demonstration transfer, world-model-based control, and skill- and agent-based execution. Comparing these interfaces clarifies their transfer assumptions and the roles of training, correspondence, and memory in making context useful. Across manipulation and navigation, we examine how these mechanisms preserve taught requirements as objects, environments, and execution conditions change. This analysis links method design to evaluation practices that distinguish responsiveness to teaching, physical transfer, and benefits from retained experience. The resulting agenda connects compositional task acquisition and faithful transfer with physical recursive self-improvement, in which experience improves the ability to learn subsequent tasks.
General-purpose vision-language models offer a promising way to zero-shot robot control: \gptastra{} excels at open-ended and language- or image-conditioned manipulation but remains substantially weaker on high-precision and long-horizon tasks. We introduce \emph{RoboICL}, an in-context robot-control framework that narrows these gaps without robot-specific parameter updates or a learned VLA. RoboICL separates \emph{demonstration context}, which provides recorded examples when available, from \emph{interaction memory}, which accumulates the model's own actions and observed outcomes. Both use a shared observation--action--receipt--observation grammar. To preserve experience across task stages, RoboICL combines sampled demonstration blocks with bounded anchored memory. Fixed anchors keep earlier rollout interactions available for in-context learning, while the latest interaction supports immediate error correction. Across 30 RoboDojo tasks, using zero shot for Open and one demonstration elsewhere, RoboICL improves on official zero-shot \gptastra{} by 20--27 progress-score points in every category. It leads the leaderboard baselines on Memory and Open, achieves comparable performance to the strongest Precision baseline, and remains competitive on Long-Horizon. Its 30-task Overall score is 50.64, versus 33.68 for the strongest baseline. On a separate ten-task subset, RoboICL scores 60.60, within 2.00 points of the π0.5 + \gptastra{} hybrid approach. On three real-robot tasks, mean progress rises from 14.45 at zero shot to 63.33 at one shot and 78.89 at three shots. On two development tasks, optional Jev-gated action reuse reduces \gptastra{} calls by 33--48%. Code is available at \href{https://github.com/Mosi-AI/RoboICL}{https://github.com/Mosi-AI/RoboICL}.