We study robotic in-context learning (ICL), an emerging paradigm that enables robots to infer and execute tasks from visual demonstrations. Despite its growing promise, the problem itself remains under-defined: a visual demonstration simultaneously conveys action trajectories, object semantics, manipulation affordances, spatial relations, and task goals, making it unclear what information the robot is actually expected to follow. In this work, we first provide a clear problem definition of robot ICL that explicitly defines its learning target and resolves this fundamental prompt ambiguity. Building on this definition, we develop a minimalist and reproducible ICL framework (SimpleICL) with a visual prompt encoder and a low-cost data collection protocol. Without massive pre-training or specialized data infrastructure, our framework achieves strong performance in both simulation and real-world environments. Extensive experiments further reveal several key properties of robot ICL, including action, semantic, composition, and affordance discrimination. We will fully open-source our data and training pipeline to facilitate systematic and reproducible research on robot ICL. The project page can be found at https://simpleicl.github.io/simpleicl.
Figures & tables
Figure 1: Overview of SimpleICL. We identify the prompt ambiguity problem in robot ICL and clarify the intrinsic semantics that the model should learn. We further improve task understanding and generalization through efficient data engineering. Our method achieves the best performance across eight evaluation tasks.
Figure 2: Four data types constructed for ICL based on our problem definition.
Figure 3: Schematic illustration of the real-world ICL instance construction.
Model
Trained on
Average
Fast-WAM
Native
37.6
SimpleICL (Ours)
Native
40.3
ICL-oriented
72.1
w/o. video cond.
ICL-oriented
61.6
w/o. future pred.
ICL-oriented
20.2
Table 1: Simulation evaluation results on the visual ambiguity test of RoboTwin tasks.
Table 3: Real-world evaluation on eight unseen tasks. Success rates (%) are reported for each method. Task abbreviations: Wipe = wiping a plate; Tea = serving tea; Drawer = pulling a drawer; Soap = pressing a soap dispenser; Fruit = placing fruit; Chopsticks = organizing chopsticks and bowls. Shelf = organizing items on a shelf. Please refer to appendices for detailed task design.
Figure 5: Qualitative comparison between our method and existing baselines. At the Medium level (top), Fast-WAM is distracted by a nearby scale, while our method correctly locates the drawer. At the Hard level (bottom), Fast-WAM mistakenly places chopsticks onto the round bowl, whereas ours reaches the correct destination.
Setting
Composition
Affordance
Snack
Cup
Bag
Cup
Measure
Tape
w/. DC†
0.97
1.00
0.90
0.87
0.77
0.83
w/o. DC
0.93
0.97
0.93
0.60
0.40
0.47
Table 4: Analysis of composition and affordance under discriminative sampling (ours) versus without discriminative sampling. The metric is intent-following rate (%). †: DC means discriminative data collection.
Method
Orig.
Spatial
Human
Scene Variations
Avg.
App.
Sty.
Light
Bg.
Clut.
Ours
0.81
0.79
0.77
0.80
0.78
0.79
0.75
0.78
w/o. CP†
0.80
0.51
0.78
0.79
0.73
0.71
0.66
0.70
Table 5: Robustness analysis. Success rates (%) are reported under multiple disturbance categories. Spatial denotes that the target object position during robot execution is shifted relative to that in the human demonstration. Human denotes human action details, where App. changes the appearance of the demonstrator’s hand (e.g., wearing gloves), while Sty. changes the human subject and motion style. For Scene Variations , Light changes the illumination condition, Bg. introduces background color distractions, and Clut. adds irrelevant clutter objects to the scene. †: CP means cross-group pairing strategy for data augmentation.
Figure 7: The timeline of attention visualization heat map for the human conditioning videos on two real tasks of Place Cup on Tray and Move Plate Between Holders .
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Novel (OOD) tasks for simulation evaluation.