Vision-language-action and world-action models have demonstrated impressive capabilities in robotics, yet generalization to unseen tasks remains challenging. More recently, general-purpose multimodal agents have shown great potential for zero-shot robotic task solving. However, they often incur high execution costs by reasoning and exploring the physical world from scratch. To reduce these costs, we introduce RoboSkill, a framework that connects skill acquisition and reuse through an Explore, Execute, Evolve loop. Within this loop, the agent explores to gather task-relevant information, executes tasks while adapting to feedback, and evolves its skill library based on execution records. It then reuses these skills to guide exploration and execution in the next cycle, closing the loop. To improve loop efficiency, we complement vision with tactile feedback to reduce uncertainty during physical interaction. We further augment textual guidance with reusable code to reduce reasoning overhead during skill reuse. On LIBERO-10, RoboSkill improves first-episode success rates by 12.5--25.0 percentage points and reduces average runtime by 7.6--72.4% across four agents. On real robots, it improves success rates by 8.3 percentage points and reduces average runtime for successful trials by at least 14.4%.
Figures & tables
Figure 1 : RoboSkill accelerates agent execution through skill reuse. In this tilted-nail example, a VLA/WAM follows the vertical hammering strategy learned from its training data. An embodied agent reasons about both grasping and striking, while RoboSkill accelerates execution by reusing the grasping skill and reasoning only about the strike direction.
Figure 2 : The RoboSkill framework. The agent explores to acquire missing information, executes with visual and tactile feedback, and produce skills to guide the next cycle. The example illustrates a nested exploration loop that verifies table height through contact before resuming grasping.
Model / setting
Final SR ↑
1st SR ↑
SR@30m ↑
Avg. Ep. ↓
Avg. Time (min) ↓
GPT-6 Astra
Baseline
100.0%
72.5%
87.5%
1.5
18.5
RoboSkill
100.0%
97.5%
95.0%
1.0
17.1
GPT-5.6 Sol
Baseline
82.5%
27.5%
32.5%
9.2
94.5
RoboSkill
97.5%
52.5%
72.5%
2.9
35.9
Table 1 : Main results on Libero-10.
Easy Tasks
Hard Tasks
Setting
1st SR ↑
Avg. Time (min) ↓
1st SR ↑
Avg. Time (min) ↓
Baseline
91.7%
20.3
80.0%
27.0
RoboSkill
100.0%
16.6
88.3%
23.1
Table 2 : Main results on real-world using GPT-6 Astra.
GPT-6 Astra
GPT-5.6 Sol
Fable 5.1
Opus 5
Setting
1st SR ↑
Avg. Time (min) ↓
1st SR ↑
Avg. Time (min) ↓
1st SR ↑
Avg. Time (min) ↓
1st SR ↑
Avg. Time (min) ↓
RoboSkill−
92.5%
16.5
67.5%
39.9
90.0%
13.0
67.5%
30.4
RoboSkill
97.5%
17.1
52.5%
35.9
97.5%
12.5
80.0%
22.3
Table 3 : Using one skill versus multiple skills on LIBERO-10. RoboSkill− uses only the skill learned on the target task. RoboSkill can additionally select skills learned on other tasks, using up to three skills in total.
GPT-5.6 Sol
Opus 5
Setting
1st SR ↑
Avg. Time (min) ↓
1st SR ↑
Avg. Time (min) ↓
Baseline
30.0%
86.0
60.0%
77.0
Other-Task Skill
44.0%
66.3
56.0%
43.5
Table 4 : Can skills transfer to a different task? Other-Task Skill provides one skill learned on a paired source task, with no skill from the target task. Baseline uses no skills.
Setting
Final SR ↑
1st SR ↑
SR@30m ↑
Avg. Ep. ↓
Avg. Time (min) ↓
Baseline
48.0%
12.0%
4.0%
3.50
180.8
RoboSkill− from Astra
80.0%
58.0%
20.0%
1.33
101.2
RoboSkill− from Sol
84.0%
42.0%
32.0%
3.56
86.5
RoboSkill− from Fable
92.0%
66.0%
60.0%
2.50
54.9
RoboSkill− from Opus
94.0%
58.0%
48.0%
1.84
52.4
Table 5 : Cross-agent skill transfer to Kimi K3 on LIBERO-10. We evaluate 5 seeds per task, totaling 50 cells per setting.
Setting
1st SR ↑
Avg. Time (min) ↓
Baseline
13.3%
34.8
Skills from Astra
70.0%
30.5
Table 6 : Cross-agent transfer on real-world easy tasks. Skills transfer from GPT-6 Astra to GPT-5.6 Sol.
Model / setting
Final SR ↑
1st SR ↑
SR@30m ↑
Avg. Ep. ↓
Avg. Time (min) ↓
GPT-5.6 Sol
Baseline
86.0%
30.0%
38.0%
9.0
80.3
RoboSkill−
96.0%
58.0%
66.0%
3.7
47.1
RoboSkill− (1st update)
94.0%
58.0%
68.0%
3.7
49.3
RoboSkill− (2nd update)
98.0%
70.0%
82.0%
3.1
30.2
Opus 5
Table 7 : Iterative skill refinement on LIBERO-10. Baseline uses no saved skills. RoboSkill− uses an initial skill constructed from successful exploration on the target task. Each update revises that skill through further successful exploration by the same model.
Setting
1st SR ↑
Avg. Time (min) ↓
Baseline
80.0%
27.0
RoboSkill−
88.3%
23.1
RoboSkill− (1st update)
93.3%
26.2
Table 8 : Skill evolution on real-world hard tasks using GPT-6 Astra.
GPT-6 Astra
GPT-5.6 Sol
Fable 5.1
Opus 5
Setting
1st SR ↑
Avg. Time (min) ↓
1st SR ↑
Avg. Time (min) ↓
1st SR ↑
Avg. Time (min) ↓
1st SR ↑
Avg. Time (min) ↓
RoboSkill− (w/o Code)
90.0%
21.1
15.0%
67.1
90.0%
16.1
70.0%
32.1
RoboSkill−
92.5%
16.5
67.5%
39.9
90.0%
13.0
67.5%
30.4
Table 9 : Effect of executable code on LIBERO-10. The w/o Code variant removes code from the skill provided to RoboSkill− while retaining its textual guidance.
LIBERO-10
Real Robot
Setting
Final SR ↑
1st SR ↑
SR@30m ↑
Avg. Ep. ↓
Avg. Time (min) ↓
1st SR ↑
Avg. Time (min) ↓
With tactile
100%
78%
94%
1.42
15.93
90%
20.00
w.o. tactile
100%
82%
74%
1.20
25.35
80%
33.63
Table 10 : Importance of tactile input during exploration using GPT-6 Astra. In simulation, we evaluate 5 seeds per task, totaling 50 cells per setting.
Figure 3 : Real-robot hardware setup. Only the right arm is used in the experiments, together with its wrist-mounted camera and tactile-equipped gripper. The left arm is shown but is not used. The upper annotation indicates the support for the overhead third-person camera, which is outside the photographed field of view. The lower panel illustrates the tactile sensor, its sensing surface, and an example tactile image.
Model
Final SR ↑
1st SR ↑
SR@30m ↑
Avg. Ep. ↓
Avg. Time (min) ↓
4 seeds per task: 40 cells
GPT-6 Astra
100.0%
72.5%
87.5%
1.50
18.48
GPT-5.6 Sol
82.5%
27.5%
32.5%
9.20
94.50
Fable 5.1
97.5%
85.0%
77.5%
2.53
34.80
Opus 5
100.0%
60.0%
17.5%
1.93
80.90
5 seeds per task: 50 cells
Table 11 : Baseline results under the two seed protocols. The 4-seed protocol excludes the skill-construction seed and use seed 1-4 for each task; the 5-seed protocol includes all seeds 0–4.
Figure 4 : Example initial states of the ten LIBERO-10 tasks. Task IDs are used consistently in the cross-task pairings and skill selection tables. Each image illustrates one initial state.
ID
Task
Objective
Easy tasks
(a)
Press button
Approach and press the button to activate it.
(b)
Lift paper bag
Grasp a handle and lift the bag from the table.
(c)
Open drawer
Grasp the drawer and pull it outward.
(d)
Remove test tube
Grasp a test tube and extract it from the rack.
(e)
Sort socks
Arrange white socks on the left and black socks on the right.
Figure 5 : Execution stages of the 12 reported real-robot tasks. Panels (a)–(f) show easy tasks and panels (g)–(l) show hard tasks. Task objectives are listed in table 12 .
Figure 5 : Execution stages of the 12 reported real-robot tasks. Panels (a)–(f) show easy tasks and panels (g)–(l) show hard tasks. Task objectives are listed in table 12 .
ID
Task
Objective
Easy tasks
(a)
Press button
Approach and press the button to activate it.
(b)
Lift paper bag
Grasp a handle and lift the bag from the table.
(c)
Open drawer
Grasp the drawer and pull it outward.
(d)
Remove test tube
Grasp a test tube and extract it from the rack.
(e)
Sort socks
Arrange white socks on the left and black socks on the right.
Table 12 : Real-robot task objectives. Letters correspond to the panels in figure 5 .
Task
Baseline
RoboSkill
Final SR ↑
Avg. Time (min) ↓
Final SR ↑
Avg. Time (min) ↓
Easy tasks
Press button
10/10
17.68
10/10
9.10
Lift paper bag
10/10
19.58
10/10
15.70
Open drawer
8/10
22.53
10/10
15.40
Remove test tube
9/10
20.00
10/10
16.00
Table 13 : Task-wise Astra performance before and after exploration. Baseline timing uses the separate successful-sample records. Setting: 12-task subset (six easy and six hard), ten trials per task for success rates, 1 h per trial.
Task
Baseline
RoboSkill from Astra
Final SR ↑
Avg. Time (min) ↓
Final SR ↑
Avg. Time (min) ↓
Press button
0/5
–
5/5
25.40
Lift paper bag
2/5
27.00
5/5
19.60
Open drawer
1/5
27.00
3/5
34.33
Remove test tube
1/5
58.00
3/5
35.67
Sort white socks left and black socks right
0/5
–
2/5
29.00
Table 14 : Task-wise GPT-5.6 Sol performance before and after receiving skills constructed by Astra. Avg. Time uses successful trials only. Setting: six easy tasks, five trials per task, tactile input, and a 1 h limit per trial.
Target task
Rule
GPT-5.6 Sol
Opus 5
Fable 5.1
GPT-6 Astra
T0: Two cans to basket
[T0,T7,T1]
[T0,T7,T1]
[T0,T7,T1]
[T0,T7,T1]
[T0,T7]
T1: Cheese and butter to basket
[T1,T7,T0]
[T1,T7]
[T1,T7,T0]
[T1,T7,T0]
[T1,T7]
T2: Turn on stove; place moka pot
[T2,T8]
[T2,T8]
[T2,T8]
[T2,T8]
[T2,T8]
T3: Black bowl to drawer; close
[T3,T9]
[T3,T9,T5]
[T3,T9]
[T3,T9]
[T3]
T4: Two cups to left/right plates
[T4,T6,T9]
[T4,T6]
[T4,T6]
[T4,T6]
[T4,T6]
T5: Book to rear caddy slot
[T5]
[T5,T3]
[T5,T7]
[T5,T3]
[T5]
Table 15 : Agent-selected versus rule-based Top-K skill sets. Each cell lists the source-task skills selected for the target task. Bold agent-selected cells differ from the fixed rule in membership or order.
Model / setting
Final SR ↑
1st SR ↑
SR@30m ↑
Avg. Ep. ↓
Avg. Time (min) ↓
GPT-6 Astra
Rule-based Top-K
100%
87.5%
85%
1.175
19.2
Agent-selected Top-K
100%
97.5%
95%
1.025
17.1
Δ agent vs. rule
+0 pp
+10 pp
+10 pp
-12.8%
-11.2%
GPT-5.6 Sol
Rule-based Top-K
97.5%
57.5%
70%
3.425
35.4
Table 16 : Rule-based versus agent-selected Top-K performance. Δ is the change from rule-based to agent-selected Top-K (percentage points for rates and relative percent otherwise). Bold marks the better non-tied value within each model and metric.
Traditional robot programming is challenging: it requires orchestrating multimodal perception, managing physical contact dynamics, and handling diverse configurations and execution failures. We introduce ASPIRE (Agentic Skill Programming through Iterative Robot Exploration), a continual learning system that autonomously writes and refines robot control programs in a code-as-policy paradigm while compounding experience into a reusable skill library. ASPIRE discovers skills that persist across tasks, simulation and real-world settings, and embodiments. It operates in an open-ended loop with three components: (1) a closed-loop robot execution engine that exposes fine-grained multimodal traces, enabling autonomous failure diagnosis, repair synthesis, and validation; (2) a continually expanding skill library that distills validated fixes into reusable, transferable knowledge; and (3) evolutionary search that generates diverse task sequences and control programs to explore beyond single-trajectory refinement. ASPIRE surpasses prior methods by up to 77% on LIBERO-Pro manipulation under perturbation, 72% on Robosuite bimanual handover, and 32% on BEHAVIOR-1K long-horizon household tasks. Its accumulated library also enables zero-shot generalization to unseen long-horizon tasks: on LIBERO-Pro Long, ASPIRE achieves 31% success versus 4% for prior methods despite their use of test-time reasoning and retries. Finally, simulation-discovered skills provide initial evidence of sim-to-real transfer, substantially reducing real-robot programming effort across different embodiments and robot APIs.
Vision-language-action (VLA) models map visual observations and language instructions directly to robot actions, but long-horizon tasks require more than action prediction. An agent must coordinate perception, planning, execution, progress verification, and recovery as the physical state evolves. An action prediction or a model-generated skill decision does not, by itself, guarantee that the proposed operation is valid in the current state or that its outcome will be verified. We propose EmbodiedSkills, a unified framework that treats each skill decision as an execution proposal: the runtime checks its prerequisites before execution and verifies the outcome afterward. A shared executable-skill interface connects high-level skill selection, bounded low-level VLA execution, and post-action verification within a single agent loop. Because this interface remains fixed, low-level VLA policies can be replaced or adapted without changing the agent loop. The interface also records planning, execution, verification, and recovery events as structured trajectories, which provide supervision for individual components and can support optional online adaptation when interactive feedback is available. We instantiate EmbodiedSkills with Qwen3-VL and OpenPI/pi0.5 on RoboTwin 2.0 and LIBERO. Task-adapted low-level VLA policies achieve an average success rate of 86.20% across 50 RoboTwin 2.0 tasks and 97.40% across the four LIBERO suites. These results establish the execution performance of the task-adapted low-level VLA policies used in EmbodiedSkills. On four memory-dependent RMBench tasks, the same task-adapted execution approach achieves 12.5% average success. The framework provides a trainable and inspectable agent layer for turning these policies into closed-loop embodied systems.
Current agentic robot systems can write executable Code-as-Policy programs, observe feedback, and revise behavior across multiple attempts, but they remain largely task-driven: reusable skills are acquired only after explicit instructions. We study Playful Agentic Robot Learning, where an embodied coding agent uses self-directed play as a continual skill-learning stage before downstream tasks arrive. We introduce RATs, Robotics Agent Teams designed for play-time skill acquisition. During play, RATs proposes novel yet learnable exploratory tasks, plans and executes robot-code policies, verifies intermediate progress, diagnoses failures, retries with dense, step-level feedback, and distills successful executions into a persistent code skill library. At test time, the agent reuses relevant skills from this frozen library to help solve new tasks. Experiments in LIBERO-PRO and MolmoSpaces show that play-learned skills improve held-out downstream tasks over no-play and random-play baselines, with 20.6 and 17.0 percentage-point gains over CaP-Agent0 on LIBERO-PRO and MolmoSpaces, respectively. Moreover, the learned skills can be plugged into other inference-time Code-as-Policy agents by simply retrieving them into the context, improving RoboSuite and real-world transfer by 8.9 and 8.8 points, respectively, without finetuning the underlying model.