Vision-language-action and world-action models have demonstrated impressive capabilities in robotics, yet generalization to unseen tasks remains challenging. More recently, general-purpose multimodal agents have shown great potential for zero-shot robotic task solving. However, they often incur high execution costs by reasoning and exploring the physical world from scratch. To reduce these costs, we introduce RoboSkill, a framework that connects skill acquisition and reuse through an Explore, Execute, Evolve loop. Within this loop, the agent explores to gather task-relevant information, executes tasks while adapting to feedback, and evolves its skill library based on execution records. It then reuses these skills to guide exploration and execution in the next cycle, closing the loop. To improve loop efficiency, we complement vision with tactile feedback to reduce uncertainty during physical interaction. We further augment textual guidance with reusable code to reduce reasoning overhead during skill reuse. On LIBERO-10, RoboSkill improves first-episode success rates by 12.5--25.0 percentage points and reduces average runtime by 7.6--72.4% across four agents. On real robots, it improves success rates by 8.3 percentage points and reduces average runtime for successful trials by at least 14.4%.
Figures & tables
Figure 1 : RoboSkill accelerates agent execution through skill reuse. In this tilted-nail example, a VLA/WAM follows the vertical hammering strategy learned from its training data. An embodied agent reasons about both grasping and striking, while RoboSkill accelerates execution by reusing the grasping skill and reasoning only about the strike direction.
Figure 2 : The RoboSkill framework. The agent explores to acquire missing information, executes with visual and tactile feedback, and produce skills to guide the next cycle. The example illustrates a nested exploration loop that verifies table height through contact before resuming grasping.
Model / setting
Final SR ↑
1st SR ↑
SR@30m ↑
Avg. Ep. ↓
Avg. Time (min) ↓
GPT-6 Astra
Baseline
100.0%
72.5%
87.5%
1.5
18.5
RoboSkill
100.0%
97.5%
95.0%
1.0
17.1
GPT-5.6 Sol
Baseline
82.5%
27.5%
32.5%
9.2
94.5
RoboSkill
97.5%
52.5%
72.5%
2.9
35.9
Table 1 : Main results on Libero-10.
Easy Tasks
Hard Tasks
Setting
1st SR ↑
Avg. Time (min) ↓
1st SR ↑
Avg. Time (min) ↓
Baseline
91.7%
20.3
80.0%
27.0
RoboSkill
100.0%
16.6
88.3%
23.1
Table 2 : Main results on real-world using GPT-6 Astra.
GPT-6 Astra
GPT-5.6 Sol
Fable 5.1
Opus 5
Setting
1st SR ↑
Avg. Time (min) ↓
1st SR ↑
Avg. Time (min) ↓
1st SR ↑
Avg. Time (min) ↓
1st SR ↑
Avg. Time (min) ↓
RoboSkill−
92.5%
16.5
67.5%
39.9
90.0%
13.0
67.5%
30.4
RoboSkill
97.5%
17.1
52.5%
35.9
97.5%
12.5
80.0%
22.3
Table 3 : Using one skill versus multiple skills on LIBERO-10. RoboSkill− uses only the skill learned on the target task. RoboSkill can additionally select skills learned on other tasks, using up to three skills in total.
GPT-5.6 Sol
Opus 5
Setting
1st SR ↑
Avg. Time (min) ↓
1st SR ↑
Avg. Time (min) ↓
Baseline
30.0%
86.0
60.0%
77.0
Other-Task Skill
44.0%
66.3
56.0%
43.5
Table 4 : Can skills transfer to a different task? Other-Task Skill provides one skill learned on a paired source task, with no skill from the target task. Baseline uses no skills.
Setting
Final SR ↑
1st SR ↑
SR@30m ↑
Avg. Ep. ↓
Avg. Time (min) ↓
Baseline
48.0%
12.0%
4.0%
3.50
180.8
RoboSkill− from Astra
80.0%
58.0%
20.0%
1.33
101.2
RoboSkill− from Sol
84.0%
42.0%
32.0%
3.56
86.5
RoboSkill− from Fable
92.0%
66.0%
60.0%
2.50
54.9
RoboSkill− from Opus
94.0%
58.0%
48.0%
1.84
52.4
Table 5 : Cross-agent skill transfer to Kimi K3 on LIBERO-10. We evaluate 5 seeds per task, totaling 50 cells per setting.
Setting
1st SR ↑
Avg. Time (min) ↓
Baseline
13.3%
34.8
Skills from Astra
70.0%
30.5
Table 6 : Cross-agent transfer on real-world easy tasks. Skills transfer from GPT-6 Astra to GPT-5.6 Sol.
Model / setting
Final SR ↑
1st SR ↑
SR@30m ↑
Avg. Ep. ↓
Avg. Time (min) ↓
GPT-5.6 Sol
Baseline
86.0%
30.0%
38.0%
9.0
80.3
RoboSkill−
96.0%
58.0%
66.0%
3.7
47.1
RoboSkill− (1st update)
94.0%
58.0%
68.0%
3.7
49.3
RoboSkill− (2nd update)
98.0%
70.0%
82.0%
3.1
30.2
Opus 5
Table 7 : Iterative skill refinement on LIBERO-10. Baseline uses no saved skills. RoboSkill− uses an initial skill constructed from successful exploration on the target task. Each update revises that skill through further successful exploration by the same model.
Setting
1st SR ↑
Avg. Time (min) ↓
Baseline
80.0%
27.0
RoboSkill−
88.3%
23.1
RoboSkill− (1st update)
93.3%
26.2
Table 8 : Skill evolution on real-world hard tasks using GPT-6 Astra.
GPT-6 Astra
GPT-5.6 Sol
Fable 5.1
Opus 5
Setting
1st SR ↑
Avg. Time (min) ↓
1st SR ↑
Avg. Time (min) ↓
1st SR ↑
Avg. Time (min) ↓
1st SR ↑
Avg. Time (min) ↓
RoboSkill− (w/o Code)
90.0%
21.1
15.0%
67.1
90.0%
16.1
70.0%
32.1
RoboSkill−
92.5%
16.5
67.5%
39.9
90.0%
13.0
67.5%
30.4
Table 9 : Effect of executable code on LIBERO-10. The w/o Code variant removes code from the skill provided to RoboSkill− while retaining its textual guidance.
LIBERO-10
Real Robot
Setting
Final SR ↑
1st SR ↑
SR@30m ↑
Avg. Ep. ↓
Avg. Time (min) ↓
1st SR ↑
Avg. Time (min) ↓
With tactile
100%
78%
94%
1.42
15.93
90%
20.00
w.o. tactile
100%
82%
74%
1.20
25.35
80%
33.63
Table 10 : Importance of tactile input during exploration using GPT-6 Astra. In simulation, we evaluate 5 seeds per task, totaling 50 cells per setting.
Figure 3 : Real-robot hardware setup. Only the right arm is used in the experiments, together with its wrist-mounted camera and tactile-equipped gripper. The left arm is shown but is not used. The upper annotation indicates the support for the overhead third-person camera, which is outside the photographed field of view. The lower panel illustrates the tactile sensor, its sensing surface, and an example tactile image.
Model
Final SR ↑
1st SR ↑
SR@30m ↑
Avg. Ep. ↓
Avg. Time (min) ↓
4 seeds per task: 40 cells
GPT-6 Astra
100.0%
72.5%
87.5%
1.50
18.48
GPT-5.6 Sol
82.5%
27.5%
32.5%
9.20
94.50
Fable 5.1
97.5%
85.0%
77.5%
2.53
34.80
Opus 5
100.0%
60.0%
17.5%
1.93
80.90
5 seeds per task: 50 cells
Table 11 : Baseline results under the two seed protocols. The 4-seed protocol excludes the skill-construction seed and use seed 1-4 for each task; the 5-seed protocol includes all seeds 0–4.
Figure 4 : Example initial states of the ten LIBERO-10 tasks. Task IDs are used consistently in the cross-task pairings and skill selection tables. Each image illustrates one initial state.
ID
Task
Objective
Easy tasks
(a)
Press button
Approach and press the button to activate it.
(b)
Lift paper bag
Grasp a handle and lift the bag from the table.
(c)
Open drawer
Grasp the drawer and pull it outward.
(d)
Remove test tube
Grasp a test tube and extract it from the rack.
(e)
Sort socks
Arrange white socks on the left and black socks on the right.
Figure 5 : Execution stages of the 12 reported real-robot tasks. Panels (a)–(f) show easy tasks and panels (g)–(l) show hard tasks. Task objectives are listed in table 12 .
Figure 5 : Execution stages of the 12 reported real-robot tasks. Panels (a)–(f) show easy tasks and panels (g)–(l) show hard tasks. Task objectives are listed in table 12 .
ID
Task
Objective
Easy tasks
(a)
Press button
Approach and press the button to activate it.
(b)
Lift paper bag
Grasp a handle and lift the bag from the table.
(c)
Open drawer
Grasp the drawer and pull it outward.
(d)
Remove test tube
Grasp a test tube and extract it from the rack.
(e)
Sort socks
Arrange white socks on the left and black socks on the right.
Table 12 : Real-robot task objectives. Letters correspond to the panels in figure 5 .
Task
Baseline
RoboSkill
Final SR ↑
Avg. Time (min) ↓
Final SR ↑
Avg. Time (min) ↓
Easy tasks
Press button
10/10
17.68
10/10
9.10
Lift paper bag
10/10
19.58
10/10
15.70
Open drawer
8/10
22.53
10/10
15.40
Remove test tube
9/10
20.00
10/10
16.00
Table 13 : Task-wise Astra performance before and after exploration. Baseline timing uses the separate successful-sample records. Setting: 12-task subset (six easy and six hard), ten trials per task for success rates, 1 h per trial.
Task
Baseline
RoboSkill from Astra
Final SR ↑
Avg. Time (min) ↓
Final SR ↑
Avg. Time (min) ↓
Press button
0/5
–
5/5
25.40
Lift paper bag
2/5
27.00
5/5
19.60
Open drawer
1/5
27.00
3/5
34.33
Remove test tube
1/5
58.00
3/5
35.67
Sort white socks left and black socks right
0/5
–
2/5
29.00
Table 14 : Task-wise GPT-5.6 Sol performance before and after receiving skills constructed by Astra. Avg. Time uses successful trials only. Setting: six easy tasks, five trials per task, tactile input, and a 1 h limit per trial.
Target task
Rule
GPT-5.6 Sol
Opus 5
Fable 5.1
GPT-6 Astra
T0: Two cans to basket
[T0,T7,T1]
[T0,T7,T1]
[T0,T7,T1]
[T0,T7,T1]
[T0,T7]
T1: Cheese and butter to basket
[T1,T7,T0]
[T1,T7]
[T1,T7,T0]
[T1,T7,T0]
[T1,T7]
T2: Turn on stove; place moka pot
[T2,T8]
[T2,T8]
[T2,T8]
[T2,T8]
[T2,T8]
T3: Black bowl to drawer; close
[T3,T9]
[T3,T9,T5]
[T3,T9]
[T3,T9]
[T3]
T4: Two cups to left/right plates
[T4,T6,T9]
[T4,T6]
[T4,T6]
[T4,T6]
[T4,T6]
T5: Book to rear caddy slot
[T5]
[T5,T3]
[T5,T7]
[T5,T3]
[T5]
Table 15 : Agent-selected versus rule-based Top-K skill sets. Each cell lists the source-task skills selected for the target task. Bold agent-selected cells differ from the fixed rule in membership or order.
Model / setting
Final SR ↑
1st SR ↑
SR@30m ↑
Avg. Ep. ↓
Avg. Time (min) ↓
GPT-6 Astra
Rule-based Top-K
100%
87.5%
85%
1.175
19.2
Agent-selected Top-K
100%
97.5%
95%
1.025
17.1
Δ agent vs. rule
+0 pp
+10 pp
+10 pp
-12.8%
-11.2%
GPT-5.6 Sol
Rule-based Top-K
97.5%
57.5%
70%
3.425
35.4
Table 16 : Rule-based versus agent-selected Top-K performance. Δ is the change from rule-based to agent-selected Top-K (percentage points for rates and relative percent otherwise). Bold marks the better non-tied value within each model and metric.