Continual skill evolution enables LLM agents to accumulate and refine reusable procedural knowledge from interaction experience without updating model parameters. Its effectiveness depends on determining not only what to change, but also why a change is justified and when it should become persistent guidance. However, existing experience-driven methods can lose the behavioral evidence and task contexts supporting edits. Moreover, a global validation outcome provides an incomplete judgment of its constituent changes: locally supported corrections may be discarded with a rejected revision, while evidence may require further experience to inform useful updates. To this end, we introduce EVISKILL, an evidence-driven framework that organizes execution observations into Replayable Evidence Cards and synthesizes edits with explicit links to their supporting contexts. Targeted replay verifies these edits through re-execution and provides feedback for correction. Across epochs, EVISKILL preserves evidence and provisionally retains supported edits for further refinement, while global validation governs their incorporation into the final skill. Experiments on three interactive benchmarks across six LLM backbones demonstrate the effectiveness of this approach.
Figures & tables
Figure 1: Experience-driven Skill Evolution vs Evidence-grounded Skill Evolution.
Figure 2: Re-execution outcomes of Experience-Derived Edits.
Figure 3: Performance of selectively retained edits from globally rejected revisions.
Figure 4: Overview of EviSkill . Current trajectories and cross-epoch comparisons provide evidence for edit synthesis. Replay verifies and refines edits, while global validation and post-rejection replay govern skill updates, provisional edit retention, and evidence reuse across epochs.
Model
Method
ALFWorld (134)
AppWorld (168)
ScienceWorld (211)
Acc. ↑
ΔNS
Acc. ↑
ΔNS
Acc. ↑
ΔNS
GPT-5.5
No Skill
90.30
0.00
85.71
0.00
76.30
0.00
LLM Skill
94.03
+3.73
94.05
+8.33
76.78
+0.47
Trace2Skill
94.03
+3.73
86.31
+0.60
65.40
-10.90
EvoSkill
94.78
+4.48
91.07
+5.36
54.50
-21.80
SkillGrad
95.52
+5.22
86.31
+0.60
67.30
-9.00
Table 1: Task-success accuracy (%) across three interactive benchmarks. ΔNS denotes the percentage-point change relative to No skill for the same model and dataset. Bold and underlined values indicate the best and second-highest distinct accuracies within each model–dataset setting.
Figure 5: Component ablations averaged over the three GPT backbones.
Figure 6: Replay-guided edit verification. (a) Initial decisions over all candidate edits. (b) Refinement outcomes among edits initially assigned to reflection, showing counts and within-group percentages. Acceptance indicates local replay support, not global revision acceptance.
Figure 7: Promotion of the 96 provisionally retained edits.
Figure 8: Number of distinct epochs in which each Evidence Card is selected for an Evidence Window.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Task Families
Train
Val.
Test
AppWorld
-
90
57
168
ScienceWorld
24
96
48
211
ALFWorld
6
96
48
134
Appendix
Table 2: Task families and dataset splits.
Method
Evolution strategy
Core mechanism
Non-Evolving Baselines
NoSkill
None
No external skill
LLM Skill
None
Fixed LLM-generated guidance
Skill-Evolution Methods
Trace2Skill
Trajectory distillation
Lesson extraction and consolidation
EvoSkill
Iterative discovery
Skill generation and refinement
Appendix
Table 3: Overview of the evaluated methods and their skill-evolution strategies.
Model
Method
ALFWorld (134)
AppWorld (168)
ScienceWorld (211)
Acc. ↑
ΔNS
Acc. ↑
ΔNS
Acc. ↑
ΔNS
GPT-5.5
No Skill
90.30
0.00
85.71
0.00
76.30
0.00
LLM Skill
94.03
+3.73
94.05
+8.33
76.78
+0.47
Trace2Skill
94.03
+3.73
86.31
+0.60
65.40
-10.90
EvoSkill
94.78
+4.48
91.07
+5.36
54.50
-21.80
SkillGrad
95.52
+5.22
86.31
+0.60
67.30
-9.00
Appendix
Table 4: Task-success accuracy (%) across three interactive benchmarks. ΔNS denotes the percentage-point change relative to No skill for the same model and dataset. Bold and underlined values indicate the best and second-highest distinct accuracies within each model–dataset setting.
Figure 9: Family-level component ablations across all six Action-Agent backbones. Each panel reports task-success accuracy averaged over the three backbones in the corresponding model family. w/o Replay removes Replay-Guided Edit Verification, while w/o Cross-Epoch disables Cross-Epoch Evidence Propagation.
Variant
Accuracy (%)
Gain
Full
88.06
+3.48
Cards Only
86.32
+1.74
Edits Only
86.82
+2.24
w/o Cross-Epoch
84.58
0.00
Appendix
Table 5: Cross-epoch ablations.
Figure 10: Replay-supported correction of origin edit lineages.
Dataset
Support
Cards
Edit-linked (of Cards)
Replay-retained (of replayed)
Final Provisional (of Cards)
Validated (of Cards)
Prov./validated (of retained)
ALFWorld
Single-task
532
151 (28.38)
139/151 (92.05)
30 (5.64)
92 (17.29)
122/139 (87.77)
Cross-task
389
154 (39.59)
136/154 (88.31)
17 (4.37)
110 (28.28)
127/136 (93.38)
AppWorld
Single-task
695
580 (83.45)
451/570 (79.12)
32 (4.60)
347 (49.93)
379/451 (84.04)
Cross-task
123
87 (70.73)
59/85 (69.41)
3 (2.44)
53 (43.09)
55/59 (93.22)
ScienceWorld
Single-task
1,223
1,182 (96.65)
818/1,159 (70.58)
0 (0.00)
732 (59.85)
730/818 (89.24)
Cross-task
351
343 (97.72)
234/338 (69.23)
0 (0.00)
212 (60.40)
211/234 (90.17)
Appendix
Table 6: Card utilization by support breadth. Percentages use the denominators specified in the column headers.
Dataset
Cards
Ranges
Steps
Steps/Card
ALFWorld
923
1,482
8,160
8.84
AppWorld
828
1,035
5,520
6.67
ScienceWorld
1,610
2,142
20,093
12.48
Overall
3,361
4,659
33,773
10.05
Appendix
Table 7: Supporting trigger-range statistics.
Figure 11: Cross-epoch lifecycle of Evidence Cards over four skill-evolution epochs. (a) Lifecycle-state shares among Cards processed from each epoch’s input Evidence Pool, pooled across 18 runs; n gives the input pool size. (b) Distribution of unique Evidence Cards by the number of distinct epochs in which they are assigned to an Evidence Window. Pre-resolved denotes Cards resolved without any Evidence-Window assignment. Percentages are computed within each benchmark, and All is the count-weighted aggregate.
Figure 12: Post-rejection replay and subsequent edit incorporation.
Method
Mean train rollouts
Mean val. gain (pp)
ALFWorld
AppWorld
ScienceWorld
Overall
EviSkill
376.00
14.25
1.99
3.01
6.33
3.77
EvoSkill
376.00
5.86
0.18
1.22
3.26
1.55
SkillGrad
119.94
1.18
0.07
0.91
1.53
0.84
Appendix
Table 8: Validation gain per 100 training rollouts on the common 18-cell subset. “Mean rollouts” and “mean gain” are cell-macro averages; gain is measured in percentage points (pp). Dataset and overall columns report E100 in validation-accuracy points per 100 training rollouts. Bold denotes the highest value in each comparison column.
Replay stage
Ranges
Full-trajectory replay steps
Localized replay steps
Copied prefix steps
Reduction
Evidence-window replay
2,505
53,242
18,007
20,714
66.18%
Post-rejection replay
207
3,863
1,205
1,629
68.81%
Appendix
Table 9: Replay workload comparison between trigger-range replay and full-trajectory replay.
Figure 13: Acceptance. The Evidence-supported search edit changes behavior on the bound replay range, discovers the previously missed target, and completes the task, leading to accept decision.
Figure 14: Rejection due to behavioral noncompliance. Replay skips the comparison required by the edit, resulting in reject .
Figure 15: Reflection followed by acceptance. Initial replay exposes a missing exact-instance binding requirement in an otherwise useful placement rule. Reflection repairs this omission, and the revised edit subsequently completes the placement successfully, resulting in an accept decision.
Figure 16: Reflection followed by rejection. Reflection removes the conflicting life-stage heuristic and repairs the written lifespan rule, but the corrected instruction remains ineffective on one replay range. The revised edit is therefore not retained.
Task
Linked trigger range (intermediate navigation omitted)
Task1
take potato 1; heat it with microwave 1; move potato 1 to diningtable 1.
Task2
take bowl 2; cool it with fridge 1; move bowl 2 to cabinet 1.
Task3
take bread 2; heat it with microwave 1; move bread 2 to diningtable 1.
Task4
take fork 1; clean it with sinkbasin 1; move fork 1 to countertop 1.
Shared rule
Maintain the exact acquired instance through transformation and final placement; do not substitute another visible instance.
Appendix
Table 10: Four task-specific trigger ranges linked to a shared Evidence Card.
Figure 17: Cross-epoch evidence reveals a persistent edit deficiency, guiding correction through reflection and replay.
Although agent skills equip LLMs with reusable procedural knowledge, manual maintenance suffers from high costs, unscalability, and misalignment. Real-world deployments thus require autonomous, on-demand skill evolution at test time, constrained by limited interaction budgets and a lack of training or validation sets. This setting introduces a severe sparse reward challenge, where outcomes conflate multiple latent failure causes. Under such ambiguity, existing methods that greedily refine a single incumbent skill are particularly vulnerable to an exploitation trap, allowing early misdiagnoses to exhaust limited trials along unproductive trajectories. To address this, we introduce SkillHEX, a closed-loop framework coupling hypothesis-driven self-verification with evidence-guided tree search. SkillHEX translates falsifiable failure hypotheses into executable tests, producing diagnostic evidence as dense reward without additional environment attempts. This evidence guides a search over persistent skill-revision branches, dynamically balancing the exploitation of supported edits with the exploration of plausible alternatives. Evaluated on 87 tasks from SkillsBench, SkillHEX outperforms existing self-evolving methods and achieves an average pass rate of 55.9% and 57.9% using GPT-5.3-Codex and Claude Opus 4.7 under a five-iteration budget, respectively.
Yuru Feng, Yaoqi Chen, Beidi Zhao +7
Microsoft · University of California, San Diego · University of Science and Technology of China +1
Agent skills are procedural artifacts that enable LLM agents to execute workflows, verify constraints, and recover from failures. Existing self-evolving methods refine skills using accumulated trajectories. However, they struggle in cold-start settings, where only an initial, imperfect skill is available. Consequently, skill construction defaults to expert authoring or one-shot LLM generation. Expert-authored skills are costly and may not align with how LLM agents actually execute tasks, while one-shot generated skills can be syntactically well formed yet behaviorally weak. To bridge this gap, we propose SkillRevise, an execution-grounded framework designed to iteratively refine these initial skills. SkillRevise diagnoses skill defects from execution evidence, retrieves relevant repair principles from a general memory, and applies execution-anchored edits. By re-executing candidates, it retains the first verifier-passing skill within the revision budget and falls back to empirical utility only when no candidate succeeds. Evaluated across three benchmarks and five LLMs, SkillRevise substantially outperforms one-shot baselines, improving the base agent's success rate on SkillsBench from 36.05% to 61.63%. Furthermore, the revised skills transfer across both executors and task environments, suggesting that SkillRevise captures reusable procedural knowledge beyond any single executor.
Yuxuan Liu, Zhaochen Su, Lingyun Xie +11
The Hong Kong University of Science and Technology · Harbin Institute of Technology · Harbin Institute of Technology, Shenzhen +2
LLM training is shifting from manual design and annotation to interaction-driven self-evolution. However, existing self-evolutionary methods face a fundamental dilemma between task diversity and verification reliability: environment-bound methods obtain precise feedback but confine learning to narrow domains, while open-ended self-generation broadens the task space but lacks reliable verification, allowing misleading rewards to pollute the training loop. We identify agent skills as a powerful middle ground to reconcile this tension: each skill ensures deep, verifiable execution in a specific scenario, while dynamic routing across skills maintains open-ended task variety. Leveraging this insight, we introduce Skill Self-Play (Skill-SP), a co-evolutionary framework comprising a proposer, a solver, and a dynamic skill controller. Orchestrated via a reinforcement learning loop, these components co-evolve in a continuous self-play loop: the proposer generates challenging tasks conditioned on dynamically sampled skills; the solver explores candidate solutions to push its capability boundaries; and the skill controller collects execution feedback to update and expand the skill library. This interactive co-evolution effectively bridges the gap between structured verification and open-ended exploration. Empirical evaluations on tool-use and reasoning benchmarks demonstrate that Skill-SP, serving as a robust evolution engine, consistently pushes the performance ceiling of competent backbones while catalyzing striking turnarounds for initially misaligned models. Our code is available at https://github.com/Qwen-Applications/skill-self-play.
Siyuan Huang, Pengyu Cheng, Haotian Liu +10
Qwen Large Model Application Team, Alibaba · The Chinese University of Hong Kong · Renmin University of China +5