Organizations: East China Normal University · Nanyang Technological University · Southern University of Science and Technology · Shanghai Innovation Institute
Agent skill evolution seeks to improve reusable procedural guidance for large language model (LLM) agents through iterative revision. Existing methods base each revision mainly on execution trajectories or feedback, leaving recurring behavioral requirements across tasks implicit and tying revision to the behavior of the current skill. We introduce SkillFocus, which decomposes recurring task requirements into a capability space that remains fixed as the skill evolves, separating what tasks require from how the current skill behaves. SkillFocus maps current task outcomes to this space to identify the capability that leaves the most tasks unresolved, then uses that capability to determine what to revise and which evidence to use. Across four benchmarks spanning heterogeneous tasks, SkillFocus achieves the best held-out accuracy on all four, outperforming the strongest competing result by 5.7 points on average while using 24% fewer evolution tokens on average than the closest iterative baseline. Controlled studies further show that capabilities derived from recurring task requirements outperform task-semantic and execution-derived alternatives, while randomizing task--capability assignments reduces final accuracy by up to 20.2 points. Matching evidence to the selected capability increases candidate gain by 4.4 points under prioritized revision.
Figures & tables
Figure 1: Existing and capability-space skill evolution. Both follow the same execution–revision loop. Existing methods base revisions on current execution evidence, whereas ours also derives a capability space from tasks before execution and uses it during revision.
Figure 2: Overview of SkillFocus. Before evolution, capability induction builds a fixed capability space (C,Q) from task specifications, from which the selection set is built. In each round, capability scores select c∗ , which sets the requirement to improve and the matched execution evidence; the resulting candidate replaces the current skill only after selection and training validation.
Method
SearchQA
SpreadsheetBench
LiveMath
IFBench
Avg.
No Skill
64.5
21.8
14.4
67.8
42.1
One-shot Skill
70.7
34.0
30.6
80.6
54.0
Trace2Skill
73.7
30.7
39.5
76.3
55.1
GEPA
74.6
58.9
37.9
73.4
61.2
SkillOpt
76.1
62.1
41.1
73.9
63.3
SkillFocus
80.1
77.1
43.6
82.0
70.7
Table 1: Test accuracy (%) of SkillFocus and baselines on four benchmarks. Bold and underlined values denote the best and second-best results, respectively; Avg.: unweighted mean over the four benchmarks.
Figure 3: Evolution trajectories of SkillFocus and SkillOpt. Each curve plots selection set accuracy gain over the initial skill in percentage points; markers show accepted updates, and horizontal segments are rejected candidates that consume tokens without changing the retained skill.
Figure 4: Task–capability assignments and their effect on final skill accuracy. (a) capabilities assigned to each task and the number of tasks per capability; (b) overlap between the task sets of different capabilities; (c) final skill accuracy when capabilities are derived from different sources. Panel annotations report labels per task, coverage, and median support overlap.
Condition
Accept (%)
Δsel
Δtr
TM: Top + Matched
40.0
+4.10
+9.37
T ¬ M: Top + Mismatched
26.7
+3.62
+4.97
RM: Random + Matched
26.7
+2.28
+5.33
R ¬ M: Random + Mismatched
20.0
+2.09
+3.39
Table 2: Focus–evidence crossover over 15 historical revision states, macro-averaged across SearchQA, SpreadsheetBench, and LiveMath. Accept: candidate acceptance rate; Δsel , Δtr : paired candidate gains (points) on the selection and training sets.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Benchmark
Tasks
K
Labels per task
Support
Mean J
Median J
Coverage
SearchQA
600
19
4.18
3–600
.055
.022
100%
SpreadsheetBench
120
34
2.46
3–89
.018
.000
99.2%
LiveMath
53
12
5.11
5–53
.202
.167
100%
IFBench
89
21
1.29
3–18
.018
.000
86.5%
Appendix
Table 3: Benchmark-level statistics of the frozen task–capability matrices. Tasks is the number of task records the matrix covers; K is the number of induced capabilities; Labels per task is the mean number of active capabilities per task; Support is the minimum–maximum number of supporting tasks; J is pairwise support-set Jaccard overlap; and Coverage is the fraction of covered tasks carrying at least one capability.
Benchmark
Representative behavioral requirement
# Tasks
SearchQA
Return exactly one non-empty answer element containing only the final answer text, with no reasoning, citations, or extra prose.
600
SearchQA
Use only the supplied question, documents, and metadata as evidence, avoiding outside knowledge, tools, and unsupported inference.
600
SearchQA
Select an answer only when the same candidate satisfies all substantive clue constraints, rather than combining separate partial matches.
191
SearchQA
Follow the relation expressed in the clue to extract the correct participant, such as an author, actor, founder, owner, or creator.
127
SpreadsheetBench
Write the exact formula or formula-based expression required by the task directly in the specified answer cells.
89
SpreadsheetBench
Generate each output with correct row-specific references, conditions, and source matches rather than positional assumptions.
31
Appendix
Table 4: Representative behavioral requirements and their task support.
capability source
SearchQA
SSB
LiveMath
SkillFocus
80.1
77.1
43.6
task-semantic
61.3
68.9
25.8
execution-derived
73.5
66.1
30.7
randomized assignment
69.0
59.6
23.4
Appendix
Table 5: Test accuracy (%) when capabilities are derived from different sources, under the same downstream SkillFocus pipeline.
Benchmark
Views
K
Common refs.
ARI
NMI
SearchQA
3
18/13/17
47.7
.883
.953
SSB
3
17/16/19
35.7
.913
.976
LiveMath
3
11/11/10
33.3
.973
.989
IFBench
3
30/24/42
112.3
.912
.974
Appendix
Table 6: Assignment agreement across three independent proposal calls on the same task specifications. K : proposal-specific capability counts; Common refs.: aligned behavioral-requirement references; ARI and NMI: adjusted Rand index and normalized mutual information between proposal assignments.
Condition
Accept (%)
Δsel
Δtr
SearchQA
TM
40.0
+5.30
+10.50
T ¬ M
20.0
+5.30
+3.42
RM
40.0
+7.50
+3.92
R ¬ M
40.0
+3.13
+1.92
SpreadsheetBench
Appendix
Table 7: Per-benchmark focus–evidence crossover (five historical revision states per benchmark). Accept: candidate acceptance rate; Δsel , Δtr : paired candidate gains (points) on the selection and training sets.
Variant
SearchQA
SSB
LiveMath
SkillFocus
80.1
77.1
43.6
w/o matched evidence
71.4
58.6
20.8
w/o capability prioritization
68.7
38.3
20.8
Appendix
Table 8: Ablations of capability-guided revision on SearchQA, SpreadsheetBench, and LiveMath (test accuracy, %).
Benchmark
capability
Rounds
Support
Share
SearchQA
c1
3
600
100.0%
SearchQA
c2
1
600
100.0%
SearchQA
c3
1
289
48.2%
SpreadsheetBench
c1
5
89
74.2%
LiveMath
c1
2
53
100.0%
LiveMath
c3
2
31
58.5%
Appendix
Table 9: Capabilities selected by the priority rule over the five rounds of each formal run. Support is the number of tasks requiring the capability in the frozen Q ; Share is its fraction of the task population. A capability is selected in consecutive rounds when it is the only eligible one.
Benchmark
Record
T/N n
ΔT/ΔN
Rd(e)
SQA
Screen.
93/107
+1.1/+0.0
+1.1
SSB
Acc.-1
89/31
+46.1/+16.1
+29.9
LM
Acc.
31/22
+6.5/+4.5
+1.9
IFB
Acc.
3/86
+0.0/+8.1
−8.1
Appendix
Table 10: Target and non-target edit responses. T/N: group sizes; ΔT/ΔN : net changes (points); Rd(e) : their difference; Screen./Acc.: screening record/accepted candidate (Acc.-1: first accepted).
SkillFocus
SkillOpt
Benchmark
Sel.
Test
Sel.
Test
SearchQA
49.5
80.1
76.0
76.1
SpreadsheetBench
80.0
77.1
85.0
62.1
LiveMath
38.9
43.6
44.4
41.1
IFBench
77.8
82.0
89.7
73.9
Appendix
Table 11: Final selection set and held-out test accuracy (%). The two columns cover different task sets, so their difference is not a calibrated generalization gap across methods.
Figure 5: Representative capability-guided revision on SpreadsheetBench. The case connects diverse task requests to a recurring behavioral requirement, an evidence-grounded revision, and paired target/non-target responses.
Group
n
Repair/reg.
Net Δ
Target
89
42/1
+46.1%
Non-target
31
7/2
+16.1%
Appendix
Table 12: Paired response of the representative SpreadsheetBench revision. Repair/reg. counts failure-to-success and success-to-failure changes; the target-minus-non-target contrast is +29.9 points.
Benchmark
Task format
Evaluation
SearchQA
Jeopardy!-style clues answered in a closed-context, single-turn setting
Deterministic answer matching
SpreadsheetBench
Spreadsheet manipulation on the Verified subset; the agent reads and writes workbook files
Target cells compared with the reference workbook
LiveMath
Multiple-choice mathematical reasoning with deterministically shuffled options; code execution is disabled
Exact match of the selected option
IFBench
Single-turn free-text responses under verifiable output constraints
Programmatic constraint verification
Appendix
Table 13: Benchmarks used in the experiments. Task format states what the agent reads and produces; Evaluation states the deterministic checker applied to each response.
Condition
No Skill
One-shot Skill
Trace2Skill
GEPA
SkillOpt
SkillFocus
Evaluator and test split
Shared
Shared
Shared
Shared
Shared
Shared
Target model
DeepSeek-V4-Flash
DeepSeek-V4-Flash
DeepSeek-V4-Flash
DeepSeek-V4-Flash
DeepSeek-V4-Flash
DeepSeek-V4-Flash
Initial skill
–
–
Shared
Shared
Shared
Shared
Optimizer
–
–
GPT-5.4
GPT-5.4
GPT-5.4
GPT-5.4
Evolution budget
–
–
One analysis and consolidation pass
160 metric calls
Official setting
At most five rounds, one screened candidate per round
Candidate retention
–
–
Single consolidated skill
Pareto-based selection
Retained-skill update
Positive paired change on selection and training sets
Appendix
Table 14: Shared and method-specific conditions across compared methods. “–” denotes a condition that does not apply to the method.
Benchmark
C/A
Init.
Iter.
Cevo
Test
SQA
5/2
5.10M
15.55M
20.66M
8.51M
SSB
5/2
6.45M
22.61M
29.06M
4.73M
LM
5/2
0.38M
5.98M
6.36M
3.40M
IFB
5/1
0.97M
2.48M
3.45M
1.58M
Appendix
Table 15: SkillFocus inference-token ledger (millions). SQA, SSB, LM, and IFB abbreviate the four benchmarks; C/A: proposed/accepted candidates; Init., Iter.: initialization and iterative evolution.
Benchmark
SkillFocus
SkillOpt
Relative change
SearchQA
20.66
45.21
−54.3%
SpreadsheetBench
29.06
40.55
−28.3%
LiveMath
6.36
8.14
−21.9%
IFBench
3.45
3.21
+7.5%
Mean
–
–
−24.0%
Appendix
Table 16: Pre-final-test evolution cost (millions of inference tokens) of SkillFocus and SkillOpt. The final row averages the per-benchmark relative changes.
Interface
Role
Calls
capability induction
ExtractRequirements
task → requirements
per task
ProposeCapabilities
cross-task proposals
3
Consensus
frozen registry
1
Assign
requirement → capability
batched
capability revision
Appendix
Table 17: Overview of the prompt interfaces and their call multiplicity in the formal runs.
Stage / perspective
Frozen operational condition
Diagnosis: missing guidance
No current-skill rule addresses the operation; the explorer leaves cited_rules empty.
Diagnosis: ineffective guidance
An existing rule addresses the operation but fails; the explorer quotes at least one such rule.
Diagnosis: behavior organization
Interacting rules create an ordering, conditioning, or precedence problem; the explorer quotes at least two rules.
Generation: minimal addition
Propose one non-duplicative standalone rule with op=add .
Generation: revise existing guidance
Replace the core content of one existing rule, quoting the rule being revised.
Generation: structural change
Change a rule’s trigger, ordering, precedence, interaction, or removal target, quoting the affected rule(s).
Appendix
Table 18: Frozen Explorer perspectives in the formal runs.
Natural-language skills are textual procedural memories through which large language model (LLM) agents retain reusable task knowledge without updating model weights. Existing methods typically treat skills as either static artifacts or monolithic documents optimized using aggregate validation scores as feedback. However, representing a skill as a monolithic document restricts optimization to its textual content, without explicitly modeling the structure through which procedural knowledge is retrieved and executed. We identify a key distinction between learning what knowledge to retain and determining how to organize it: textual updates should first be validated through execution evidence, after which the retained knowledge should be structured according to its procedural dependencies and retrieval requirements. To this end, we introduce SkillSpec, a two-phase framework comprising consensus-gated evolution and representation specialization. In the consensus-gated phase, complementary editing intents generate complete candidate skills. An update is committed only when paired evaluations reach consensus, requiring sufficient overall improvement and non-negative aggregate paired gain in every repeated evaluation. In the specialization phase, signals of process and redundancy sensitivity derived from the full optimization trajectory, including accepted and rejected candidates, guide the selection of a flat, graph, or hybrid representation.Across six benchmarks and three target language models, SkillSpec improves average success rate over SkillOpt by 6.89%, averaged across the three models. These results demonstrate that reliable skill evolution and representation specialization address complementary objectives: deciding what knowledge to retain and how to structure it for inference.
Skill documents, structured natural-language instructions that guide Large Language Model (LLM) agents, are critical to modern agent frameworks, yet LLMs struggle to write skills that actually work. On SkillsBench, human-authored skills improve pass rates by 16.2 percentage points, while LLM-authored skills provide no measurable gain. We introduce SkillAxe, a fully unsupervised framework that enables LLMs to iteratively diagnose and refine their own skills. SkillAxe decomposes skill quality into four interpretable dimensions (quality impact, trigger precision, instruction compliance with fault attribution, and solution-path coverage), producing structured improvement briefs that require no ground-truth labels, test suites, or environment rewards. On SkillsBench, SkillAxe improves pass rates by 28% relative over unimproved LLM skills and closes 47--67% of the gap to human-authored skills. We validate the approach as a continuous improvement engine in the wild on SpreadsheetBench, where a SkillAxe-built skill library learns from past agent trajectories and raises pass rate from 16.0% to 52.0% using only 22 skills.
Agent skills provide a lightweight way to adapt LLM agents to specialized domains by storing reusable procedural knowledge in structured files. However, whether downloaded from third parties or self-generated, these skills are often unreliable, incomplete, or outdated. Existing skill-evolution methods often address these deficiencies through heuristic reflections without an explicit optimization formulation. In this paper, we propose SkillGrad, a gradient-descent-inspired framework for optimizing agent skills. SkillGrad treats the skill package as a structured parameter to optimize in a gradient descent fashion: task executions provide trajectory-level loss evidence, automatic diagnoses then provide text-based gradients that indicate the correction directions. To stabilize optimization across iterations, a momentum agent accumulates recurring diagnostic patterns into a persistent memory overlay. Finally, an LLM-based patcher executes the parameter update by applying layer-aware edits to the skill package. Evaluated on SpreadsheetBench Verified and WikiTableQuestions, SkillGrad consistently outperforms training-based skill evolution baselines across two backbone LLMs, improving over the strongest training-based baseline by 6.7 percentage points on average. Ablations further show that momentum and contrastive diagnosis both contribute to the final skill quality.
Hanyu Wang, Yifan Lan, Bochuan Cao +2
College of Information Sciences and Technology The Pennsylvania State University University Park, PA, USA