Natural-language skills are textual procedural memories through which large language model (LLM) agents retain reusable task knowledge without updating model weights. Existing methods typically treat skills as either static artifacts or monolithic documents optimized using aggregate validation scores as feedback. However, representing a skill as a monolithic document restricts optimization to its textual content, without explicitly modeling the structure through which procedural knowledge is retrieved and executed. We identify a key distinction between learning what knowledge to retain and determining how to organize it: textual updates should first be validated through execution evidence, after which the retained knowledge should be structured according to its procedural dependencies and retrieval requirements. To this end, we introduce SkillSpec, a two-phase framework comprising consensus-gated evolution and representation specialization. In the consensus-gated phase, complementary editing intents generate complete candidate skills. An update is committed only when paired evaluations reach consensus, requiring sufficient overall improvement and non-negative aggregate paired gain in every repeated evaluation. In the specialization phase, signals of process and redundancy sensitivity derived from the full optimization trajectory, including accepted and rejected candidates, guide the selection of a flat, graph, or hybrid representation.Across six benchmarks and three target language models, SkillSpec improves average success rate over SkillOpt by 6.89%, averaged across the three models. These results demonstrate that reliable skill evolution and representation specialization address complementary objectives: deciding what knowledge to retain and how to structure it for inference.
Figures & tables
Figure 1 : Overview of SkillSpec. Phase I combines multi-intent candidate generation with repeated paired validation to produce a validated backbone and optimization trajectory. Phase II estimates PSS and RSS from accepted and rejected candidates, selects a flat , graph , or hybrid representation, and specializes a derived artifact while preserving the source backbone.
Figure 2 : Illustrative skill representations. Flat organizes guidance as independently retrievable units; graph encodes explicit dependencies among units; and hybrid combines an ordered core with auxiliary modules.
Structure
Condition
Specialization
flat
ψ≤ζψ
Initialize a derived copy from Φ⋆ , split it into self-contained guidance units, and add, revise, merge, or remove units without explicit dependency edges or a prescribed order.
graph
ψ>ζψ , ρ≤ζρ
Initialize a derived graph from Φ⋆ ; jointly add, revise, or remove nodes and directed dependency edges; and preserve a valid execution order.
hybrid
Otherwise
Initialize an ordered core from Φ⋆ ; add, revise, or remove auxiliary modules and update their attachment points.
Table 1 : Adaptive representation selection from PSS ( ψ ) and RSS ( ρ ), where ζψ and ζρ denote the corresponding thresholds and Φ⋆ denotes the frozen validated backbone.
Model
Method
SearchQA
Spreadsheet
OfficeQA
DocVQA
LiveMath
ALFWorld
Average
Improvement
GPT–4.1
Baseline
69.57
36.07
1.16
66.31
25.00
44.78
40.48
–
GEPA
79.71
35.36
12.79
80.75
29.03
64.93
50.43
+9.95
Trace2Skill
75.50
41.79
13.95
78.88
33.06
44.78
47.99
+7.51
SkillAdam
80.93
46.79
10.47
82.62
29.03
55.22
50.84
+10.36
SkillOpt
79.93
46.43
11.05
73.26
29.03
52.99
48.78
+8.30
SkillOpt-lite
75.71
48.93
12.21
69.79
28.23
64.93
49.96
+9.48
Table 2 : Success rate (%) on the complete test set and six-benchmark macro averages. Improvement is the absolute success-rate difference (%) from the no-skill baseline. SkillSpec-I denotes the consensus-gated Phase I output, and SkillSpec denotes its Phase II specialized result. The GPT–4.1 LiveMath pair uses the legacy confidence-gated routing protocol, which selected Flat, with GPT–4.1 as the Phase II optimizer. Best and second-best results are bolded and underlined, respectively.
Figure 3 : Optimization trajectories during Phase I across six benchmarks for GPT–4.1, GPT–5.4 Nano, and GPT–5.4. All trajectories use 10 rounds and three validation repeats.
Variant
SearchQA
Spreadsheet
OfficeQA
DocVQA
LiveMath
ALFWorld
Avg.
Phase I, equal compute
66.57
46.79
4.07
65.24
45.97
51.49
46.69
Phase II, Flat only
68.50
46.79
4.65
66.58
41.94
52.99
46.91
Phase II, Graph only
67.86
43.57
6.98
67.11
45.97
55.97
47.91
Phase II, Hybrid only
66.93
49.29
5.81
62.30
53.23
57.01
49.10
PSS ( ψ )
0.428
0.925
0.619
0.411
0.998
0.778
–
RSS ( ρ )
0.918
0.837
0.499
0.653
1.000
0.889
–
Table 3 : Controlled representation ablation on GPT–5.4 Nano, reported separately from the main-table runs. Within each benchmark, Flat, Graph, and Hybrid are initialized from the same Phase I backbone and use the same specialization budget. Boxes indicate the branches selected by PSS/RSS. Avg. is the unweighted mean across six benchmarks.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter
Value
Evaluated models
GPT–4.1 / GPT–5.4 Nano / GPT–5.4
Optimization data
Official training and validation splits
Final evaluation
Complete official test split
Primary metric
Hard accuracy
Aggregation
Macro average across six benchmarks
Failure accounting
Agent, setup, and timeout failures count as incorrect
Appendix
Table 4 : Experimental and specialization settings used for SkillSpec unless noted otherwise.
Benchmark
Content role
Purpose
Example content
LiveMath
Procedural specialization
Distinguish theorem-style statements with subtle strength differences.
Before choosing, rewrite each option as a formal claim: hypotheses, quantifiers, conclusion strength, and endpoint cases. Rank options by logical implication, not by thematic similarity. Do not stop at the first true-looking statement; choose the strongest statement actually proved.
DocVQA
Evidence specialization
Preserve literal values and provenance from a document image.
Ground the answer in visible document evidence. Prefer the exact string, unit, date, or entity shown in the source. If multiple nearby values appear, compare labels and table headers before answering.
OfficeQA
Backbone-anchored process
Support multi-turn file/search tasks that require verified evidence.
Search broadly, then narrow by entity/date/table. Read the source document before answering. If evidence is missing, reformulate the query. Track provenance and answer only from retrieved evidence; do not answer from memory.
SpreadsheetBench
Transformation specialization
Preserve unrelated workbook state during a requested edit.
Use the full workbook, not only the preview. Iterate over all actual rows and sheets. Preserve unrelated cells and formulas. Write outputs to the requested sheet/cell range and validate against the requested answer position.
ALFWorld
Backbone-anchored process
Preserve state across long action sequences and recovery steps.
Track the goal, current room, inventory, and object state. Explore systematically; after a failed action, inspect the observation and choose a recovery action. Use containers and object affordances explicitly; stop only after the goal is satisfied.
Appendix
Table 5 : Representative specialization content. Text is lightly abbreviated for readability.
Score
Process structure
Interference
0
No grounded addition or strengthening of a dependency.
No grounded addition of interfering guidance.
1
Weak or local prerequisite, state handoff, or conditional dependency.
Weak or local duplication, conflict, irrelevance, or excessive detail.
2
Clear dependency in the revised solution procedure.
Clear interfering guidance in the revision.
3
Dependency structure is a dominant feature of the edit.
Interfering guidance is a dominant feature of the edit.
Appendix
Table 6 : Candidate-diff annotation rubric. Strength describes the visible edit, not its measured benefit or the intrinsic complexity of the task.
Form
Organization of shared guidance
Flat
Independently retrievable units for locating, acquiring, transforming, and delivering the object, alongside exploration and progress/loop guidance. No dependency graph schedules retrieval; a retrieved unit can still state a local prerequisite such as “transform before placing.”
Graph
Explicit dependencies connect locate → acquire → transform → deliver. Container inspection refines locating, while state and progress checks provide guidance at the relevant nodes. Edges make prerequisite order explicit.
Hybrid
The locate–acquire–transform–deliver chain forms the ordered core. Container exploration, progress tracking, and loop recovery remain auxiliary modules that can be retrieved when relevant, rather than expanding every core step with all guidance.
Appendix
Table 7 : Three schematic organizations of the same archived ALFWorld guidance. The content is held fixed to make structural differences explicit.
Step
Action
Observed result (condensed)
0
go to fridge 1
The fridge is closed.
1
open fridge 1
Several objects are visible, but no egg.
2
go to microwave 1
The microwave is closed.
3
open microwave 1
The microwave is empty.
4
go to fridge 1
The same fridge contents are observed again.
5
go to countertop 1
An egg is visible among other objects.
Appendix
Table 8 : Archived training execution illustrating state dependencies in the shared backbone. This is one observed episode, not a three-way representation comparison.
Agent skill evolution seeks to improve reusable procedural guidance for large language model (LLM) agents through iterative revision. Existing methods base each revision mainly on execution trajectories or feedback, leaving recurring behavioral requirements across tasks implicit and tying revision to the behavior of the current skill. We introduce SkillFocus, which decomposes recurring task requirements into a capability space that remains fixed as the skill evolves, separating what tasks require from how the current skill behaves. SkillFocus maps current task outcomes to this space to identify the capability that leaves the most tasks unresolved, then uses that capability to determine what to revise and which evidence to use. Across four benchmarks spanning heterogeneous tasks, SkillFocus achieves the best held-out accuracy on all four, outperforming the strongest competing result by 5.7 points on average while using 24% fewer evolution tokens on average than the closest iterative baseline. Controlled studies further show that capabilities derived from recurring task requirements outperform task-semantic and execution-derived alternatives, while randomizing task--capability assignments reduces final accuracy by up to 20.2 points. Matching evidence to the selected capability increases candidate gain by 4.4 points under prioritized revision.
Ning Wang, Zhiren Gong, Bingdong Li +2
East China Normal University · Nanyang Technological University · Southern University of Science and Technology +1
Agent skills provide a lightweight way to adapt LLM agents to specialized domains by storing reusable procedural knowledge in structured files. However, whether downloaded from third parties or self-generated, these skills are often unreliable, incomplete, or outdated. Existing skill-evolution methods often address these deficiencies through heuristic reflections without an explicit optimization formulation. In this paper, we propose SkillGrad, a gradient-descent-inspired framework for optimizing agent skills. SkillGrad treats the skill package as a structured parameter to optimize in a gradient descent fashion: task executions provide trajectory-level loss evidence, automatic diagnoses then provide text-based gradients that indicate the correction directions. To stabilize optimization across iterations, a momentum agent accumulates recurring diagnostic patterns into a persistent memory overlay. Finally, an LLM-based patcher executes the parameter update by applying layer-aware edits to the skill package. Evaluated on SpreadsheetBench Verified and WikiTableQuestions, SkillGrad consistently outperforms training-based skill evolution baselines across two backbone LLMs, improving over the strongest training-based baseline by 6.7 percentage points on average. Ablations further show that momentum and contrastive diagnosis both contribute to the final skill quality.
Hanyu Wang, Yifan Lan, Bochuan Cao +2
College of Information Sciences and Technology The Pennsylvania State University University Park, PA, USA
Textual skills enable large language model (LLM) based agents to accumulate reusable procedural knowledge without updating model parameters. Yet existing skill evolution remains largely confined to the text space: an optimizer must diagnose success and failure patterns, and revise skills solely from long execution trajectories and sparse task outcomes. This text-only paradigm leaves the agent's internal representations, which contain rich records of its evolving execution state, outside the skill optimization loop. We ask whether an agent can improve its external textual skills by reflecting on its own internal representations. We introduce Rep2Skill, a representation-guided framework for self-evolution on agent skills. Specifically, upon the collected agent rollouts, Rep2Skill models their internal model representation trajectories to localize turns that deviate from successful execution dynamics, and it further interprets these signals alongside the execution contexts as actionable textual feedback for targeted skill revision. Experiments on two agent environments with two open-source LLMs show that Rep2Skill consistently outperforms text-only approaches in the self-evolution setting, where the same LLM serves as both executor and optimizer without a stronger external model. This establishes a promising direction moving agent self-improvement beyond text-only reflection.