Textual skills enable large language model (LLM) based agents to accumulate reusable procedural knowledge without updating model parameters. Yet existing skill evolution remains largely confined to the text space: an optimizer must diagnose success and failure patterns, and revise skills solely from long execution trajectories and sparse task outcomes. This text-only paradigm leaves the agent's internal representations, which contain rich records of its evolving execution state, outside the skill optimization loop. We ask whether an agent can improve its external textual skills by reflecting on its own internal representations. We introduce Rep2Skill, a representation-guided framework for self-evolution on agent skills. Specifically, upon the collected agent rollouts, Rep2Skill models their internal model representation trajectories to localize turns that deviate from successful execution dynamics, and it further interprets these signals alongside the execution contexts as actionable textual feedback for targeted skill revision. Experiments on two agent environments with two open-source LLMs show that Rep2Skill consistently outperforms text-only approaches in the self-evolution setting, where the same LLM serves as both executor and optimizer without a stronger external model. This establishes a promising direction moving agent self-improvement beyond text-only reflection.
Figures & tables
Figure 1: Motivation of Rep2Skill . Top: Rep2Skill introduces internal representation signals into the iterative skill optimization loop to provide fine-grained guidance. Bottom: standard skill evolution relies on textual trajectories alone, while Rep2Skill leverages representation-guided evidence from grouped rollouts to support more targeted skill updates.
Figure 2: Overview of Rep2Skill . Given grouped rollouts under the current skill sn , a trajectory model trained on successful executions scores each turn by its representation prediction error, and the top- Kc critical turns are verbalized into textual feedback F(k) . The optimizer then revises sn with this feedback through a validation-gated update. The same frozen LLM πθ serves as executor, analyzer and optimizer.
Model
Method
ALFWorld
WebShop
Succ (%) ↑
Score ↑
Succ (%) ↑
Qwen3-4B
No Skill
22.64 ± 1.76
41.86 ± 1.65
13.67 ± 1.53
Trace2Skill ( Ni et al., 2026 )
47.26 ± 5.18
45.68 ± 3.40
13.00 ± 1.41
EvoSkill ( Alzubi et al., 2026 )
43.03 ± 3.52
49.48 ± 2.02
13.33 ± 0.94
SkillOpt ( Yang et al., 2026a )
46.51 ± 3.12
39.04 ± 2.49
11.67 ± 3.40
Rep2Skill (Ours)
52.24 ± 3.22
49.79 ± 5.95
16.33 ± 3.51
Table 1: Self-evolving skill performance on ALFWorld and WebShop. All results are mean ± standard deviation over 3 independent evolution runs.
Input
AUROC ↑
AUPRC ↑
P@Top- 3↑
Trajectory Text
0.494
0.736
63.66%
Rep.
0.592
0.789
79.26%
Text + Rep. ( Ours )
0.838
0.876
88.22%
Table 2: Turn-level error detection under different attribution signals. Representation-space signals provide complementary information to trajectory text. Both the execution model and analyzer model are Qwen3-4B ( Yang et al., 2025 ) .
Figure 3: Turn-level non-progress detection with different attribution inputs. Precision@K and Recall@K are macro-averaged over eligible trajectories for K∈{1,3,5,7,10} . Text + Rep. Guidance achieves the highest precision and recall across all evaluated selection budgets.
Figure 4: Diversity of trajectories in text and representation space. Within-query action diversity and representation trajectories for four rollouts of one ALFWorld task.
Variant
ALFWorld SR (%)
No Skill
22.64
Rep2Skill (full)
52.24
w/o representation verbalization
47.76 ( ↓ 4.48)
w/o representation analyzer
47.02 ( ↓ 5.22)
w/o group-wise rollout evidence
46.77 ( ↓ 5.47)
Table 3: Ablation study of Rep2Skill using Qwen3-4B on ALFWorld. We report the success rate (SR), averaged over three independent runs. Colored values in parentheses indicate the performance change relative to the full Rep2Skill .
Figure 5: Case Study of Rep2Skill . Attribution localizes the failure to repeated cabinet search, turning a generic reflection into a concrete, reusable skill rule.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
ALFWorld
WebShop
Benchmark
Train / selection / test tasks
39 / 18 / 134
50 / 20 / 100
Max interaction steps per episode
30 / 50
15
History window in actor prompt (turns)
2
5
Selection metric of validation gate
Success rate
Task score
Skill evolution
Appendix
Table 4: Benchmark and model hyperparameters. Values separated by “/” are for Qwen3-4B / Qwen3.5-9B.
Turn
Action disagreement
Unique prefix
Text distance
Rep. distance
1
35.3%
42.1%
0.388
0.040
5
82.7%
88.4%
0.658
0.154
10
84.3%
98.5%
0.728
0.151
15
84.5%
99.7%
0.735
0.130
Appendix
Table 5: Within-group diversity of Qwen3-4B rollouts on ALFWorld. Action disagreement is the fraction of rollout pairs taking different actions at the turn. Unique prefix is the fraction of rollouts whose action prefix up to the turn is unique within the group. Text distance is the TF-IDF cosine distance of turn contents. Rep. distance is the mean pairwise cosine distance of hidden states.
Figure 6: Representation trajectories of four same-task rollouts for six ALFWorld task families (Qwen3-4B, empty skill). Each panel shows the first 15 turns in the top three principal components of that group’s hidden states. Circles and diamonds mark the first and last displayed turns, respectively, and the legend gives each rollout’s final outcome.
Natural-language skills are textual procedural memories through which large language model (LLM) agents retain reusable task knowledge without updating model weights. Existing methods typically treat skills as either static artifacts or monolithic documents optimized using aggregate validation scores as feedback. However, representing a skill as a monolithic document restricts optimization to its textual content, without explicitly modeling the structure through which procedural knowledge is retrieved and executed. We identify a key distinction between learning what knowledge to retain and determining how to organize it: textual updates should first be validated through execution evidence, after which the retained knowledge should be structured according to its procedural dependencies and retrieval requirements. To this end, we introduce SkillSpec, a two-phase framework comprising consensus-gated evolution and representation specialization. In the consensus-gated phase, complementary editing intents generate complete candidate skills. An update is committed only when paired evaluations reach consensus, requiring sufficient overall improvement and non-negative aggregate paired gain in every repeated evaluation. In the specialization phase, signals of process and redundancy sensitivity derived from the full optimization trajectory, including accepted and rejected candidates, guide the selection of a flat, graph, or hybrid representation.Across six benchmarks and three target language models, SkillSpec improves average success rate over SkillOpt by 6.89%, averaged across the three models. These results demonstrate that reliable skill evolution and representation specialization address complementary objectives: deciding what knowledge to retain and how to structure it for inference.
Agent skills are procedural artifacts that enable LLM agents to execute workflows, verify constraints, and recover from failures. Existing self-evolving methods refine skills using accumulated trajectories. However, they struggle in cold-start settings, where only an initial, imperfect skill is available. Consequently, skill construction defaults to expert authoring or one-shot LLM generation. Expert-authored skills are costly and may not align with how LLM agents actually execute tasks, while one-shot generated skills can be syntactically well formed yet behaviorally weak. To bridge this gap, we propose SkillRevise, an execution-grounded framework designed to iteratively refine these initial skills. SkillRevise diagnoses skill defects from execution evidence, retrieves relevant repair principles from a general memory, and applies execution-anchored edits. By re-executing candidates, it retains the first verifier-passing skill within the revision budget and falls back to empirical utility only when no candidate succeeds. Evaluated across three benchmarks and five LLMs, SkillRevise substantially outperforms one-shot baselines, improving the base agent's success rate on SkillsBench from 36.05% to 61.63%. Furthermore, the revised skills transfer across both executors and task environments, suggesting that SkillRevise captures reusable procedural knowledge beyond any single executor.
Yuxuan Liu, Zhaochen Su, Lingyun Xie +11
The Hong Kong University of Science and Technology · Harbin Institute of Technology · Harbin Institute of Technology, Shenzhen +2
Agent skills provide a lightweight way to adapt LLM agents to specialized domains by storing reusable procedural knowledge in structured files. However, whether downloaded from third parties or self-generated, these skills are often unreliable, incomplete, or outdated. Existing skill-evolution methods often address these deficiencies through heuristic reflections without an explicit optimization formulation. In this paper, we propose SkillGrad, a gradient-descent-inspired framework for optimizing agent skills. SkillGrad treats the skill package as a structured parameter to optimize in a gradient descent fashion: task executions provide trajectory-level loss evidence, automatic diagnoses then provide text-based gradients that indicate the correction directions. To stabilize optimization across iterations, a momentum agent accumulates recurring diagnostic patterns into a persistent memory overlay. Finally, an LLM-based patcher executes the parameter update by applying layer-aware edits to the skill package. Evaluated on SpreadsheetBench Verified and WikiTableQuestions, SkillGrad consistently outperforms training-based skill evolution baselines across two backbone LLMs, improving over the strongest training-based baseline by 6.7 percentage points on average. Ablations further show that momentum and contrastive diagnosis both contribute to the final skill quality.
Hanyu Wang, Yifan Lan, Bochuan Cao +2
College of Information Sciences and Technology The Pennsylvania State University University Park, PA, USA