Textual skills enable large language model (LLM) based agents to accumulate reusable procedural knowledge without updating model parameters. Yet existing skill evolution remains largely confined to the text space: an optimizer must diagnose success and failure patterns, and revise skills solely from long execution trajectories and sparse task outcomes. This text-only paradigm leaves the agent's internal representations, which contain rich records of its evolving execution state, outside the skill optimization loop. We ask whether an agent can improve its external textual skills by reflecting on its own internal representations. We introduce Rep2Skill, a representation-guided framework for self-evolution on agent skills. Specifically, upon the collected agent rollouts, Rep2Skill models their internal model representation trajectories to localize turns that deviate from successful execution dynamics, and it further interprets these signals alongside the execution contexts as actionable textual feedback for targeted skill revision. Experiments on two agent environments with two open-source LLMs show that Rep2Skill consistently outperforms text-only approaches in the self-evolution setting, where the same LLM serves as both executor and optimizer without a stronger external model. This establishes a promising direction moving agent self-improvement beyond text-only reflection.
Figures & tables
Figure 1: Motivation of Rep2Skill . Top: Rep2Skill introduces internal representation signals into the iterative skill optimization loop to provide fine-grained guidance. Bottom: standard skill evolution relies on textual trajectories alone, while Rep2Skill leverages representation-guided evidence from grouped rollouts to support more targeted skill updates.
Figure 2: Overview of Rep2Skill . Given grouped rollouts under the current skill sn , a trajectory model trained on successful executions scores each turn by its representation prediction error, and the top- Kc critical turns are verbalized into textual feedback F(k) . The optimizer then revises sn with this feedback through a validation-gated update. The same frozen LLM πθ serves as executor, analyzer and optimizer.
Model
Method
ALFWorld
WebShop
Succ (%) ↑
Score ↑
Succ (%) ↑
Qwen3-4B
No Skill
22.64 ± 1.76
41.86 ± 1.65
13.67 ± 1.53
Trace2Skill ( Ni et al., 2026 )
47.26 ± 5.18
45.68 ± 3.40
13.00 ± 1.41
EvoSkill ( Alzubi et al., 2026 )
43.03 ± 3.52
49.48 ± 2.02
13.33 ± 0.94
SkillOpt ( Yang et al., 2026a )
46.51 ± 3.12
39.04 ± 2.49
11.67 ± 3.40
Rep2Skill (Ours)
52.24 ± 3.22
49.79 ± 5.95
16.33 ± 3.51
Table 1: Self-evolving skill performance on ALFWorld and WebShop. All results are mean ± standard deviation over 3 independent evolution runs.
Input
AUROC ↑
AUPRC ↑
P@Top- 3↑
Trajectory Text
0.494
0.736
63.66%
Rep.
0.592
0.789
79.26%
Text + Rep. ( Ours )
0.838
0.876
88.22%
Table 2: Turn-level error detection under different attribution signals. Representation-space signals provide complementary information to trajectory text. Both the execution model and analyzer model are Qwen3-4B ( Yang et al., 2025 ) .
Figure 3: Turn-level non-progress detection with different attribution inputs. Precision@K and Recall@K are macro-averaged over eligible trajectories for K∈{1,3,5,7,10} . Text + Rep. Guidance achieves the highest precision and recall across all evaluated selection budgets.
Figure 4: Diversity of trajectories in text and representation space. Within-query action diversity and representation trajectories for four rollouts of one ALFWorld task.
Variant
ALFWorld SR (%)
No Skill
22.64
Rep2Skill (full)
52.24
w/o representation verbalization
47.76 ( ↓ 4.48)
w/o representation analyzer
47.02 ( ↓ 5.22)
w/o group-wise rollout evidence
46.77 ( ↓ 5.47)
Table 3: Ablation study of Rep2Skill using Qwen3-4B on ALFWorld. We report the success rate (SR), averaged over three independent runs. Colored values in parentheses indicate the performance change relative to the full Rep2Skill .
Figure 5: Case Study of Rep2Skill . Attribution localizes the failure to repeated cabinet search, turning a generic reflection into a concrete, reusable skill rule.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
ALFWorld
WebShop
Benchmark
Train / selection / test tasks
39 / 18 / 134
50 / 20 / 100
Max interaction steps per episode
30 / 50
15
History window in actor prompt (turns)
2
5
Selection metric of validation gate
Success rate
Task score
Skill evolution
Appendix
Table 4: Benchmark and model hyperparameters. Values separated by “/” are for Qwen3-4B / Qwen3.5-9B.
Turn
Action disagreement
Unique prefix
Text distance
Rep. distance
1
35.3%
42.1%
0.388
0.040
5
82.7%
88.4%
0.658
0.154
10
84.3%
98.5%
0.728
0.151
15
84.5%
99.7%
0.735
0.130
Appendix
Table 5: Within-group diversity of Qwen3-4B rollouts on ALFWorld. Action disagreement is the fraction of rollout pairs taking different actions at the turn. Unique prefix is the fraction of rollouts whose action prefix up to the turn is unique within the group. Text distance is the TF-IDF cosine distance of turn contents. Rep. distance is the mean pairwise cosine distance of hidden states.
Figure 6: Representation trajectories of four same-task rollouts for six ALFWorld task families (Qwen3-4B, empty skill). Each panel shows the first 15 turns in the top three principal components of that group’s hidden states. Circles and diamonds mark the first and last displayed turns, respectively, and the legend gives each rollout’s final outcome.