Large language model agents can improve across tasks by retaining reusable skills distilled from prior interactions. Recent work jointly optimizes task execution and skill extraction, enabling the policy and skillbank to co-evolve. However, as the actor continues learning, rewarding skill proposals through their reuse in subsequent training steps may conflate skill benefits with actor improvement, while directly testing each proposed skill requires costly additional actor rollouts. In this paper, we introduce UniSkill, which uses a shared policy to interact with the environment and propose skillbank edits (Add, Update, or No Edit) from the resulting trajectories. Specifically, the actor learns from environment rewards, while contrastive action feedback guides skill proposal learning. This feedback provides an actor-alignment signal by measuring how replacing the retrieved skill with a proposed skill changes the current actor's action log-likelihood gap between previously collected successful and failed trajectories from the same task, thereby avoiding new rollouts for each proposal. Since proposal-level feedback may suppress an otherwise appropriate edit operation when the proposed skill content scores poorly, we further apply skill-edit support regularization to preserve exploration. Empirically, UniSkill achieves strong performance, reaching 98.4% success on ALFWorld and 84.7% on WebShop while maintaining stable joint training. Further ALFWorld experiments show that UniSkill remains effective when the shared policy uses a smaller backbone. Our implementation is available at https://github.com/LimOkii/UniSKill.
Figures & tables
Figure 1: ALFWorld training dynamics. (a) UniSkill continues improving later in training, whereas Evolving-RL collapses after strong early gains. (b) The learned skillbank grows rapidly early and continues evolving through occasional additions and updates (non-overlapping 10-step means).
Figure 2: Overview of UniSkill. (1–2) The shared policy interacts with the environment and proposes edits from the resulting trajectories, while a skill critic checks edit appropriateness and content support. (3) Contrastive action feedback measures actor alignment on fixed trajectories. (4–5) Actor and proposer objectives jointly update the policy. (6) Eligible edits update the skillbank.
Figure 3
Setting
Training
ALFWorld Success Rate (%)
Actor
Skill Proposer
Ralign reward
w/o Retrieval
w/ Retrieval
Actor Only
Updated
Frozen
×
84.4
87.5
Skill Proposer Only
Frozen
Updated
✓
34.4
32.8
UniSkill (w/o Ralign reward)
Updated
Updated
×
84.4
89.1
UniSkill
Updated
Updated
✓
96.1
98.4
Table 2: Ablation results on ALFWorld with and without skill retrieval at evaluation.
Figure 5: Actor alignment on ALFWorld with linear fits and pointwise 95% bootstrap CIs.
Figure 6: Support regularization on ALFWorld. (a–b) Skill-edit distributions (one run per setting; 10-step windows). (c) Validation success (mean ± std; three independent runs per setting).
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter
ALFWorld
WebShop
Optimization
Optimizer
AdamW
AdamW
AdamW (β1,β2)
(0.9,0.999)
(0.9,0.999)
Learning rate (constant)
1×10−6
1×10−6
Weight decay
0.01
0.01
Gradient norm clipping
1.0
1.0
Appendix
Table 3: Training and generation settings.
Percentiles of ∣Ralign∣
Score proportions (%)
50th
90th
95th
Ralign<−η
∣Ralign∣≤η
Ralign>η
0.0060
0.0358
0.0615
3.0
93.3
3.7
Appendix
Table 4: Unclipped Ralign statistics from a preliminary ALFWorld run ( n=566 ).
Setting
Pick
Look
Clean
Heat
Cool
Pick2
All
w/ Retrieval
100.0 ±0.0
100.0 ±0.0
94.1 ±5.9
100.0 ±0.0
100.0 ±0.0
98.2 ±3.0
98.2 ±1.2
w/o Retrieval
100.0 ±0.0
92.9 ±0.0
92.2 ±3.4
100.0 ±0.0
98.0 ±3.4
98.2 ±3.0
96.6 ±0.9
Appendix
Table 5: UniSkill’s ALFWorld out-of-distribution success rates (%), reported as mean ± sample std over three independent training runs.
Skill Critic
Task Score
Success (%)
DeepSeek-V4-Pro
90.5 ±1.4
84.7 ±0.5
DeepSeek-V4-Flash
88.9
83.6
Qwen3.5-35B-A3B
90.4
85.2
Appendix
Table 6: WebShop performance with different skill critics.
Figure 7: WebShop validation success with different skill critics.
Figure 8: Skill-edit distributions for two additional independent runs without support regularization on ALFWorld. Shares are computed from recognized proposals pooled over non-overlapping 10-step windows.
Edit
Target
Interaction evidence
Key skill change
Ralign
Add
Bowl
No prior skill; the lamp was found on a dresser.
Add the dresser as a lamp-search location.
0.09850
Update
Pen
Dresser-only guidance missed a lamp located on a shelf.
Add shelves alongside the dresser.
0.01724
Update
Bowl
The lamp was found on a side table, which the current skill omitted.
Add side tables alongside dressers and shelves.
0.11585
Appendix
Table 7: Evolution of one ALFWorld lamp-search skill. Each row is a critic-accepted edit committed to the skillbank; the last column reports raw alignment scores.
Task
Δ+
Δ−
Ralign
Success
Admit
Clean kettle → stove burner
−0.00449
−0.01515
0.01067
15/32 → 8/32
No
Two kettles → one cabinet
0.00087
−0.02682
0.02769
30/32 → 24/32
Yes
Two remotes → armchair
0.00280
−0.01030
0.01311
16/32 → 6/32
Yes
Appendix
Table 8: Representative cases with positive alignment and negative fixed-actor rollout gain.
A persistent skill library allows language model agents to reuse successful strategies across tasks. Maintaining such a library requires three coupled capabilities. The agent selects a relevant skill, utilizes it during execution, and distills new skills from experience. Existing methods optimize these capabilities in isolation or with separate reward sources, resulting in partial and conflicting evolution. We propose Skill1, a framework that trains a single policy to co-evolve skill selection, utilization, and distillation toward a shared task-outcome objective. The policy generates a query to search the skill library, re-ranks candidates to select one, solves the task conditioned on it, and distills a new skill from the trajectory. All learning derives from a single task-outcome signal. Its low-frequency trend credits selection and its high-frequency variation credits distillation. Experiments on ALFWorld and WebShop show that Skill1 outperforms prior skill-based and reinforcement learning baselines. Training dynamics confirm the co-evolution of the three capabilities, and ablations show that removing any credit signal degrades the evolution.
Yaorui Shi, Yuxin Chen, Zhengxi Lu +6
University of Science and Technology of China · 2Meituan · 3National University of Singapore +2
Large Language Model (LLM) agents have shown stunning results in complex tasks, yet they often operate in isolation, failing to learn from past experiences. Existing memory-based methods primarily store raw trajectories, which are often redundant and noise-heavy. This prevents agents from extracting high-level, reusable behavioral patterns that are essential for generalization. In this paper, we propose SkillRL, a framework that bridges the gap between raw experience and policy improvement through automatic skill discovery and recursive evolution. Our approach introduces an experience-based distillation mechanism to build a hierarchical skill library SkillBank, an adaptive retrieval strategy for general and task-specific heuristics, and a recursive evolution mechanism that allows the skill library to co-evolve with the agent's policy during reinforcement learning. These innovations significantly reduce the token footprint while enhancing reasoning utility. Experimental results on ALFWorld, WebShop and seven search-augmented tasks demonstrate that SkillRL achieves state-of-the-art performance, outperforming strong baselines over 15.3% and maintaining robustness as task complexity increases. Code is available at this https://github.com/aiming-lab/SkillRL.
Peng Xia, Jianwen Chen, Hanyang Wang +10
UNC-Chapel Hill · University of Chicago · University of California San Diego +3
Agentic reinforcement learning (RL) enables LLM agents to improve continuously from environment rewards, yet the resulting policies do not systematically accumulate reusable strategies that generalize across tasks. Modular skills can provide such reusable strategies, yet existing skill-augmented RL methods decouple skill creation from policy optimization, risking adopting skills that conflict with the evolving policy. Inspired by Anthropic's Skill Creator, we introduce ReSkill, an RL-in-the-loop skill creation framework that reconciles skill evolution with policy learning. ReSkill exploits the group-wise structure of GRPO to naturally embed three mechanisms with only marginal additional overhead: (1) an assertion-driven skill creator that diagnoses failures from past experience and proposes conditional, trigger-based skill revisions; (2) within-group rollout sampling that enables controlled comparison of skill versions, capturing which version best supports the policy's ongoing learning; and (3) Thompson Sampling with adaptive discounting to balance exploration and exploitation in skill version selection as the policy evolves. Across several domains, ReSkill consistently outperforms existing memory and skill-based RL methods, with the largest gains on unseen tasks. Analysis of the skill lifecycle shows skills being automatically created, tested, refined, and pruned as the policy improves, demonstrating reconciled skill-policy co-evolution.
Zelin He, Haotian Lin, Boran Han +6
The Pennsylvania State University · † Work done during an internship at Amazon · Amazon IntelliHub +1