Large language model agents can improve across tasks by retaining reusable skills distilled from prior interactions. Recent work jointly optimizes task execution and skill extraction, enabling the policy and skillbank to co-evolve. However, as the actor continues learning, rewarding skill proposals through their reuse in subsequent training steps may conflate skill benefits with actor improvement, while directly testing each proposed skill requires costly additional actor rollouts. In this paper, we introduce UniSkill, which uses a shared policy to interact with the environment and propose skillbank edits (Add, Update, or No Edit) from the resulting trajectories. Specifically, the actor learns from environment rewards, while contrastive action feedback guides skill proposal learning. This feedback provides an actor-alignment signal by measuring how replacing the retrieved skill with a proposed skill changes the current actor's action log-likelihood gap between previously collected successful and failed trajectories from the same task, thereby avoiding new rollouts for each proposal. Since proposal-level feedback may suppress an otherwise appropriate edit operation when the proposed skill content scores poorly, we further apply skill-edit support regularization to preserve exploration. Empirically, UniSkill achieves strong performance, reaching 98.4% success on ALFWorld and 84.7% on WebShop while maintaining stable joint training. Further ALFWorld experiments show that UniSkill remains effective when the shared policy uses a smaller backbone. Our implementation is available at https://github.com/LimOkii/UniSKill.
Figures & tables
Figure 1: ALFWorld training dynamics. (a) UniSkill continues improving later in training, whereas Evolving-RL collapses after strong early gains. (b) The learned skillbank grows rapidly early and continues evolving through occasional additions and updates (non-overlapping 10-step means).
Figure 2: Overview of UniSkill. (1–2) The shared policy interacts with the environment and proposes edits from the resulting trajectories, while a skill critic checks edit appropriateness and content support. (3) Contrastive action feedback measures actor alignment on fixed trajectories. (4–5) Actor and proposer objectives jointly update the policy. (6) Eligible edits update the skillbank.
Figure 3
Setting
Training
ALFWorld Success Rate (%)
Actor
Skill Proposer
Ralign reward
w/o Retrieval
w/ Retrieval
Actor Only
Updated
Frozen
×
84.4
87.5
Skill Proposer Only
Frozen
Updated
✓
34.4
32.8
UniSkill (w/o Ralign reward)
Updated
Updated
×
84.4
89.1
UniSkill
Updated
Updated
✓
96.1
98.4
Table 2: Ablation results on ALFWorld with and without skill retrieval at evaluation.
Figure 5: Actor alignment on ALFWorld with linear fits and pointwise 95% bootstrap CIs.
Figure 6: Support regularization on ALFWorld. (a–b) Skill-edit distributions (one run per setting; 10-step windows). (c) Validation success (mean ± std; three independent runs per setting).
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter
ALFWorld
WebShop
Optimization
Optimizer
AdamW
AdamW
AdamW (β1,β2)
(0.9,0.999)
(0.9,0.999)
Learning rate (constant)
1×10−6
1×10−6
Weight decay
0.01
0.01
Gradient norm clipping
1.0
1.0
Appendix
Table 3: Training and generation settings.
Percentiles of ∣Ralign∣
Score proportions (%)
50th
90th
95th
Ralign<−η
∣Ralign∣≤η
Ralign>η
0.0060
0.0358
0.0615
3.0
93.3
3.7
Appendix
Table 4: Unclipped Ralign statistics from a preliminary ALFWorld run ( n=566 ).
Setting
Pick
Look
Clean
Heat
Cool
Pick2
All
w/ Retrieval
100.0 ±0.0
100.0 ±0.0
94.1 ±5.9
100.0 ±0.0
100.0 ±0.0
98.2 ±3.0
98.2 ±1.2
w/o Retrieval
100.0 ±0.0
92.9 ±0.0
92.2 ±3.4
100.0 ±0.0
98.0 ±3.4
98.2 ±3.0
96.6 ±0.9
Appendix
Table 5: UniSkill’s ALFWorld out-of-distribution success rates (%), reported as mean ± sample std over three independent training runs.
Skill Critic
Task Score
Success (%)
DeepSeek-V4-Pro
90.5 ±1.4
84.7 ±0.5
DeepSeek-V4-Flash
88.9
83.6
Qwen3.5-35B-A3B
90.4
85.2
Appendix
Table 6: WebShop performance with different skill critics.
Figure 7: WebShop validation success with different skill critics.
Figure 8: Skill-edit distributions for two additional independent runs without support regularization on ALFWorld. Shares are computed from recognized proposals pooled over non-overlapping 10-step windows.
Edit
Target
Interaction evidence
Key skill change
Ralign
Add
Bowl
No prior skill; the lamp was found on a dresser.
Add the dresser as a lamp-search location.
0.09850
Update
Pen
Dresser-only guidance missed a lamp located on a shelf.
Add shelves alongside the dresser.
0.01724
Update
Bowl
The lamp was found on a side table, which the current skill omitted.
Add side tables alongside dressers and shelves.
0.11585
Appendix
Table 7: Evolution of one ALFWorld lamp-search skill. Each row is a critic-accepted edit committed to the skillbank; the last column reports raw alignment scores.
Task
Δ+
Δ−
Ralign
Success
Admit
Clean kettle → stove burner
−0.00449
−0.01515
0.01067
15/32 → 8/32
No
Two kettles → one cabinet
0.00087
−0.02682
0.02769
30/32 → 24/32
Yes
Two remotes → armchair
0.00280
−0.01030
0.01311
16/32 → 6/32
Yes
Appendix
Table 8: Representative cases with positive alignment and negative fixed-actor rollout gain.