SkillCycle: Co-Evolving Agent Policies and Skill Banks
Organizations: Tsinghua University · Northwestern Polytechnical University
Abstract
Internalizing external skills changes a language agent's capabilities and, with them, the value of its remaining guidance: rules can become redundant, misleading, or insufficient for newly encountered decisions. This creates a coupled problem of learning from skills and adapting the skills that supervise further learning. We introduce SkillCycle, a framework for co-evolving agent policies and skill banks through a feedback loop between skill internalization and rule revision. Our central contribution is to give distillation feedback a second role: token-level contextual differences help locate rules for inspection, while interaction outcomes guide edits to their content and applicability. SkillCycle alternates between two phases: policy learning with a fixed skill bank and router, and rule revision with a frozen policy. Candidate edits undergo rule-level and whole-bank environment comparisons before they guide the next learning cycle. On WebShop, SkillCycle with a 3B model achieves a success rate of 74.74% and a score of 88.37 without inference-time skill inputs, representing relative improvements of 0.73% and 3.96% over the state-of-the-art (SOTA) model, respectively. In Cycle 3 ablations on ALFWorld and WebShop, SkillCycle's no-skill success rates improve by 10.18% and 18.11% relative to a static skill bank, and by 2.41% and 2.50% relative to a single bank update, respectively. These results show that continually revising skill guidance as the agent's capabilities change helps transform external skills into policy capabilities that require no skill inputs at inference. We will release code, configurations, skill banks, and evaluation protocols.
Figures & tables
| ALFWorld | WebShop | ||||||||||
| Method | C | S/T | Pick | Look | Clean | Heat | Cool | Pick2 | Avg. | Score | SR |
| (a) Qwen3-1.7B-Instruct | |||||||||||
| Vanilla ( Yang et al., 2025 ) | – | 25.00 | 22.20 | 3.10 | 0.00 | 21.40 | 4.20 | 12.50 | 46.50 | 4.70 | |
| Skill-Prompt ∗ ( Yang et al., 2026b ) | – | 10.30 | 50.00 | 16.10 | 0.00 | 0.00 | 5.00 | 9.40 | 23.00 | 2.30 | |
| OPSD ( Zhao et al., 2026 ) | – | 26.30 | 33.30 | 9.10 | 0.00 | 4.50 | 5.30 | 14.10 | 47.40 | 9.30 | |
| GRPO ( Shao et al., 2024 ) | – | 71.10 | 41.70 | 36.40 | 40.00 | 31.80 | 31.60 | 46.10 | 67.30 | 38.30 | |
| ALFWorld Macro SR | WebShop SR | |||
| Condition | Student | Teacher | Student | Teacher |
| GiGPO ( Feng et al., 2025 ) | 55.03 [50.03, 60.03] | 61.29 [56.29, 66.29] | 63.54 [58.62, 68.20] | 64.84 [59.94, 69.45] |
| Static | 48.61 [43.61, 53.61] | 51.82 [46.82, 56.82] | 63.28 [58.35, 67.95] | 66.41 [61.54, 70.95] |
| Refresh once | 52.30 [47.30, 57.30] | 58.81 [53.81, 63.81] | 72.92 [68.26, 77.12] | 76.82 [72.35, 80.77] |
| SkillCycle | 53.56 [49.16, 63.81] | 58.28 [53.24, 63.78] | 74.74 [70.16, 78.83] | 76.82 [72.35, 80.77] |
| vs. Static | +4.95 | +6.46 | +11.46 | +10.42 |
| (a) ALFWorld: frozen Qwen3-1.7B | ||||||||
|---|---|---|---|---|---|---|---|---|
| Revision | Pick | Look | Clean | Heat | Cool | Pick2 | Macro SR | |
| 1 | 32.14 | 43.81 | 60.47 | 70.95 | 61.03 | 28.16 | 49.43 | – |
| 2 | 37.84 | 44.71 | 61.94 | 72.38 | 57.77 | 34.06 | 51.45 | +2.02 |
| 3 | 40.12 | 44.93 | 63.25 | 72.10 | 57.99 | 36.28 | 52.45 | +3.02 |
| Operation | Pairs | Gains | Losses | Net | Decision |
|---|---|---|---|---|---|
| Add simple-search recovery | 48 | 1 | 1 | 0 | Reject |
| Rewrite clean readiness | 48 | 0 | 1 | Reject | |
| Rewrite cool-search frontier | 48 | 3 | 6 | Reject | |
| Rewrite heat visible acquisition | 48 | 0 | 1 | Reject | |
| Rewrite heat retention | 48 | 0 | 4 | Reject | |
| Rewrite invalid-action contract | 240 | 4 | 5 | Reject |
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Symbol | Meaning |
|---|---|
| Task instruction and task distribution. | |
| Latent environment state and exposed observation. | |
| Observable interaction history and constructed model context. | |
| Number of training segments and the context constructor. | |
| Generated response and its token at position . | |
| Executed action and response-to-action interface. |
| Setting | WebShop 1.7B | WebShop 3B | ALFWorld 3B |
|---|---|---|---|
| Demo tokens / reduction | Action / | Full response / | Disabled |
| Formal actor updates | 8 | 8 | 8 |
| Learning rate | |||
| Reference coefficient | |||
| Configuration | Auxiliary + demo | Demo | Advantage injection |
| 1 | for do |
|---|---|
| 2 | ; Phase 1: fix |
| 3 | for do |
| 4 | , ; |
| 5 | for each decision |
| 6 | |
| 7 | ; |
| ALFWorld | WebShop | |||||||||
| Condition | Pick | Look | Clean | Heat | Cool | Pick2 | Overall | Macro | Score | SR |
| DeepSeek-V4-Pro | ||||||||||
| No Skill | 86.11 | 50.00 | 69.57 | 58.33 | 21.74 | 57.69 | 60.94 | 57.24 | 7.54 | 5.73 |
| Static General Skill | 88.89 | 62.50 | 69.57 | 36.11 | 63.77 | 56.41 | 67.71 | 62.87 | 11.86 | 8.33 |
| Skill + Router | 86.11 | 75.00 | 69.57 | 50.00 | 47.83 | 57.69 | 66.41 | 64.37 | 15.59 | 11.20 |
| GLM-5.1 | ||||||||||
| ALFWorld ( ) | WebShop ( ) | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Method | Cycle | Pick | Look | Clean | Heat | Cool | Pick2 | Avg. | Score | SR |
| Vanilla ( Yang et al., 2025 ) | – | 25.00 | 22.20 | 3.10 | 0.00 | 21.40 | 4.20 | 12.50 | 46.50 | 4.70 |
| Skill-Prompt ∗ ( Yang et al., 2026b ) | – | 10.30 | 50.00 | 16.10 | 0.00 | 0.00 | 5.00 | 9.40 | 23.00 | 2.30 |
| OPSD ( Zhao et al., 2026 ) | – | 26.30 | 33.30 | 9.10 | 0.00 | 4.50 | 5.30 | 14.10 | 47.40 | 9.30 |
| GRPO ( Shao et al., 2024 ) | – | 71.10 | 41.70 | 36.40 | 40.00 | 31.80 | 31.60 | 46.10 | 67.30 | 38.30 |
| Skill-GRPO ( Yang et al., 2026b ) | – | 27.60 | 54.50 | 22.70 | 27.30 | 0.00 | 19.20 | 21.10 | 73.40 | 46.10 |
| ALFWorld ( ) | WebShop ( ) | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Method | Cycle | Pick | Look | Clean | Heat | Cool | Pick2 | Avg. | Score | SR |
| Vanilla ( Yang et al., 2024 ) | – | 44.40 | 11.10 | 6.20 | 15.40 | 28.60 | 12.50 | 21.90 | 6.70 | 0.80 |
| Skill-Prompt ∗ ( Yang et al., 2026b ) | – | 51.70 | 66.70 | 48.40 | 0.00 | 4.30 | 10.00 | 28.90 | 0.20 | 0.80 |
| OPSD ( Zhao et al., 2026 ) | – | 48.80 | 41.70 | 16.70 | 0.00 | 15.80 | 16.70 | 28.10 | 11.30 | 3.10 |
| GRPO ( Shao et al., 2024 ) | – | 91.20 | 62.50 | 96.20 | 61.90 | 65.00 | 47.40 | 75.00 | 79.80 | 63.30 |
| Skill-GRPO ( Yang et al., 2026b ) | – | 88.90 | 71.40 | 58.80 | 70.60 | 40.70 | 29.20 | 60.20 | 77.30 | 60.90 |
| ALFWorld ( ) | WebShop ( ) | |||||||||
| Method | Cycle | Pick | Look | Clean | Heat | Cool | Pick2 | Avg. | Score | SR |
| Vanilla ( Yang et al., 2024 ) | – | 36.10 | 22.20 | 3.10 | 0.00 | 0.00 | 0.00 | 12.50 | 5.90 | 1.60 |
| Skill-Prompt ∗ ( Yang et al., 2026b ) | – | 51.70 | 50.00 | 32.30 | 5.30 | 4.30 | 0.00 | 23.40 | 1.70 | 0.80 |
| OPSD ( Zhao et al., 2026 ) | – | 50.00 | 60.00 | 22.70 | 21.40 | 17.60 | 9.50 | 32.80 | 4.50 | 2.30 |
| GRPO ( Shao et al., 2024 ) | – | 91.20 | 87.50 | 96.20 | 81.00 | 65.00 | 57.90 | 81.20 | 80.90 | 72.60 |
| Skill-GRPO ( Yang et al., 2026b ) | – | 88.50 | 66.70 | 65.20 | 61.10 | 57.70 | 73.10 | 69.50 | 80.40 | 71.90 |
| Condition | Cycle | Pick | Look | Clean | Heat | Cool | Pick2 | Overall |
|---|---|---|---|---|---|---|---|---|
| GiGPO | 1 | 63.02 [58.07, 67.97] | 53.12 [42.97, 63.02] | 57.03 [52.08, 61.98] | 64.06 [59.11, 69.01] | 33.07 [21.09, 45.05] | 27.08 [15.10, 39.06] | 49.56 [44.56, 54.56] |
| 2 | 65.89 [60.94, 71.09] | 54.95 [45.05, 65.10] | 59.90 [54.95, 65.10] | 66.93 [61.98, 71.88] | 36.98 [25.00, 48.96] | 30.99 [19.01, 42.97] | 52.61 [47.61, 57.61] | |
| 3 | 67.97 [63.02, 72.92] | 57.03 [46.88, 66.93] | 61.98 [57.03, 66.93] | 69.01 [64.06, 73.96] | 40.10 [28.12, 52.08] | 34.11 [21.88, 46.09] | 55.03 [50.03, 60.03] | |
| Static | 1 | 61.46 [56.51, 66.41] | 50.00 [40.10, 59.90] | 53.91 [48.96, 59.11] | 62.50 [57.55, 67.45] | 29.43 [17.45, 41.41] | 23.96 [11.98, 35.94] | 46.88 [41.88, 51.88] |
| 2 | 62.50 [57.55, 67.45] | 50.78 [40.89, 60.68] | 54.69 [49.74, 59.90] | 63.28 [58.07, 68.23] | 30.47 [18.49, 42.45] | 25.00 [13.02, 36.98] | 47.79 [42.79, 52.79] | |
| 3 | 63.02 [58.07, 67.97] | 51.56 [41.41, 61.46] | 55.47 [50.52, 60.42] | 64.06 [59.11, 69.01] | 31.51 [19.53, 43.49] | 26.04 [14.06, 38.02] | 48.61 [43.61, 53.61] |
| Condition | Cycle | Pick | Look | Clean | Heat | Cool | Pick2 | Overall |
|---|---|---|---|---|---|---|---|---|
| GiGPO | 1 | 67.97 [63.02, 72.92] | 58.07 [47.92, 67.97] | 61.98 [57.03, 66.93] | 67.97 [63.02, 72.92] | 40.10 [28.12, 52.08] | 34.11 [21.88, 46.09] | 55.03 [50.03, 60.03] |
| 2 | 71.88 [66.93, 77.08] | 60.94 [51.04, 71.09] | 65.89 [60.94, 71.09] | 71.09 [65.89, 76.04] | 45.05 [33.07, 57.03] | 38.02 [26.04, 50.00] | 58.81 [53.81, 63.81] | |
| 3 | 73.96 [69.01, 78.91] | 63.02 [53.12, 72.92] | 67.97 [63.02, 72.92] | 72.92 [67.97, 78.12] | 47.92 [35.94, 59.90] | 41.93 [29.95, 53.91] | 61.29 [56.29, 66.29] | |
| Static | 1 | 65.10 [59.90, 70.05] | 53.12 [42.97, 63.02] | 57.03 [52.08, 61.98] | 65.10 [59.90, 70.05] | 34.11 [21.88, 46.09] | 28.12 [15.89, 40.10] | 50.43 [45.43, 55.43] |
| 2 | 65.10 [59.90, 70.05] | 53.91 [44.01, 64.06] | 58.07 [53.12, 63.02] | 65.89 [60.94, 71.09] | 34.90 [22.92, 46.88] | 28.91 [16.93, 40.89] | 51.13 [46.13, 56.13] | |
| 3 | 65.62 [60.42, 70.57] | 54.43 [44.53, 64.58] | 58.59 [53.39, 63.54] | 66.41 [61.46, 71.61] | 35.94 [23.96, 47.92] | 29.95 [17.97, 41.93] | 51.82 [46.82, 56.82] |
| Condition | Cycle | Student SR (%) | Teacher SR (%) |
|---|---|---|---|
| GiGPO | 1 | 54.17 [49.17, 59.08] | 66.67 [61.81, 71.20] |
| 2 | 63.02 [58.09, 67.70] | 64.58 [59.68, 69.20] | |
| 3 | 63.54 [58.62, 68.20] | 64.84 [59.94, 69.45] | |
| Static | 1 | 60.42 [55.45, 65.18] | 65.89 [61.01, 70.45] |
| 2 | 66.67 [61.81, 71.20] | 66.67 [61.81, 71.20] | |
| 3 | 63.28 [58.35, 67.95] | 66.41 [61.54, 70.95] |
| Task | Seed | No Skill | Router + Skill |
|---|---|---|---|
| A: clean soapbar, place in cabinet | 1 | F / 30 | S / 10 |
| 2 | F / 30 | S / 11 | |
| 3 | F / 30 | S / 11 | |
| B: cool cup, place in microwave | 1 | F / 30 | S / 8 |
| 2 | F / 30 | S / 8 | |
| 3 | S / 13 | S / 15 |