SkillGym: Training Skill-Use Agents with Automatic Verifiable Environment Generation
Organizations: Mohamed bin Zayed University of Artificial Intelligence
Abstract
Skills equip LLM agents with professional knowledge and guidance to complete long-horizon and complex tasks. Although skills have been widely adopted in recent agent paradigms and harnesses, how to synthesize reliable training data and how to train agents for skill use remain underexplored. In this work, we propose SkillGym, an automatic pipeline to build verifiable environments, collect trajectories, and train skill-use agents. SkillGym first crawls a large volume of skills from the internet, then keeps those whose workflows can run reproducibly offline. A builder-reviewer pipeline is used to construct difficulty-controlled tasks, spanning four task types, each with a reference solution and an executable verifier. With this pipeline, we build 6.8k environments and collect 19k verified successful trajectories for supervised finetuning. Finetuning on these trajectories improves LLMs of different families and sizes, from 2B to 122B parameters across four skill-use benchmarks; Our Qwen3.5-9B SFT model outperforms the 397B untrained model on two of them. Further analysis shows that training teaches agents to invoke skills, raising the rate of reading the relevant skill from 28% to 96%, and that the gains hold across reasoning structures, extending to task types that form a minority of the training data and to skills held out from training
Figures & tables
| Model | Size | SkillGym | SkillEval | SkillsBench | Skill-Use-Bench |
|---|---|---|---|---|---|
| Qwen3.5 | 397B/17B | ||||
| GLM-5.2 ⋄ | 753B | ||||
| DeepSeek-V4-Flash ⋄ | 284B/13B | ||||
| Kimi-K3 ⋄ | 2.8T/104B | ||||
| Base and SkillGym SFT comparisons | |||||
| MiniCPM5 | 2B | ||||
| Ours | SkillsBench | SkillEval | Skill-Use-Bench | |||
|---|---|---|---|---|---|---|
| Training data | Pass | Strict | SU | Compl. | ||
| Qwen3.5-9B (no SFT) | 41.3 | 14.8 | 65.7 | 17.3 | 14.8 | 51.4 |
| SkillGym (full) | 59.5 | 22.4 | 74.7 | 33.3 | 49.6 | 46.9 |
| Validated only | 40.5 | 10.8 | 41.2 | 18.3 | 45.1 | 43.1 |
| Review-approved | 45.2 | 12.0 | 47.9 | 24.7 | 47.4 | 52.4 |
| Procedural only | 42.5 | 17.4 | 56.0 | 24.7 | 52.2 | 56.2 |
| Task type | w/o | w/ | Gain |
|---|---|---|---|
| Procedural | 31.0 | 47.0 | +16.0 |
| Constraint sat. | 79.8 | 89.9 | +10.1 |
| Abductive | 65.7 | 69.7 | +4.0 |
| Partial order | 65.7 | 68.7 | +3.0 |
| All | 60.5 | 68.8 | +8.3 |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Level | Field | Values (%) | Selection |
| (i) Screening ( ) | Label | real skill 88.9, unclear 5.5, meta/doc 3.3, template 1.0, test/demo 0.7, placeholder 0.5 | real skill |
| Language | English 92.4, Chinese 3.0, Japanese 1.8, Korean 0.9, mixed 0.9, other 1.0 | English | |
| Repository ‡ | archived flag, bulk-publisher flag, license, stars | not archived or bulk | |
| Size ‡ | number of files in the skill folder | 300 | |
| (ii) Identity ( ) | Domain | 21 domains; software engineering 53.0, agents/meta 8.8, productivity 3.9, infrastructure 3.9, writing 3.8, media 3.7, … | – |
| Task type † | generate 51.7, validate 48.4, analyze 46.8, integrate 30.7, plan 26.7, orchestrate 25.2, transform 18.4, … | – |
| Model | Size | SkillGym | SkillEval | SkillsBench | Skill-Use-Bench |
|---|---|---|---|---|---|
| Qwen3.5 | 397B/17B | ||||
| Qwen3.8 | 27B | ||||
| GLM-5.2 ⋄ | 753B | ||||
| GLM-5.3 | 753B | ||||
| DeepSeek-V4-Pro | 1.6T/49B | ||||
| DeepSeek-V4-Flash ⋄ | 284B/13B |
| Stage | skills.sh | Registry | Total |
|---|---|---|---|
| Crawled (manifest) | 9,724 | 174,706 | 184,430 |
| Fetched and deduplicated | 8,593 | 42,538 | 51,131 |
| Screened as a real skill | 8,104 | 37,333 | 45,437 |
| Annotated (not archived or bulk-published) | 6,776 | 20,793 | 27,569 |
| Selected (rules in Table 4 ) | 2,838 | 9,268 | 12,106 |
| After cross-source deduplication | 2,838 | 9,059 | 11,897 |
| Domain | Selected skills | Skills with tasks | Tasks |
|---|---|---|---|
| Marketing | 406 | 374 | 798 |
| Software engineering | 7,726 | 707 | 791 |
| Productivity | 321 | 285 | 618 |
| Cybersecurity | 341 | 289 | 602 |
| ML / AI tooling | 286 | 263 | 598 |
| Infrastructure | 317 | 283 | 594 |
| Profile | Passed | Unresolved | Total |
|---|---|---|---|
| Procedural execution | 1,293 | 869 | 2,162 |
| Abductive diagnosis | 2,400 | 430 | 2,830 |
| Constraint satisfaction | 1,498 | 282 | 1,780 |
| Total | 5,191 | 1,581 | 6,772 |
| Size | Median | Mean | P10–P90 |
| Instruction length (words) | 336 | 365.4 | 215–556 |
| Successful | Unsuccessful | ||
| Trajectories | 19,070 | 15,649 | |
| Harness | MiniSwe-Agent | 6,433 | 4,360 |
| AgentFly | 5,936 | 6,909 | |
| Terminus-2 | 6,701 | 4,380 | |
| Teacher | Kimi-K3 | 10,603 | 7,649 |
| DeepSeek-V4-Flash | 6,521 | 5,507 | |
| Term | Coefficient | 95% CI |
|---|---|---|
| Sequential procedure | ||
| Dependency ordering | ||
| Diagnosis and repair | ||
| Interacting constraints | ||
| Rule application | ||
| SkillEval |
| Stage | Priced as | Cached in | New in | Output | Cost (USD) |
|---|---|---|---|---|---|
| Task construction, measured | V4-Pro | 58.4 | 0.54 | 1.00 | 3.6k |
| Task construction, all attempts | V4-Pro | extrapolated | up to 6.3k | ||
| Quality review, assumed | V4-Pro | 0.28 | 0.39 | 0.84 | 1.9k |
| Trajectory collection | V4-Flash | 6.44 | 0.48 | 0.08 | 0.14k |
| Total | 5.7k–8.4k | ||||