cs.AISep 29, 2026

SkillGym: Training Skill-Use Agents with Automatic Verifiable Environment Generation

Authors: Renxi Wang, Mingshan Hee, Fajri Koto, Timothy Baldwin, Haonan Li

Organizations: Mohamed bin Zayed University of Artificial Intelligence

Abstract

Skills equip LLM agents with professional knowledge and guidance to complete long-horizon and complex tasks. Although skills have been widely adopted in recent agent paradigms and harnesses, how to synthesize reliable training data and how to train agents for skill use remain underexplored. In this work, we propose SkillGym, an automatic pipeline to build verifiable environments, collect trajectories, and train skill-use agents. SkillGym first crawls a large volume of skills from the internet, then keeps those whose workflows can run reproducibly offline. A builder-reviewer pipeline is used to construct difficulty-controlled tasks, spanning four task types, each with a reference solution and an executable verifier. With this pipeline, we build 6.8k environments and collect 19k verified successful trajectories for supervised finetuning. Finetuning on these trajectories improves LLMs of different families and sizes, from 2B to 122B parameters across four skill-use benchmarks; Our Qwen3.5-9B SFT model outperforms the 397B untrained model on two of them. Further analysis shows that training teaches agents to invoke skills, raising the rate of reading the relevant skill from 28% to 96%, and that the gains hold across reasoning structures, extending to task types that form a minority of the training data and to skills held out from training

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. SkillGym: Internalizing Human Skills into LLMs for Real-World Problem Solving

    Sep 23, 2026Zhilong Ge, Yuting Shao, Yutao Yang +6Reusable Agent SkillsSkills

  2. SkillGen: Verified Inference-Time Agent Skill Synthesis

    May 9, 2026Yuchen Ma, Yue Huang, Han Bao +5Reusable Agent SkillsSkills

  3. SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation

    Aug 3, 2026Zelin Tan, Yiqun Zhang, Hao Li +11Synthetic TaskLanguage-Model Agents