Executable environments are critical for post-training agents on tasks that require tool use and multi-step interaction, but constructing executable tasks together with their environments remains difficult to scale. Skills provide reusable domain knowledge, operational procedures, and tool-use instructions, but a substantial gap remains between the information contained in a skill and a concrete, challenging task with a complete executable environment. To address this gap, we introduce Skill2Env, a capability-oriented framework that starts from a skill and uses agent capability demands to guide task and environment synthesis. Skill2Env represents these demands through reusable difficulty patterns and instantiates them into task blueprints that specify objectives, challenges, environment facts, information boundaries, and acceptance criteria. These blueprints guide the joint construction of task instructions, execution substrates, workspaces, and rubric-based evaluators around source skills. We further propose Iterative Task Hardening, which uses solver execution evidence to identify insufficiently challenging task designs, strengthen or extend their difficulty-pattern instantiations, and revise the corresponding blueprints and environments. Using 1.5K high-scoring trajectories generated from Skill2Env environments for supervised fine-tuning, we observe consistent improvements across a broad range of agent benchmarks, demonstrating the effectiveness of capability-oriented environment synthesis for agent post-training.
Figures & tables
Figure 1: Skill2Env performance on Terminal-Bench 2.1 and SkillsBench.
Figure 2: Overview of Skill2Env. Curated skills and capability demands are connected through difficulty patterns and organized into task blueprints, which guide the construction of executable tasks and their environments. Solver rollouts provide rubric results, and behavioral evidence for environment validation and task diagnosis. This feedback further supports difficulty-pattern expansion and Iterative Task Hardening through coordinated updates to task blueprints and environments.
Model
Terminal- Bench 2.1
SWE-bench Multilingual
Skills Bench
Claw- Eval
τ3 - Banking
Automation Bench
VitaBench
Avg.
Frontier Closed-Source Models
GPT-5.4
78.3
71.7
51.7
60.3
28.5
27.7
47.1
52.2
Claude Opus 4.6
71.2
77.8
50.2
70.4
20.3
25.5
38.3
50.5
Gemini-3.1 Pro
73.8
44.0
60.8
57.8
23.7
28.2
52.7
48.7
Open-Weight Models
DeepSeek-V4-Flash-0731
78.7
76.0
53.8
49.3
30.3
36.3
56.3
54.4
Table 1: Main results on seven agent benchmarks.
Figure 3: SkillsBench improvements after SFT under OpenHands. (a) Overall scores before and after fine-tuning; blue annotations show absolute gains. (b) Absolute score gains across eight domains; parentheses give the number of tasks. Scores are mean task reward ×100 ; gains are differences in percentage points. Skills off* disables automatic loading while leaving skill files accessible.
Figure 4: Terminal-Bench 2.1 Performance. Blue annotations show gains (percentage points).
Figure 5: Task hardening and downstream training utility. (a) DeepSeek-V4-Flash full-credit rate ( r=1 , left axis) and mean recorded assistant turns (right axis) over 500 tasks. (b–c) Assistant-turn distributions for 500 tasks at all three iterations, comparing iterations 0–1 and 1–2, respectively. (d) Seven-benchmark mean after SFT on 500 trajectories per iteration.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Capability Demand
Agent Requirements
Environment Understanding
Identify task-relevant facts from the information provided by the environment and understand the task objectives and constraints.
Planning
Identify dependencies between operations and continually adapt the execution plan based on evolving environment states.
Skill Usage
Understand and appropriately apply the tool-use instructions and execution procedures provided by the skill according to the current task requirements.
Long-Horizon Consistency
Continuously track state changes during multi-step execution and maintain consistency between intermediate results and the environment state.
Error Recovery
Locate errors using execution feedback, adjust operations, and resume task execution.
Appendix
Table 2: Capability demands used to guide environment synthesis.
Pattern
Construction and hardening control
U01 Target discrimination
Provide near-identical files or records; relational clues uniquely identify the requested target. Control: candidate similarity.
U02 Distributed evidence
Split complementary facts across known sources so no single source supports the complete answer. Control: source count.
U03 Nested discovery
Place accessible evidence inside nested archives, linked records, or embedded objects. Control: access depth.
U04 Scope enumeration
Require all qualifying objects across partitions or pages, with a discoverable completeness criterion. Control: partition count.
U05 Sparse-signal retrieval
Embed a small set of relevant facts among topically similar but irrelevant passages. Control: distractor ratio.
U06 Authority resolution
Provide conflicting claims with an explicit, accessible hierarchy of authoritative sources. Control: authority tiers.
Appendix
Table 3: Difficulty patterns for environment understanding.
Pattern
Construction and hardening control
P01 State dependencies
Make later actions require specific earlier outputs or state transitions, forming a nontrivial dependency graph. Control: dependency depth.
P02 Bootstrap dependencies
Introduce an apparent dependency cycle with a documented bootstrap action that makes the workflow feasible. Control: cycle size.
P03 Conditional workflows
Make the correct continuation depend on an earlier result, requiring an explicit contingent plan. Control: branch depth.
P04 Subgoal decomposition
Specify one coherent objective that must be decomposed into intermediate milestones and actionable work units. Control: milestone count.
P05 Probe scheduling
Offer several informative probes with different costs; require choosing which uncertainty to resolve next. Control: probe alternatives.
P06 Coupled constraints
Make individually feasible requirements share decision variables, so they must be solved jointly. Control: coupling density.
Appendix
Table 4: Difficulty patterns for planning.
Pattern
Construction and hardening control
S01 Capability routing
Provide several documented skill entry points; require selecting the one implementing an already chosen operation. Control: entry-point similarity.
S02 Applicability checks
Make a skill procedure valid only under stated input assumptions that must be checked before applying it. Control: precondition count.
S03 Procedural exceptions
Provide a default procedure with explicit exceptions; some inputs require the exception-specific local handling. Control: exception density.
S04 Example adaptation
Supply a worked skill example whose incidental constants differ from the current case; require transferring the procedure correctly. Control: example mismatch.
S05 Version-specific behavior
Fix a known installed version whose documented operation differs from another version or common example. Control: version divergence.
S06 Parameter binding
Require exact binding of arguments, flags, or option values for a selected operation whose meaning is already known. Control: parameter coupling.
Appendix
Table 5: Difficulty patterns for skill usage.
Pattern
Construction and hardening control
L01 Identity continuity
Rename or relocate an object during a valid workflow while later steps must continue tracking the same logical object. Control: identity transitions.
L02 Revision coherence
Produce several valid revisions; downstream steps must consistently consume the designated current revision. Control: revision distance.
L03 Derivation lineage
Require intermediate and final artifacts to retain traceable links to the exact source records used to generate them. Control: lineage depth.
L04 Assumption continuity
Choose a convention or mapping early in the workflow and require later steps to preserve that choice. Control: retention span.
L05 Configuration continuity
Run valid steps across sessions or tools whose defaults differ; maintain the intended shared configuration throughout. Control: context switches.
L06 Reference integrity
Update linked objects while preserving valid cross-references, keys, paths, or foreign-key relationships. Control: reference fan-out.
Appendix
Table 6: Difficulty patterns for long-horizon consistency.
Pattern
Construction and hardening control
R01 Failure recognition
Return a success-like status despite a violated postcondition; expose independent evidence enabling detection of the deviation. Control: signal discrepancy.
R02 Fault localization
Surface an error downstream from its cause; provide logs or checks that identify the faulty component or step. Control: propagation distance.
R03 Cause discrimination
Make several causes produce the same symptom within a localized component; provide probes distinguishing the actual cause. Control: competing causes.
R04 Persistence diagnosis
Provide observable evidence distinguishing a transient interruption from a persistent defect requiring an altered action. Control: diagnostic delay.
R05 Diagnostic preservation
Make inspection or repair alter ephemeral recovery evidence; require preserving a faithful copy before intervening. Control: evidence lifetime.
R06 Input repair
Supply recoverably malformed input with sufficient repair evidence; require restoring validity without inventing missing semantics. Control: defect extent.
Appendix
Table 7: Difficulty patterns for error recovery.
Figure 6: Skill-resource diversity across 2,963 task-associated source-skill packages. (a) Resource types after removing catalog auxiliary files. (b) All 13 source-language or format labels. (c) Resource placement under fixed path and filename rules. (d) The 15 most frequently referenced external programs in the Python subprocess analysis, ranked by the number of packages containing a literal program-name reference; ties are ordered alphabetically. Panels (a)–(c) use logarithmic count axes; (d) uses a linear axis.
Figure 7: Execution-substrate diversity. (a) Number of distinct dependency packages across eight package ecosystems. (b) Task occurrence frequencies of the 4,242 identified packages, ranked by frequency. (c) Task occurrence frequencies of the 823 identified CLI names, ranked by frequency.
Figure 8: Diversity of 2,963 synthesized tasks. (a) Task counts across 25 source-skill domains, shown with abbreviated labels. (b) Percentage of tasks in each domain containing each workspace asset family; rows align with (a).
Figure 9: Per-benchmark scores across hardening iterations. Bars show scores and lines connect successive iterations. Labels report the scores to two decimal places.
Skills equip LLM agents with professional knowledge and guidance to complete long-horizon and complex tasks. Although skills have been widely adopted in recent agent paradigms and harnesses, how to synthesize reliable training data and how to train agents for skill use remain underexplored. In this work, we propose SkillGym, an automatic pipeline to build verifiable environments, collect trajectories, and train skill-use agents. SkillGym first crawls a large volume of skills from the internet, then keeps those whose workflows can run reproducibly offline. A builder-reviewer pipeline is used to construct difficulty-controlled tasks, spanning four task types, each with a reference solution and an executable verifier. With this pipeline, we build 6.8k environments and collect 19k verified successful trajectories for supervised finetuning. Finetuning on these trajectories improves LLMs of different families and sizes, from 2B to 122B parameters across four skill-use benchmarks; Our Qwen3.5-9B SFT model outperforms the 397B untrained model on two of them. Further analysis shows that training teaches agents to invoke skills, raising the rate of reading the relevant skill from 28% to 96%, and that the gains hold across reasoning structures, extending to task types that form a minority of the training data and to skills held out from training
Renxi Wang, Mingshan Hee, Fajri Koto +2
Mohamed bin Zayed University of Artificial Intelligence
Skills are a promising way to improve LLM agent capabilities without retraining, while keeping the added procedure reusable and controllable. However, high-quality skills are still largely written by hand. We introduce SkillGen, a multi-agent framework that synthesizes a single auditable skill from trajectories generated by a base agent. The output is a human-readable artifact that can be inspected before use. Rather than merely summarizing trajectories, SkillGen leverages contrastive induction over both successful and failed trajectories to identify reusable success patterns, recurring failure modes, and behaviors that appear in nearby successes but are missing from failures. SkillGen then generates candidate skills and iteratively refines the skill. A key novelty in SkillGen is that we model agent skills as interventions to empirically verify the net effect of skills on the overall performance. Specifically, we compare outcomes on the same instances with and without the skill, so that we account for both repairs (cases where the skill fixes a baseline failure) and regressions (cases where the skill breaks a baseline success). Across a broad range of agents and datasets, SkillGen consistently improves held-out performance, outperforms existing skill-generation baselines, and produces skills that transfer across models.
Yuchen Ma, Yue Huang, Han Bao +5
Munich Center for Machine Learning, LMU Munich · University of Notre Dame · Microsoft Research
Agent skills, which consist of reusable strategies that guide agent reasoning and action, have shown strong potential for improving model capability at inference time. However, current skill construction methods treat the problem as one-shot extraction, overlooking a fundamental tension: a skill tailored to the specific task fails to transfer, while the abstracted skill often provides insufficient guidance. We attribute this fragility to the absence of explicit mechanisms for skill specification and generalization. To address this gap, we introduce SkillComposer, a framework that decomposes skill construction into three learnable operations: create, improve, and merge. Trained via systematic rejection sampling recipe, SkillComposer enables language models to self-evolve skills at inference time and supports three deployment modes: offline for building generalized libraries, online for task-specific refinement, and hybrid for combining both. Comprehensive experiments on τ2-Bench, LiveCodeBench v6, and AppWorld show that SkillComposer consistently outperforms baselines. Our SkillComposer-4B improves a 27B executor by up to +4.5 on agent tasks and +3.4 on code tasks, while generalizing across domains and task types unseen during training. Analysis reveals that merge and improve address orthogonal quality dimensions and that skill composition is a transferable meta-ability, providing a practical recipe for skill-augmented inference.
Qi Zhang, Zhaopeng Feng, Xiaonan Shi +8
1Zhejiang University · 2Tongyi Lab · 3National University of Singapore