Agent Skills package procedural guidance and resources for reuse, but a relevant Skill does not necessarily improve task performance. Existing studies characterize Skill content and evaluate downstream performance, yet provide limited explanations of how utility depends on content, execution configuration, and multi-Skill organization. We conduct an empirical study on 87 SkillsBench tasks, defining downstream utility as the pass-rate difference from No-Skill on the same tasks under the same model--harness configuration. We compare the same Skills across nine configurations, then examine alternative published Skills and organizations of fixed Skill sets under three selected configurations. We retrieve marketplace candidates from a curated corpus of 37,596 Skills. LLM-assisted analysis of content, execution traces, and final artifacts, followed by author review, relates provided support to actual use and task outcomes. The same Skills help some configurations and hurt others on 36.78% of tasks, with trajectories showing that recommended procedures can become an execution burden. Relevance rankings overlook more useful candidates. Within the evaluated candidate sets, reranking by support for required operations raises first-choice pass rates by 4.35--5.80 percentage points across the three configurations. We derive 17 authoring practices linking executable procedures to recovery, preservation of task requirements, and checks on final artifacts. Stage Plan and Dependency DAG outperform use order alone, with DAG's additional benefits concentrated in tasks supplied with five or six Skills. These findings guide developers to assess usable operation support, allow procedure adaptation while preserving task requirements, and make artifact dependencies explicit when organizing Skills.
Figures & tables
Figure 1. Overview of the empirical study on Agent Skill utility. Overview of the empirical study. A shared Skill corpus, task benchmark, and experimental setup support three research questions on model–harness configurations, alternative Skills for the same task, and multi-Skill organization. Content and trajectory analysis link observed utility differences to Skill authoring practices.
Stage
Collection
Completeness
Popularity
Language
Deduplication
Records
124,425
113,972
65,048
40,412
37,596
Table 1. Skill corpus construction.
Skills per task
1
2
3
4
5
6
7
Tasks
23
23
20
9
6
5
1
Table 2. Distribution of tasks by supplied Skill count.
Configuration
No-Skill (%)
Benchmark-Skill (%)
Gain (pp)
GPT + Codex
42.15
54.02
11.88
GPT + OpenClaw
30.65
47.89
17.24
GPT + OpenCode
24.14
43.68
19.54
DeepSeek + Codex
34.10
44.44
10.34
DeepSeek + OpenClaw
35.63
50.19
14.56
DeepSeek + OpenCode
36.40
48.28
11.88
Table 3. Pass rates and gains over No-Skill across model–harness configurations. Bold values indicate the highest Benchmark-Skill pass rate and gain within each model.
Figure 2. Task-level gains over No-Skill from the same Benchmark Skills across nine configurations. Rows and columns represent configurations and tasks, respectively. Tasks are grouped by gains and losses across configurations, with group sizes in parentheses. A heatmap with nine configuration rows and 87 task columns. Teal indicates a pass-rate increase relative to No-Skill, red a decrease, and light gray no change. Thirty-two tasks show both gains and losses across configurations; 32 show gains with no losses; eight show losses with no gains; and 15 show no change in any configuration. Tasks are ordered alphabetically within each group.
Configuration
No-Skill
Benchmark- Skill
M1
M2
M3
M4
M5
GPT + Codex
36.23
49.28
42.03
40.58
43.48
37.68
37.68
DeepSeek + OpenClaw
28.99
44.93
40.58
39.13
37.68
42.03
36.23
Qwen + OpenCode
17.39
18.84
14.49
20.29
14.49
15.94
17.39
Table 4. Pass rates (%) by original relevance rank.
Configuration
No-Skill
Benchmark- Skill
M1
MR1
Gain (pp)
GPT + Codex
36.23
49.28
42.03
46.38
+4.35
DeepSeek + OpenClaw
28.99
44.93
40.58
44.93
+4.35
Qwen + OpenCode
17.39
18.84
14.49
20.29
+5.80
Table 5. Pass rates (%) and reranking gains over M1.
Table 6. Five-stage Skill authoring framework with 17 practices.
Configuration
No-Skill
Flat
Sequence
Stage Plan
Dependency DAG
GPT + Codex
44.72
52.03
52.85
55.28
58.54
DeepSeek + OpenClaw
34.96
39.84
47.97
51.22
51.22
Qwen + OpenCode
11.38
21.14
21.95
25.20
27.64
Table 7. Pass rates (%) under alternative organizations of the same Skill set.
Figure 3. Pass rates by the number of supplied Skills. Three line charts compare No-Skill, Flat, Sequence, Stage Plan, and Dependency DAG across groups with three to seven supplied Skills. Stage Plan and DAG have identical overall pass rates in the four-Skill group in all configurations and in the three-Skill group for GPT and Qwen, with a small difference for DeepSeek. Their gains over Flat are larger in the combined five- and six-Skill group than in the combined three- and four-Skill group in each configuration.
Agent Skills augment large language model (LLM) agents with procedural knowledge at inference time, but current benchmarks rarely distinguish what a Skill says from how it is organized. We study this distinction through Progressive Disclosure, where a concise root file points agents to supporting resources on demand, and compare it with a normalized flat baseline. We present SkillJuror, a framework for evaluating Skill writing paradigms through semantically controlled variants, matched multi-trial evaluations, and trajectory evidence while holding task knowledge fixed. In an 82-task SkillsBench study, Progressive Disclosure changes runtime behavior before aggregate outcomes: distinct Skill resources touched per trajectory rise from 1.18 to 3.85, and effective uptake events rise from 1.33 to 3.92. It also yields 17 additional verifier-passing trials out of 410 matched trials (+4.1%) over the normalized flat baseline. The benefit is task-dependent. Progressive Disclosure helps when supporting resources guide implementation, checking, or repair, but is weaker when success hinges on exact output conventions, numerical thresholds, or long artifact-generation pipelines. These results show that Skill organization is not mere presentation: it can change how agents search and apply procedural knowledge, while outcome gains depend on whether the exposed resources are actionable for the task. Code is available at https://github.com/zhiyuchen-ai/skill-juror.
Zhiyu Chen, Zihan Guo, Bo Huang +4
Tongji University · Shanghai Innovation Institute · Sun Yat-sen University +1
Agent skills are the de facto mechanism for extending LLM agents with reusable guidance. A skill can shape the agent's task execution, including planning, tool use, problem-solving, and validation. Prior work reported mixed results of agent skills: some skills improve task success rates, while others have no effect, increase token use and execution time, and even reduce success rates. This paper presents a comprehensive analysis of skill-induced agent failures by attributing task failures and cost regressions to specific loaded skills. We introduce a differential analysis framework that attributes a failure or regression to a skill by comparing a target skill-guided run against a no-skill or semantically matched skill reference run that solves the same task, or solves it more cheaply. We instantiate this framework on SkillsBench and SWE-Skills-Bench, yielding 307 skill-induced failures, including 125 functional failures and 182 efficiency regressions. We also build SkillTriage, a taxonomy-guided attribution tool that normalizes paired cases, extracts differential evidence, and produces triage reports. Our major findings include: (1) Skill induced functional failures are rarely caused by obviously irrelevant skills; instead, seemingly relevant skills often make the agent incorrectly implement or omit task-required implementation elements. (2) Skill-induced efficiency regressions are not explained by prompt length alone. (3) The largest sources within Excessive Procedure are excessive verification and heavy implementation pipelines, contributing 67 and 30 cases, respectively. This shows that skills often turn validation checklists and construction recipes into mandatory work. Based on our findings, we propose research topics and tooling improvements for safer and more cost-aware skill reuse.
Gen Dong, Yanjie Gao, Liqun Li +3
Huazhong University of Science and Technology · Microsoft Research · Microsoft +1
Agent skills encode reusable procedural knowledge for large language model (LLM) agents, and existing benchmarks show that such skills can improve task-level performance. However, a task outcome does not reveal which parts of a reusable skill were exercised, nor whether the agent followed the relevant skill instructions when those parts were exercised. This gap makes it unclear whether a skill has been adequately tested, or whether observed task failures provide actionable evidence for improving agent skill effectiveness. To fill this gap, we introduce skill coverage, a trajectory-based test-adequacy metric for reusable agent skills. Our framework extracts skill behavior constraints from each skill, translating natural-language skill instructions into semi-structured constraints that specify the expected agent behavior under particular conditions. It then determines whether each constraint is covered by an agent trajectory and, for covered constraints, assigns a Pass or Fail verdict according to the agent behavior. We apply this framework to SkillsBench. The results show that agent trajectories on the benchmark leaderboard cover only 38.66 to 45.51% of the extracted skill behavior constraints on average. We then use Fail verdicts to strengthen the corresponding skill content only by emphasizing the original instructions that the agent failed to follow, and run the same tasks with the strengthened skills. This emphasis yields an average 16.0% recovery rate of the failed tasks across the five agent-model rows. These results show that skill coverage is both a test-adequacy metric and a fine-grained signal for observing skill-use behavior. In failed tasks, failed constraint labels provide actionable evidence for improving agent skill effectiveness. A project website accompanies the paper.
Boyin Tan, Xiaowei Huang, Youcheng Sun
1Mohamed bin Zayed University of Artificial Intelligence Abu Dhabi, United Arab Emirates · University of Liverpool Liverpool, United Kingdom