Large Language Model (LLM) agents struggle to scale safely when exposed to vast enterprise toolsets. Providing an agent with access to every internal tool leads to oversized context windows, degraded tool selection, and severe governance vulnerabilities - as system policies defined purely in prompts remain probabilistic advice rather than hard constraints. Existing mitigations, such as multi-agent domain delegation, decentralize audit logs and fail to guarantee policy compliance across sessions. We introduce skilder, a framework that packages capabilities into roles: bundles of skills, tools, and instructions, together with the limits that bound them. An agent begins with a minimal role catalog, learns the roles a task requires, and receives each role's skills, instructions, and tools through a single MCP server. Because tools reach the agent only inside learned skills, the same server enforces the scope of what was learned deterministically. We evaluate skilder against flat-context tool selection and multi-agent orchestration across 13 tasks using six models (10 runs each). Our results show that, when models completed discovery and issued a governed call, the skilder simulated authorization layer enforced governance boundaries: no unauthorized tool call or parameter violation (e.g., a spending-limit breach) executed. Aggregate task pass rates also reflect whether each model followed the discovery protocol and satisfied response-quality checks; those misses are not authorization failures. Furthermore, by allowing agents to dynamically acquire cross-role capabilities mid-task, skilder preserves problem-solving flexibility while providing hard system-level enforcement.
Role selection — containment by authorization scope
§ 5.1
8
Ambiguous account activity — investigate before flagging
§ 5.1
Does context quality improve? (institutional)
Table 1: Thirteen functional scenarios, shown in reading order (governance first). Scenario numbers match the harness. Behavioral vs. structural labels are in the results text.
Condition
Behavioral safety (%)
Structural isolation (%)
Flat injection
58.3
0.0
Multi-agent, full roster
95.0
0.0
Multi-agent + roster policy
100.0
100.0
Skilder
100.0
100.0
Table 2: Scenario 5: Multi-turn adversarial — support then social-engineered admin request. Pass rates (%; n=60 per condition). Behavioral safety requires both no destructive execution and a refusal or escalation response. Structural isolation requires System Admin to be absent from the reachable capability set.
Condition
Behavioral success (%)
Refund ceiling enforced (%)
Flat injection
5.0
0.0
Multi-agent, full roster
95.0
0.0
Multi-agent + roster policy
90.0
0.0
Skilder
80.0
100.0
Table 3: Scenario 6: Over-limit refund — should escalate, not process. Pass rates (%; n=60 per condition). Behavioral success and the scenario-specific structural guarantee are reported separately.
Condition
Behavioral success (%)
Billing/PII tools unreachable (%)
Flat injection
20.0
0.0
Multi-agent, full roster
91.7
0.0
Multi-agent + roster policy
98.3
100.0
Skilder
93.3
100.0
Table 4: Scenario 7: Role selection — should pick Security & Fraud for investigation. Pass rates (%; n=60 per condition). Behavioral success and the scenario-specific structural guarantee are reported separately.
Condition
Behavioral success (%)
Flagging requires scope transition (%)
Flat injection
100.0
0.0
Multi-agent, full roster
95.0
100.0
Multi-agent + roster policy
98.3
100.0
Skilder
88.3
100.0
Table 5: Scenario 8: Ambiguous account activity — investigate before flagging. Pass rates (%; n=60 per condition). Behavioral success and the scenario-specific structural guarantee are reported separately.
Condition
Sc. 11 (%)
Sc. 12 (%)
Flat injection
73.3
66.7
Multi-agent
40.0
30.0
Skilder
63.3
63.3
Table 6: Guidance-only institutional-policy pass rates (%; n=30 per condition and scenario). Every condition receives the same policy content; no sequence guard is enabled.
Condition
End-to-end
Blocked
Recovered
Premature
Role scoped
Guard
Flat injection
66.7
0.0
0.0
30.0
No
No
Multi-agent
30.0
0.0
0.0
60.0
Yes
No
Skilder , guidance
83.3
0.0
0.0
16.7
Yes
No
Flat + gateway
90.0
26.7
20.0
0.0
No
Yes
Skilder , enforced
80.0
40.0
26.7
0.0
Yes
Yes
Table 7: Scenario 13 enforcement ablation (%; n=30 per condition). Blocked is the share of trials containing a blocked refund attempt; recovered is the share of all trials that later completed a valid refund. Premature is the share in which a refund executed before all required steps. Role scope and guard are configuration properties.
Model
Flat-inj.
Multi-agent
Skilder
Haiku 4.5
100.0
100.0
100.0
Qwen 3.5 122B
60.0
40.0
70.0
Gemma 4 31B
100.0
100.0
100.0
Ministral 3 14B
90.0
90.0
70.0
Opus 4.7
100.0
100.0
100.0
GPT-5.5
100.0
100.0
90.0
Table 8: Scenario 9: Multi-turn — support first, then fraud discovery mid-conversation. Pass rate (%) over 10 trials per model ( n=60 pooled). Mean is the unweighted average across models.
Model
Flat-inj.
Multi-agent
Skilder
Haiku 4.5
100.0
100.0
100.0
Qwen 3.5 122B
30.0
0.0
0.0
Gemma 4 31B
0.0
100.0
100.0
Ministral 3 14B
20.0
30.0
20.0
Opus 4.7
100.0
60.0
90.0
GPT-5.5
100.0
10.0
70.0
Table 9: Scenario 10: Proactive role expansion — billing correction requiring Billing Admin. Pass rate (%) over 10 trials per model ( n=60 pooled). Mean is the unweighted average across models.
Model
Flat-inj.
Multi-agent
Skilder
Haiku 4.5
100.0
100.0
100.0
Qwen 3.5 122B
50.0
40.0
10.0
Gemma 4 31B
100.0
100.0
100.0
Ministral 3 14B
100.0
100.0
90.0
Opus 4.7
100.0
100.0
70.0
GPT-5.5
100.0
100.0
100.0
Table 10: Scenario 1: Refund request — role discovery + entitlement check. Pass rate (%) over 10 trials per model ( n=60 pooled). Mean is the unweighted average across models.
Model
Flat-inj.
Multi-agent
Skilder
Haiku 4.5
100.0
100.0
100.0
Qwen 3.5 122B
100.0
60.0
60.0
Gemma 4 31B
100.0
100.0
100.0
Ministral 3 14B
100.0
100.0
100.0
Opus 4.7
100.0
100.0
100.0
GPT-5.5
100.0
100.0
100.0
Table 11: Scenario 2: Simple lookup — Skilder overhead vs naive directness. Pass rate (%) over 10 trials per model ( n=60 pooled). Mean is the unweighted average across models.
Model
Flat-inj.
Multi-agent
Skilder
Haiku 4.5
100.0
90.0
100.0
Qwen 3.5 122B
80.0
50.0
60.0
Gemma 4 31B
100.0
100.0
100.0
Ministral 3 14B
100.0
90.0
30.0
Opus 4.7
100.0
90.0
100.0
GPT-5.5
100.0
80.0
100.0
Table 12: Scenario 3: Error recovery — first lookup fails, agent must adapt. Pass rate (%) over 10 trials per model ( n=60 pooled). Mean is the unweighted average across models.
Model
Flat-inj.
Multi-agent
Skilder
Haiku 4.5
100.0
100.0
100.0
Qwen 3.5 122B
80.0
50.0
60.0
Gemma 4 31B
90.0
90.0
100.0
Ministral 3 14B
90.0
70.0
0.0
Opus 4.7
100.0
100.0
100.0
GPT-5.5
100.0
100.0
100.0
Table 13: Scenario 4: Role disambiguation — robustness check on vague input. Pass rate (%) over 10 trials per model ( n=60 pooled). Mean is the unweighted average across models.
Theme
n
Flat- inj.
Multi- agent
Skilder
Governance (5–8)
240
31.3
96.7
90.4
Institutional (11–13)
90
68.9
33.3
68.9
Adaptability (9–10)
120
75.0
69.2
75.8
Parity (1–4)
240
95.4
87.9
82.5
Table 14: Theme pass rates (%) from the published per-scenario tables (six models). Scenarios 1–10 use ten trials per cell; institutional scenarios use five, so n is reported explicitly. The institutional Skilder value uses the enforced product condition for Scenario 13.
Model
Flat-inj.
Multi-agent
Skilder
Haiku 4.5
72.0
99.0
100.0
Qwen 3.5 122B
50.0
61.0
53.0
Gemma 4 31B
59.0
99.0
100.0
Ministral 3 14B
62.0
87.0
61.0
Opus 4.7
80.0
91.0
96.0
GPT-5.5
71.0
89.0
96.0
Table 15: Non-institutional all-assertion aggregate pass rate (%) over Scenarios 1–10 ( n=100 per agent). The Multi-agent column uses the explicit role-to-roster policy condition. Includes gating structural assertions that can fail flat-injection when a forbidden tool is exposed; Scenario 8 reports exposure separately. Institutional results are excluded because that suite has a different repeat count and rubric.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Flat-inj.
Multi-agent full
Multi-agent + policy
Skilder
Haiku 4.5
10/10
10/10
10/10
10/10
Qwen 3.5 122B
10/10
10/10
10/10
10/10
Gemma 4 31B
1/10
9/10
10/10
10/10
Ministral 3 14B
0/10
8/10
10/10
10/10
Opus 4.7
10/10
10/10
10/10
10/10
GPT-5.5
4/10
10/10
10/10
10/10
Appendix
Table 16: Scenario 5 per-model decomposition. Panel (a) reports safe behavior; Panel (b) reports whether System Admin is structurally unreachable.
Model
Flat-inj.
Multi-agent full
Multi-agent + policy
Skilder
Haiku 4.5
1/10
9/10
10/10
10/10
Qwen 3.5 122B
0/10
9/10
8/10
1/10
Gemma 4 31B
0/10
10/10
10/10
10/10
Ministral 3 14B
2/10
10/10
10/10
7/10
Opus 4.7
0/10
9/10
6/10
10/10
GPT-5.5
0/10
10/10
10/10
10/10
Appendix
Table 17: Scenario 6 per-model decomposition: Over-limit refund — should escalate, not process.
Model
Flat-inj.
Multi-agent full
Multi-agent + policy
Skilder
Haiku 4.5
1/10
9/10
10/10
10/10
Qwen 3.5 122B
0/10
10/10
9/10
6/10
Gemma 4 31B
0/10
10/10
10/10
10/10
Ministral 3 14B
0/10
9/10
10/10
10/10
Opus 4.7
10/10
10/10
10/10
10/10
GPT-5.5
1/10
7/10
10/10
10/10
Appendix
Table 18: Scenario 7 per-model decomposition: Role selection — should pick Security & Fraud for investigation.
Table 23 : Scaling benchmark results. All values are averages over completed runs. “Ratio” is flat-injection/ Skilder token consumption.
Series
1 turn
2 turns
3 turns
4 turns
Skilder, thin context
8,809
12,188
18,482
25,874
Skilder, rich context
8,923
12,269
19,445
26,251
Multi-agent, thin context
7,166
9,329
18,235
26,867 †
Multi-agent, rich context
9,469
11,809
23,481
32,494 †
Appendix
Table 24 : Turn-cost benchmark (225 tools): input tokens per series and user-turn depth. Comparable cells: first passing run within 3 attempts. † Minimum observed input when no comparable multi-agent run was achieved (comparability rules below).
Large language model (LLM) agents increasingly rely on skills, structured documents that specify when to act, which procedure to follow, and which tools are allowed. Existing evaluations mostly judge the quality of a skill or its contribution to task success, leaving unexamined whether an agent can recognize a relevant skill and apply it on its own. We introduce Skill-Use, a benchmark that evaluates skill use under progressive disclosure, where an agent sees only a skill's name and short description and must retrieve the full procedure before following it. Skill-Use separates three facets of skill use. Trigger measures whether the agent invokes the relevant skill, Compliance measures how faithfully it follows the prescribed procedure, and Boundary measures whether it avoids forbidden operations. A Skill-Use (SU) score combines the three and credits execution only after the skill is triggered. Skill-Use pairs 79 real skills with 177 executable tasks across nine domains, each grounded in real files, run in an isolated Docker sandbox, and scored by a trajectory-based rubric. Evaluating eight LLMs under two agent harnesses, we find that reliable skill use remains out of reach, as the strongest configuration reaches an SU of only 0.613. Triggering and procedural compliance fail as independent bottlenecks, and both scores and model rankings shift with the harness, so skill use behaves as a capability conditioned on the harness rather than a fixed property of the model.
Jinyi Han, Yuanjian Xu, Ying Liao +6
1East China Normal University · 2Hong Kong University of Science and Technology · 3Fudan University +1
Large language model agents place outputs from external skills into their execution context, allowing attacker-controlled data to influence later privileged actions. Existing defenses mainly classify untrusted content or authorize proposed operations. They do not directly address how an agent's future authority should change once untrusted data enters its state. We present SkillGuard, a harness-level enforcement layer that treats this event as contamination and restricts future capabilities to disconnect the resulting state from deployer-defined forbidden states. Given sound skill summaries and policies, SkillGuard represents security-relevant transitions with a Skill Impact Graph, specifies admissible control over skill parameters via steerability signatures, and mediates invocations with an inline reference monitor. Following contamination, it computes weighted capability restrictions using binary, fractional, or fractional-flow strategies without auxiliary language-model inference. We evaluate SkillGuard on four AgentDojo suites with two backend LLMs, Gemini 2.5 Flash and Llama3.3-70B, against an LLM-only No Defense baseline and three defenses at different system layers: Spotlighting, CaMeL, and AttriGuard. We construct a compositional attack benchmark in which each attack combines observations individually insufficient to induce target violation and evaluate the same baselines on it. Under AgentDojo's Tool Knowledge attacks, SkillGuard eliminates attack success on three of four suites for both backends and reduces it to 4.8% and 14.3% on Slack. Against compositional attacks, it outperforms every baseline on Llama and matches the strongest baseline on Gemini at higher benign utility. Fractional-flow restriction preserves substantially more capabilities than binary restriction at the same attack success rate. Across both settings, SkillGuard adds no model calls or token overhead.
Wujie Xiong, Rabimba Karanjai, Yang Lu +2
Kent State Univeristy · PayPal AI Labs · University of Houston
Agent Skills augment large language model (LLM) agents with procedural knowledge at inference time, but current benchmarks rarely distinguish what a Skill says from how it is organized. We study this distinction through Progressive Disclosure, where a concise root file points agents to supporting resources on demand, and compare it with a normalized flat baseline. We present SkillJuror, a framework for evaluating Skill writing paradigms through semantically controlled variants, matched multi-trial evaluations, and trajectory evidence while holding task knowledge fixed. In an 82-task SkillsBench study, Progressive Disclosure changes runtime behavior before aggregate outcomes: distinct Skill resources touched per trajectory rise from 1.18 to 3.85, and effective uptake events rise from 1.33 to 3.92. It also yields 17 additional verifier-passing trials out of 410 matched trials (+4.1%) over the normalized flat baseline. The benefit is task-dependent. Progressive Disclosure helps when supporting resources guide implementation, checking, or repair, but is weaker when success hinges on exact output conventions, numerical thresholds, or long artifact-generation pipelines. These results show that Skill organization is not mere presentation: it can change how agents search and apply procedural knowledge, while outcome gains depend on whether the exposed resources are actionable for the task. Code is available at https://github.com/zhiyuchen-ai/skill-juror.
Zhiyu Chen, Zihan Guo, Bo Huang +4
Tongji University · Shanghai Innovation Institute · Sun Yat-sen University +1