Agent skills extend coding agents with task-specific instructions, scripts, and resources, but they also create a trusted instruction channel that can be abused beyond conventional security attacks. This paper studies token amplification through skill injection: an economic resource-abuse threat in which a malicious skill causes an agent to consume substantially more tokens than needed for normal task execution. We present SkillBloat, a two-phase framework that first screens a library of diverse attack-type conditions across multiple amplification mechanisms and then refines the strongest candidate through LLM-guided full-document skill rewriting. Evaluated on a real-world skill benchmark, SkillBloat achieves 5.4184x-10.1455x average best amplification across multiple coding-agent target configurations. An ablation shows that the second-stage refinement loop consistently improves average best amplification over Phase 1 attack-type screening alone, demonstrating that iterative optimization provides additional benefit beyond initial attack-type selection. These results show that skill ecosystems expose a practical resource-amplification attack surface that is orthogonal to existing security-oriented skill poisoning.
Figures & tables
Figure 1: Overview of the SkillBloat two-phase pipeline. Starting from an original agent execution on the analyzing-projects skill, Phase 1 screens attack-type conditions that target output inflation, tool-driven amplification, and context amplification, then selects the strongest candidate according to total-token amplification. Phase 2 uses the Phase 1 trace and accumulated feedback to iteratively rewrite the full SKILL.md, execute the target agent, diagnose the outcome, and retain the best adversarial skill. The rightmost panel illustrates how the optimized skill preserves the original request while expanding the agent’s behavior into staged analysis, repeated verification, and quality-assurance passes.
Agent frontend
Backend model
Avg. best amplification
Max
Median
Baseline completion
Best-attack completion
Claude Code
GLM-4.7-Flash
9.3105 ×
71.59 ×
4.89 ×
76.00%
80.00%
GLM-5
5.4184 ×
21.49 ×
4.02 ×
90.00%
92.00%
Codex
GPT-5.4 Mini
10.1455 ×
75.86 ×
6.10 ×
86.00%
92.00%
GPT-5.5
6.0063 ×
26.39 ×
4.36 ×
94.00%
88.00%
Table 1: Main results across coding-agent frontends and backend models. For each skill–task pair, the best attack is the evaluated adversarial run with the highest total-token amplification. Average best amplification is measured relative to the benign baseline, and best-attack completion is estimated using a binary label assigned by a GLM-5 judge.
Agent frontend
Backend model
Phase 1 avg.
No-feedback avg.
Phase 2 avg.
Claude Code
GLM-4.7-Flash
7.5105 ×
8.9700 ×
9.3105 ×
GLM-5
4.0220 ×
5.1518 ×
5.4184 ×
Codex
GPT-5.4 Mini
7.5101 ×
8.8427 ×
10.1455 ×
GPT-5.5
4.0573 ×
5.7904 ×
6.0063 ×
Table 2: Ablation of the two-stage pipeline. Phase 1 and Phase 2 report the average best amplification after attack-type screening and after five rounds of feedback-guided refinement, respectively. No-feedback starts from the Phase 1 best skill, asks the Attack Agent to generate five new rewrites without execution feedback, and selects the best candidate among the Phase 1 winner and these five rewrites as a resampling baseline with the same five additional evaluations as Phase 2.
Figure 2: Example execution traces for the consciousness-principles task. The selected Phase 1 and Phase 2 attacks induce additional validation, retry, and revision steps while receiving a completed label from the task judge.
Stage
Amplification
Total tokens
Total cost
Commands
File edits
Baseline
1.00×
38,410
$0.21
3
0
After Phase 1
10.07×
386,807
$2.10
9
1
After Phase 2
26.39×
1,013,561
$5.41
28
10
Table 3: Stage-level behavior in the consciousness-principles case analysis. The table reports amplification ratios, total tokens, total costs, command executions, and file edits for each stage.
Agent frontend
Backend model
Completion under cap
Avg. total-token amplification
Claude Code
GLM-4.7-Flash
62.00%
1.36 ×
GLM-5
48.00%
1.24 ×
Codex
GPT-5.4 Mini
68.00%
1.39 ×
GPT-5.5
64.00%
1.34 ×
Table 4: Output-token capping defense. For each task, we run the best amplified skill but cap cumulative output tokens at 2× the benign baseline output tokens. Completion is judged by GLM-5 ; total-token amplification is measured relative to the benign baseline total tokens.
Mandates multi-perspective analysis from 5 viewpoints (security, performance, maintainability, UX, reliability) with conflict reconciliation.
Task Decomposition
task_planner
Breaks the task into granular micro-steps using WBS methodology, each requiring individual execution and verification.
Tool-Driven Amplification
Tool Pollution
verbose_checker , qa_pipeline
Exposes multiple quality-assurance tools so that the rewritten skill presents tool execution as routine validation work.
Appendix
Table 5: Complete list of evaluated attack types. Each attack type selects a tool context from the manifest and guides the Attack Agent’s full-document rewrite.
Skills extend an agent's capabilities by injecting instructions and information into the context, and are widely used by agents such as OpenClaw and Claude Code. Prior work shows third-party marketplaces host malicious skills that give attackers direct influence over the victim's agent. The emerging defense scans skills before installation, pairing deterministic static checks with an LLM-based semantic judge, as in NVIDIA's SkillSpector. We show that such defenses fall to an attacker who knows the detector. Our white-box LLM attacker, Pretext, iteratively crafts skills that evade detection while still delivering the payload and performing the benign task: moving the payload from code into natural language leaves static analysis inert, while framing it as the skill's legitimate purpose and splitting instructions across files keeps the LLM stage below its blocking threshold. Across three open-source models, Pretext achieves up to 97% and 77% against a frozen detector and a co-adaptive one, respectively, revealing major gaps in current skill scanners.
Agent skills provide a lightweight mechanism for extending general-purpose agents, but their open format exposes them to skill-poisoning attacks. A practically dangerous injection must stay invisible: if executing the payload derails the user's legitimate task, the resulting failure signal invites inspection of the skill. We therefore evaluate attacks by Attack Success Rate, which requires the injected payload to execute and the user's task to still pass its verifier in the same trial. Prior skill-poisoning attacks face a reliability-stealth trade-off under this lens: YAML-header injections are reliably loaded but easily inspected, whereas stealthier body injections that place explicit malicious commands in the skill prose are less reliable because out-of-context commands invite the agent's own suspicion. We introduce POISE, a position-aware attack that compresses the trigger into a single, benign-looking body instruction, placing it at a feasible position and using a context-aware generator to blend it with nearby setup or prerequisite steps. On Skill-Inject with codex+gpt-5.2, POISE achieves an 89.3% ASR, 28.0 points above a random-placement body baseline and 2.6 points above a YAML-only baseline, while retaining the stealth advantage of body placement. That stealth is the decisive margin: because legitimate skill bodies naturally require privileged tool operations, LLM scanners are hyper-sensitive, falsely flagging 74.6% of clean skills on average across four judges and both benchmarks. Blending into these false alarms, POISE causes only 5.6% of poisoned variants to gain a new high-risk alert over their clean baselines, rendering current static defenses ineffective.
Haochang Hao, Dehai Min, Zhifang Zhang +4
University of Illinois at Chicago · University of Queensland · 3Tulane University +1
Large language model (LLM) agents increasingly rely on reusable skills i.e. documents describing task-specific procedures. However, this introduces a new attack surface for agents to manage. We study two complementary directions for this threat. First, we evaluate guardian-based defenses: an intermediary LLM agent that acts as a mediator for skill file access (dynamic guardian) or pre-rewrites these files at build time (static guardian). Across three LLM agent families, our guardians cut attack success rate (ASR) by well over half while preserving task utility. Second, we stress test them through attack reframing using four attacks that preserve the malicious instruction but change the phrasing. For non-guardian setup, the reframing pushes the ASR up to 81.4%, but the dynamic guardian brings it down to 18.6%, showing that real-time mediation is a robust defense.