Agent skills extend coding agents with task-specific instructions, scripts, and resources, but they also create a trusted instruction channel that can be abused beyond conventional security attacks. This paper studies token amplification through skill injection: an economic resource-abuse threat in which a malicious skill causes an agent to consume substantially more tokens than needed for normal task execution. We present SkillBloat, a two-phase framework that first screens a library of diverse attack-type conditions across multiple amplification mechanisms and then refines the strongest candidate through LLM-guided full-document skill rewriting. Evaluated on a real-world skill benchmark, SkillBloat achieves 5.4184x-10.1455x average best amplification across multiple coding-agent target configurations. An ablation shows that the second-stage refinement loop consistently improves average best amplification over Phase 1 attack-type screening alone, demonstrating that iterative optimization provides additional benefit beyond initial attack-type selection. These results show that skill ecosystems expose a practical resource-amplification attack surface that is orthogonal to existing security-oriented skill poisoning.
Figures & tables
Figure 1: Overview of the SkillBloat two-phase pipeline. Starting from an original agent execution on the analyzing-projects skill, Phase 1 screens attack-type conditions that target output inflation, tool-driven amplification, and context amplification, then selects the strongest candidate according to total-token amplification. Phase 2 uses the Phase 1 trace and accumulated feedback to iteratively rewrite the full SKILL.md, execute the target agent, diagnose the outcome, and retain the best adversarial skill. The rightmost panel illustrates how the optimized skill preserves the original request while expanding the agent’s behavior into staged analysis, repeated verification, and quality-assurance passes.
Agent frontend
Backend model
Avg. best amplification
Max
Median
Baseline completion
Best-attack completion
Claude Code
GLM-4.7-Flash
9.3105 ×
71.59 ×
4.89 ×
76.00%
80.00%
GLM-5
5.4184 ×
21.49 ×
4.02 ×
90.00%
92.00%
Codex
GPT-5.4 Mini
10.1455 ×
75.86 ×
6.10 ×
86.00%
92.00%
GPT-5.5
6.0063 ×
26.39 ×
4.36 ×
94.00%
88.00%
Table 1: Main results across coding-agent frontends and backend models. For each skill–task pair, the best attack is the evaluated adversarial run with the highest total-token amplification. Average best amplification is measured relative to the benign baseline, and best-attack completion is estimated using a binary label assigned by a GLM-5 judge.
Agent frontend
Backend model
Phase 1 avg.
No-feedback avg.
Phase 2 avg.
Claude Code
GLM-4.7-Flash
7.5105 ×
8.9700 ×
9.3105 ×
GLM-5
4.0220 ×
5.1518 ×
5.4184 ×
Codex
GPT-5.4 Mini
7.5101 ×
8.8427 ×
10.1455 ×
GPT-5.5
4.0573 ×
5.7904 ×
6.0063 ×
Table 2: Ablation of the two-stage pipeline. Phase 1 and Phase 2 report the average best amplification after attack-type screening and after five rounds of feedback-guided refinement, respectively. No-feedback starts from the Phase 1 best skill, asks the Attack Agent to generate five new rewrites without execution feedback, and selects the best candidate among the Phase 1 winner and these five rewrites as a resampling baseline with the same five additional evaluations as Phase 2.
Figure 2: Example execution traces for the consciousness-principles task. The selected Phase 1 and Phase 2 attacks induce additional validation, retry, and revision steps while receiving a completed label from the task judge.
Stage
Amplification
Total tokens
Total cost
Commands
File edits
Baseline
1.00×
38,410
$0.21
3
0
After Phase 1
10.07×
386,807
$2.10
9
1
After Phase 2
26.39×
1,013,561
$5.41
28
10
Table 3: Stage-level behavior in the consciousness-principles case analysis. The table reports amplification ratios, total tokens, total costs, command executions, and file edits for each stage.
Agent frontend
Backend model
Completion under cap
Avg. total-token amplification
Claude Code
GLM-4.7-Flash
62.00%
1.36 ×
GLM-5
48.00%
1.24 ×
Codex
GPT-5.4 Mini
68.00%
1.39 ×
GPT-5.5
64.00%
1.34 ×
Table 4: Output-token capping defense. For each task, we run the best amplified skill but cap cumulative output tokens at 2× the benign baseline output tokens. Completion is judged by GLM-5 ; total-token amplification is measured relative to the benign baseline total tokens.
Mandates multi-perspective analysis from 5 viewpoints (security, performance, maintainability, UX, reliability) with conflict reconciliation.
Task Decomposition
task_planner
Breaks the task into granular micro-steps using WBS methodology, each requiring individual execution and verification.
Tool-Driven Amplification
Tool Pollution
verbose_checker , qa_pipeline
Exposes multiple quality-assurance tools so that the rewritten skill presents tool execution as routine validation work.
Appendix
Table 5: Complete list of evaluated attack types. Each attack type selects a tool context from the manifest and guides the Attack Agent’s full-document rewrite.