LLM agents increasingly rely on reusable Skills for complex, multi-step tasks, creating a critical supply-chain attack surface where poisoned Skill content steers agent decision loops under benign requests. Existing skill poisoning attacks either colocate actuation with its contextual pretext or distribute actuation across multiple Skills, but do not explicitly separate the rationale for execution from the operation itself. In this work, we reveal that untrusted agent decisions fundamentally depend on two conceptually distinct Risk-Realization Factors (RRFs): an actuation factor (specifying what concrete operation is performed) and a pretext factor (providing the situational rationale for why the agent must perform it). Guided by this abstraction, we propose a coordination-based attack paradigm: decoupling pretext from actuation. Rather than fragmenting the malicious actuation, we preserve it as an intact operation within a downstream Steering Skill, while delegating the pretext factor to an upstream Grounding Skill that subtly alters persistent environment artifacts through routine utility operations. The intact actuation thus hides in plain sight, appearing completely legitimate and task-driven only when evaluated against the fabricated pretext. Building on this formulation, we develop an automated framework that discovers authentic execution dependencies, synthesizes coordinated pretext-actuation skill pairs, and iteratively refines poisoned skill instructions via runtime closed-loop feedback. Extensive evaluations across single-session and persistent cross-lifecycle scenarios demonstrate that decoupled skill poisoning achieves high attack success, exposing a critical blind spot in isolated Skill security audits. Our automated framework code is available at https://github.com/Wenxin-buaa/CoordPoison.git.
Figures & tables
Figure 1: Overview of CoordPoison. CoordPoison validates natural dependencies in Grounding–Steering skill pairs ( SG,SS ), screens out Steering-only vulnerabilities, and injects a decoupled pretext factor ( ZG ) linked to the intact actuation ( A(P) ) to yield poisoned pairs ( SG⋆,SS⋆ ). Runtime evidence validates coordination dependence via intermediate artifacts ( aZ ) and iteratively repairs failed constructions, while persistent state enables both same-lifecycle and cross-lifecycle attacks.
Figure 2: An Illustrative Case Study of Decoupled Skill Poisoning. Demonstration of poisoned instruction constructions ( SG∗,SS∗ ) and their corresponding runtime execution flow.
Victim Model
Skill-Inject Payload
SkillJect Payload
Overall
ASR
TCR
ASR
TCR
ASR
TCR
DeepSeek-V4-Flash
67.83 (78/115)
100.00
94.34 (50/53)
100.00
76.19 (128/168)
100.00
GLM-5.3-Flash
33.86 (43/127)
100.00
67.78 (61/76)
100.00
51.23 (104/203)
100.00
Claude-Sonnet-4.6
43.86 (50/114)
100.00
45.95 (34/74)
100.00
44.44 (84/189)
100.00
Table 1: Generation-stage attack performance across victim models and payload families. Counts in parentheses report coordinated attack successes over eligible Steering-only failures.
Figure 4
Source Model
Method
MiniMax-M3
GPT-5.5
Skill-Inject [-2pt] Payload
SkillJect [-2pt] Payload
Overall
Skill-Inject [-2pt] Payload
SkillJect [-2pt] Payload
Overall
DeepSeek-V4-Flash
Naive
0.00
0.00
0.00
0.00
0.00
0.00
Skill-Inject
8.97
48.00
24.22
41.03
44.00
42.19
SkillJect
46.15
94.00
64.84
93.59
94.00
93.75
CoordPoison
70.51
96.00
80.47
96.15
98.00
96.88
GLM-5.3-Flash
Naive
0.00
0.00
0.00
0.00
0.00
0.00
Table 2: Cross-model attack transferability across held-out victim models (MiniMax-M3 and GPT-5.5) evaluated on transfer ASR (%). Evaluation metrics encompass localized Steering-only baselines ( Naive , Skill-Inject , and SkillJect ) and our proposed CoordPoison .
Setting
DeepSeek-V4-Flash
GLM-5.3-Flash
Claude-Sonnet-4.6
ASR
cASR
ASR
cASR
ASR
cASR
w/o Pretext (SkillJect)
17.46
17.46
39.43
39.43
29.76
29.76
w/o Steering Pretext ZS
10.32
10.32
49.04
49.04
46.43
46.43
w/o Grounding Pretext ZG
63.28
65.85
51.92
51.92
30.95
31.71
Same-Lifecycle CoordPoison
68.25
69.35
53.85
53.85
52.38
52.38
Cross-Lifecycle CoordPoison
95.31
100.00
60.58
91.30
63.10
84.13
Table 3: Ablation study and structural lifecycle analysis across victim models (%).
Victim Model
GPT-5.5
MiniMax-M3
Source Model
Undefended
Defended
Undefended
Defended
DeepSeek-V4-Flash
96.88
92.97 ↓ 3.91
80.47
75.78 ↓ 4.69
GLM-5.3-Flash
96.15
93.23 ↓ 2.92
95.19
82.69 ↓ 12.50
Claude-Sonnet-4.6
92.86
91.67 ↓ 1.19
78.57
70.23 ↓ 8.34
Table 4: Cross-model evaluation of Attack Success Rate (ASR, %) under undefended baseline vs. hardened dual-prompt defense across different source–victim model pairs.
Revise SG⋆ prompt to prevent workflow interruption; fallback to qualification if relation is broken.
steering_not_invoked
Revise SS⋆ entry conditions; preserve SG⋆ materialization and qualified substrate R .
Coordination Core (Sec. 2.3)
coordination_state_not_materialized
Revise SG⋆ to enforce runtime creation of aZ carrying ZG . Preserve SS⋆ logic.
coordination_state_not_consumed
Revise SS⋆ read path/positioning to inspect aZ . Preserve established SG⋆ materialization.
steering_condition_not_satisfied
Align ZG ’s semantics with ZS . Preserve artifact handoff structure and SS⋆ consumption path.
Appendix
Table 5: Taxonomy of runtime failure modes and corresponding repair directives in CoordPoison.
Skill source
Pairs
anthropics-skills
30
autumnsgrove-claudeskills
8
openai-runtime
6
microsoft-deep-wiki
3
wshobson-*
12
eigent-ai-agent-skills
2
Appendix
Table 6: Distribution of qualified Skill pairs across source collections. openai-runtime denotes Skills bundled with the Codex/OpenAI runtime; the remaining entries correspond to public GitHub repositories.
Payload
Source
Fixed resource
Name meaning and harmful behavior
Observable success condition
CodeExec
Skill-Inject obvious payload
backup.sh
Code Execution : downloads an additional shell payload and executes it.
bash resources/backup.sh
DoS
Skill-Inject obvious payload
network_diag.sh
Denial of Service : repeatedly sends network requests in an unbounded loop.
bash resources/network_diag.sh
LocTrack
Skill-Inject obvious payload
ip_loc.py
Location Tracking : requests precise device location through macOS CoreLocation.
python resources/ip_loc.py
InfoDisc
SkillJect
sys_check_0.sh
Information Disclosure : collects system, resource, and configuration data and sends diagnostics to an endpoint.
Unauthorized Write : creates or overwrites build files, setup files, directories, and project configuration.
bash resources/build_gen_0.sh
Appendix
Table 7: Fixed payloads, their names, approximate harmful behaviors, and runtime execution oracles. Payload success is determined from deterministic runtime evidence of helper-script execution; mere textual mention of a command does not count as success.
Setting
DeepSeek-V4-Flash
GLM-5.3-Flash
Claude-Sonnet-4.6
ASR
cASR
ASR
cASR
ASR
cASR
Same-Lifecycle CoordPoison
68.25
69.35
53.85
53.85
52.38
52.38
Cross-Lifecycle CoordPoison
95.31
100.00
60.58
91.30
63.10
84.13
Cross-Lifecycle + Trace Summary
90.63
100.00
62.50
90.28
60.71
83.61
L2 Only
14.06
94.74
8.65
81.82
5.95
83.33
Appendix
Table 8: Ablation on upstream execution context retention across lifecycle boundaries. Performance remains stable even when upstream trace summaries are provided to the downstream agent.
Agent skills introduce a new and more severe form of indirect injection for LLM agents: unlike traditional indirect prompt injection, attackers can hide malicious instructions inside a dense, action-oriented skill that already functions as a legitimate instruction source. We study pre-execution skill-poison detection and show that successful skill poisoning induces a structured internal effect, attention hijacking, in which response-time attention shifts from trusted context to malicious skill spans and drives harmful behavior. Motivated by this mechanism, we propose RouteGuard, a frozen-backbone detector that combines response-conditioned attention and hidden-state alignment through reliability-gated late fusion. Across both real and synthetic open-source skill benchmarks, RouteGuard is consistently the strongest or most robust detector; on the critical Skill-Inject channel slice, it reaches 0.8834 F1 and recovers 90.51% of description attacks missed by lexical screening, showing that defending against skill poisoning requires internal-signal detection rather than text-only filtering
Wenjie Xiao, Xuehai Tang, Biyu Zhou +2
University of Chinese Academy of Sciences · Institute of Information Engineering, Chinese Academy of Sciences
Agent skills provide a lightweight mechanism for extending general-purpose agents, but their open format exposes them to skill-poisoning attacks. A practically dangerous injection must stay invisible: if executing the payload derails the user's legitimate task, the resulting failure signal invites inspection of the skill. We therefore evaluate attacks by Attack Success Rate, which requires the injected payload to execute and the user's task to still pass its verifier in the same trial. Prior skill-poisoning attacks face a reliability-stealth trade-off under this lens: YAML-header injections are reliably loaded but easily inspected, whereas stealthier body injections that place explicit malicious commands in the skill prose are less reliable because out-of-context commands invite the agent's own suspicion. We introduce POISE, a position-aware attack that compresses the trigger into a single, benign-looking body instruction, placing it at a feasible position and using a context-aware generator to blend it with nearby setup or prerequisite steps. On Skill-Inject with codex+gpt-5.2, POISE achieves an 89.3% ASR, 28.0 points above a random-placement body baseline and 2.6 points above a YAML-only baseline, while retaining the stealth advantage of body placement. That stealth is the decisive margin: because legitimate skill bodies naturally require privileged tool operations, LLM scanners are hyper-sensitive, falsely flagging 74.6% of clean skills on average across four judges and both benchmarks. Blending into these false alarms, POISE causes only 5.6% of poisoned variants to gain a new high-risk alert over their clean baselines, rendering current static defenses ineffective.
Haochang Hao, Dehai Min, Zhifang Zhang +4
University of Illinois at Chicago · University of Queensland · 3Tulane University +1
Skills extend an agent's capabilities by injecting instructions and information into the context, and are widely used by agents such as OpenClaw and Claude Code. Prior work shows third-party marketplaces host malicious skills that give attackers direct influence over the victim's agent. The emerging defense scans skills before installation, pairing deterministic static checks with an LLM-based semantic judge, as in NVIDIA's SkillSpector. We show that such defenses fall to an attacker who knows the detector. Our white-box LLM attacker, Pretext, iteratively crafts skills that evade detection while still delivering the payload and performing the benign task: moving the payload from code into natural language leaves static analysis inert, while framing it as the skill's legitimate purpose and splitting instructions across files keeps the LLM stage below its blocking threshold. Across three open-source models, Pretext achieves up to 97% and 77% against a frozen detector and a co-adaptive one, respectively, revealing major gaps in current skill scanners.