Skills extend an agent's capabilities by injecting instructions and information into the context, and are widely used by agents such as OpenClaw and Claude Code. Prior work shows third-party marketplaces host malicious skills that give attackers direct influence over the victim's agent. The emerging defense scans skills before installation, pairing deterministic static checks with an LLM-based semantic judge, as in NVIDIA's SkillSpector. We show that such defenses fall to an attacker who knows the detector. Our white-box LLM attacker, Pretext, iteratively crafts skills that evade detection while still delivering the payload and performing the benign task: moving the payload from code into natural language leaves static analysis inert, while framing it as the skill's legitimate purpose and splitting instructions across files keeps the LLM stage below its blocking threshold. Across three open-source models, Pretext achieves up to 97% and 77% against a frozen detector and a co-adaptive one, respectively, revealing major gaps in current skill scanners.
Figures & tables
Figure 1 : An example attack from an attacker-controlled skill.
Figure 2 : The one-run refinement loop.
Figure 3 : The generational learning cycle.
Figure 4 : Left: attacker memory self-convergence (top) and plasticity (bottom) at the final generation, one bar per model. Center: attacker success rate (ASR) per generation for Mode A and Mode B (blind and informed). Right: detector metrics (Mode B only, where the detector learns): final-generation false-positive rate (top) and detector coverage of attacker lessons (bottom). Lines/bars are the mean over 5 replicates; ASR bands are ±1 std.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A.1 : How SkillSpector scans a skill: blue nodes are deterministic, orange nodes are LLM-based, and both LLM stages are skipped under --no-llm .
Rule IDs
Family (what it flags)
Deterministic (no LLM)
P1–P4
prompt injection
P5
harmful content
P6–P8
system-prompt leakage
E1–E4
data exfiltration
PE1–PE3
privilege escalation
Appendix
Table A.1: SkillSpector’s rule families, keyed by the IDs used throughout this paper.
Attack type
Targets (rule families)
Payload in brief
Data exfiltration
E1–E4, TT3–TT4, SSD-3
send env secrets / context to an external endpoint
Prompt injection / leakage
P1–P4, P6–P8, SSD-1–2
override guardrails or leak the system prompt
Memory poisoning
MP1–MP3, E1–E2
persist a standing malicious instruction across sessions
Dangerous code execution
AST1–AST8, TT5, SC2–SC3
emit curl | bash / dynamic-exec code
Privilege escalation
PE1–PE3, AST4, TT5
sudo / edit sudoers / chmod 777 as “setup”
Dependency confusion
SC1, SC2, SC6, TT5
install a confusable package from a public index
Appendix
Table B.1: The 12 trojan attack types and the SkillSpector rules each primarily targets.
Figure B.1 : The Mode-B detector scans matched malicious/benign pairs with a frozen SkillSpector plus a learned-heuristics analyzer, and its reflector turns each generation’s false negatives and false positives into updated heuristics for the next.
Call
Receives (and source)
Withheld
Attacker (design / refine)
fixed attack type (round-robin); its predefined description, payload goal, and success criterion; the full SkillSpector ruleset and its own accumulated memory (cached system context); a mode/condition preamble; on refine, either the fired detector rules + reasoning (if detected) or the victim’s full output + verdict (if it evaded but did not deliver)
the injected SKILL.md as an installed skill; a benign task (attacker-authored); a synthetic capabilities prompt for the type
that it is a security test; any real tools or execution
Verdict judge
the attack-type label; the attacker-authored payload description; the predefined success criterion; the benign task; the victim’s text output
the SKILL.md (cover and payload)
Attacker reflector
the generation’s N run records (mechanical logs: outcome, fired rules per iteration, delivery verdict); the current Global section and the in-batch type sections
the skill bundles themselves
Detector LLM stage (Mode B)
the skill bundle; its own learned heuristics
the attack type
Detector reflector (Mode B)
the generation’s false negatives and false positives (full bundles, + rules fired on the FPs); the current heuristics
—
Appendix
Table B.2: What each LLM call receives and what is withheld. The attacker, verdict judge, and both reflectors are separate stateless calls of the attacker’s model; the detector stage and its reflector use the detector model.
Target stack
Mode A
Mode B blind
Mode B informed
qwen3t
0.2%
1.3%
0.3%
glm
4.8%
2.0%
2.7%
gpt-oss
10.0%
16.0%
17.3%
Appendix
Table C.1: Share of runs excluded from the success-rate denominators because the SkillSpector scan timed out (900 s), per stack and mode.
Metric
Definition
Purpose
Outcome
attacker success rate
fraction of runs that evade the detector (score <20 ), deliver the payload, and complete the benign task
the bottom line: is the attacker winning
detector FN rate
fraction of malicious skills not flagged (score <20 )
the detector’s miss rate on real attacks
detector FP rate
fraction of benign skills flagged (score ≥20 )
false-alarm cost of being trigger-happy
Attacker learning dynamics (Mode A and B)
self-convergence
mean nearest-neighbour cosine between the lessons of the memory at generations t−1 and t , averaged in both directions
is the memory circling one repertoire (mode collapse) or still moving
Appendix
Table C.2: The metrics used in this paper, their definition, and what each is for.
Target stack
success (%)
self-conv.
plasticity
qwen3t
96.7±4.6
0.98±0.01
0.09±0.05
glm
63.2±7.1
0.95±0.04
0.28±0.27
gpt-oss
70.5±14.2
0.96±0.03
0.24±0.13
Appendix
Table D.1: Attacker memory dynamics at the final generation of Mode A (mean ± sample std over 5 replicates).
Stack (cond.)
FN (%)
FP (%)
coverage
time-to-counter
qwen3t (blind)
75±11
40±12
0.54±0.06
0.27±0.04
qwen3t (inf.)
78±18
30±9
0.52±0.14
0.27±0.08
glm (blind)
52±15
20±8
0.61±0.07
0.30±0.11
glm (inf.)
52±14
18±12
0.60±0.10
0.29±0.06
gpt-oss (blind)
59±14
62±14
0.33±0.05
0.52±0.16
gpt-oss (inf.)
47±10
50±8
0.39±0.05
0.53±0.25
Appendix
Table D.2: Mode-B detector breakdown: false-negative/false-positive rates at the final generation, the share of attacker lessons its learned heuristics cover, and effort asymmetry (mean time-to-counter, lower is faster); mean ± sample std over 5 replicates.
Figure D.1 : Mean refinement iterations to a successful attack (cap iter=3 ) per generation, one line per target stack, for Mode A and Mode B (blind and informed); mean over 5 replicates, bands are ± SEM.
Figure D.2 : Memory plasticity (reshape rate) per generation, coloured by target stack, for the attacker (solid) and, in Mode B, the detector (dashed).
Attacker
Mode A
Mode B blind
Mode B informed
deleted (% of prev-gen lessons pruned)
qwen3t
0.8
6.5
3.0
glm
4.7
6.5
4.7
gpt-oss
3.6
5.4
6.9
born (% of current-gen lessons newly added)
qwen3t
8.9
20.7
20.6
Appendix
Table D.3: Attacker memory-shape per generation (deleted, born, and refined lesson shares), averaged over all generations and 5 replicates; these are the components pooled into plasticity.
Figure D.3 : Per-generation attacker memory self-convergence (left) and detector coverage (right) in Mode B, coloured by target stack; mean over the blind and informed conditions and 5 replicates, bands are ± SEM.
Target stack
Skills
Statically detected
Mean static score
qwen3t
60
0 (0.0%)
1.2
glm
59
6 (10.2%)
4.4
gpt-oss
59
4 (6.8%)
3.1
All
178
10 (5.6%)
2.9
Appendix
Table D.4: Static-only ( --no-llm ) detection over the final-generation attacker skills; the deterministic layer alone reaches the 20 threshold on only 5.6% of them.
LLM agents increasingly load skills, file-based packages of natural-language instructions written by third parties and distributed through marketplaces, that execute with the user's privileges. A single malicious skill can exfiltrate data, hijack the agent, or persist as a supply-chain foothold, which turns the skill marketplace into a new attack surface for agentic systems. Prompt-injection defenses do not carry over to this setting. They rely on a boundary between trusted instructions and untrusted data, whereas a skill is itself a body of instructions, so an injected command sits among many legitimate ones and inherits their authority. We present Locate-and-Judge, a two-stage detector designed for this regime. A lightweight locator scores the structural spans of a skill by the instruction-following attention each span draws and retains only the top-K. A judge then examines the retained spans in detail. Concentrating the costly judgment on a few high-attention spans lets the detector audit an entire marketplace instead of a sample. Compared to direct LLM-based scanning, this approach offers an order-of-magnitude cost reduction, dramatically increasing its scalability at a small cost to recall, and it dominates keyword and regex baselines at comparable expense. Deployed at marketplace scale and at negligible cost, Locate-and-Judge flags skills with high precision, the majority of which we manually confirmed as malicious, surfacing dozens of live malicious skills, including several disguised as benign functionality and many that SkillSpector and Cisco Skill Scanner fail to detect. We release the resulting labeled dataset.
Bacem Etteib, Daniele Lunghi, Tégawendé F. Bissyandé
Agent skills let LLM agents reuse instructions, resources, tools, and workflows, but they also create a new place for malicious behavior to hide. A skill may look benign in its documentation or code while becoming harmful only when it is invoked with particular user requests, local assets, persistent state, or multi-step tool interactions. This makes purely static vetting brittle. We present Runtime Skill Audit (RSA), a dynamic analysis method that audits skills by asking what the skill-mediated agent actually does under targeted runtime conditions. Instead of testing every skill with the same generic tasks, RSA profiles risk-relevant interfaces, prepares the execution context needed to exercise them, and assigns security labels from the resulting trace evidence. We instantiate RSA on OpenClaw and evaluate it on 100 skills against representative static baselines. RSA achieves 90.0% accuracy with an 88.0% true positive rate and an 8.0% false positive rate, improving accuracy by 13.0 percentage points over the best static baseline. Under self-evolving attacks, static detectors collapse after one or two rounds, while RSA continues to detect 19--20 out of 20 malicious skills across rounds.
Agent skills extend LLM agents with privileged third-party capabilities such as filesystem access, credentials, network calls, and shell execution. Existing safety work catches malicious prompts and risky runtime actions, but the skill artifact itself goes unverified. We formalize this as the behavioral integrity verification (BIV) problem: a typed set comparison between declared and actual capabilities over a shared taxonomy that bridges code, instructions, and metadata. The BIV framework instantiates this comparison by pairing deterministic code analysis with LLM-assisted capability extraction. The resulting structured evidence supports three downstream analyses: deviation taxonomy, root-cause classification, and malicious-skill detection. On 49,943 skills from the OpenClaw registry, the deviation taxonomy reveals a pervasive description-implementation gap: 80.0% of skills deviate from declared behavior, with four novel compound-threat categories surfaced. Root-cause classification finds that deviations are mostly oversight, not malice: 81.1% trace to developer oversight and 18.9% to adversarial intent, with 5.0% of skills carrying predicted multi-stage attack chains. On a 906-skill malicious-skill detection benchmark, BIV reaches an F1 of 0.946, outperforming state-of-the-art rule-based and single-pass LLM baselines. These results demonstrate behavioral integrity auditing for agent skills at scale.