Skills extend an agent's capabilities by injecting instructions and information into the context, and are widely used by agents such as OpenClaw and Claude Code. Prior work shows third-party marketplaces host malicious skills that give attackers direct influence over the victim's agent. The emerging defense scans skills before installation, pairing deterministic static checks with an LLM-based semantic judge, as in NVIDIA's SkillSpector. We show that such defenses fall to an attacker who knows the detector. Our white-box LLM attacker, Pretext, iteratively crafts skills that evade detection while still delivering the payload and performing the benign task: moving the payload from code into natural language leaves static analysis inert, while framing it as the skill's legitimate purpose and splitting instructions across files keeps the LLM stage below its blocking threshold. Across three open-source models, Pretext achieves up to 97% and 77% against a frozen detector and a co-adaptive one, respectively, revealing major gaps in current skill scanners.
Figures & tables
Figure 1 : An example attack from an attacker-controlled skill.
Figure 2 : The one-run refinement loop.
Figure 3 : The generational learning cycle.
Figure 4 : Left: attacker memory self-convergence (top) and plasticity (bottom) at the final generation, one bar per model. Center: attacker success rate (ASR) per generation for Mode A and Mode B (blind and informed). Right: detector metrics (Mode B only, where the detector learns): final-generation false-positive rate (top) and detector coverage of attacker lessons (bottom). Lines/bars are the mean over 5 replicates; ASR bands are ±1 std.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A.1 : How SkillSpector scans a skill: blue nodes are deterministic, orange nodes are LLM-based, and both LLM stages are skipped under --no-llm .
Rule IDs
Family (what it flags)
Deterministic (no LLM)
P1–P4
prompt injection
P5
harmful content
P6–P8
system-prompt leakage
E1–E4
data exfiltration
PE1–PE3
privilege escalation
Appendix
Table A.1: SkillSpector’s rule families, keyed by the IDs used throughout this paper.
Attack type
Targets (rule families)
Payload in brief
Data exfiltration
E1–E4, TT3–TT4, SSD-3
send env secrets / context to an external endpoint
Prompt injection / leakage
P1–P4, P6–P8, SSD-1–2
override guardrails or leak the system prompt
Memory poisoning
MP1–MP3, E1–E2
persist a standing malicious instruction across sessions
Dangerous code execution
AST1–AST8, TT5, SC2–SC3
emit curl | bash / dynamic-exec code
Privilege escalation
PE1–PE3, AST4, TT5
sudo / edit sudoers / chmod 777 as “setup”
Dependency confusion
SC1, SC2, SC6, TT5
install a confusable package from a public index
Appendix
Table B.1: The 12 trojan attack types and the SkillSpector rules each primarily targets.
Figure B.1 : The Mode-B detector scans matched malicious/benign pairs with a frozen SkillSpector plus a learned-heuristics analyzer, and its reflector turns each generation’s false negatives and false positives into updated heuristics for the next.
Call
Receives (and source)
Withheld
Attacker (design / refine)
fixed attack type (round-robin); its predefined description, payload goal, and success criterion; the full SkillSpector ruleset and its own accumulated memory (cached system context); a mode/condition preamble; on refine, either the fired detector rules + reasoning (if detected) or the victim’s full output + verdict (if it evaded but did not deliver)
the injected SKILL.md as an installed skill; a benign task (attacker-authored); a synthetic capabilities prompt for the type
that it is a security test; any real tools or execution
Verdict judge
the attack-type label; the attacker-authored payload description; the predefined success criterion; the benign task; the victim’s text output
the SKILL.md (cover and payload)
Attacker reflector
the generation’s N run records (mechanical logs: outcome, fired rules per iteration, delivery verdict); the current Global section and the in-batch type sections
the skill bundles themselves
Detector LLM stage (Mode B)
the skill bundle; its own learned heuristics
the attack type
Detector reflector (Mode B)
the generation’s false negatives and false positives (full bundles, + rules fired on the FPs); the current heuristics
—
Appendix
Table B.2: What each LLM call receives and what is withheld. The attacker, verdict judge, and both reflectors are separate stateless calls of the attacker’s model; the detector stage and its reflector use the detector model.
Target stack
Mode A
Mode B blind
Mode B informed
qwen3t
0.2%
1.3%
0.3%
glm
4.8%
2.0%
2.7%
gpt-oss
10.0%
16.0%
17.3%
Appendix
Table C.1: Share of runs excluded from the success-rate denominators because the SkillSpector scan timed out (900 s), per stack and mode.
Metric
Definition
Purpose
Outcome
attacker success rate
fraction of runs that evade the detector (score <20 ), deliver the payload, and complete the benign task
the bottom line: is the attacker winning
detector FN rate
fraction of malicious skills not flagged (score <20 )
the detector’s miss rate on real attacks
detector FP rate
fraction of benign skills flagged (score ≥20 )
false-alarm cost of being trigger-happy
Attacker learning dynamics (Mode A and B)
self-convergence
mean nearest-neighbour cosine between the lessons of the memory at generations t−1 and t , averaged in both directions
is the memory circling one repertoire (mode collapse) or still moving
Appendix
Table C.2: The metrics used in this paper, their definition, and what each is for.
Target stack
success (%)
self-conv.
plasticity
qwen3t
96.7±4.6
0.98±0.01
0.09±0.05
glm
63.2±7.1
0.95±0.04
0.28±0.27
gpt-oss
70.5±14.2
0.96±0.03
0.24±0.13
Appendix
Table D.1: Attacker memory dynamics at the final generation of Mode A (mean ± sample std over 5 replicates).
Stack (cond.)
FN (%)
FP (%)
coverage
time-to-counter
qwen3t (blind)
75±11
40±12
0.54±0.06
0.27±0.04
qwen3t (inf.)
78±18
30±9
0.52±0.14
0.27±0.08
glm (blind)
52±15
20±8
0.61±0.07
0.30±0.11
glm (inf.)
52±14
18±12
0.60±0.10
0.29±0.06
gpt-oss (blind)
59±14
62±14
0.33±0.05
0.52±0.16
gpt-oss (inf.)
47±10
50±8
0.39±0.05
0.53±0.25
Appendix
Table D.2: Mode-B detector breakdown: false-negative/false-positive rates at the final generation, the share of attacker lessons its learned heuristics cover, and effort asymmetry (mean time-to-counter, lower is faster); mean ± sample std over 5 replicates.
Figure D.1 : Mean refinement iterations to a successful attack (cap iter=3 ) per generation, one line per target stack, for Mode A and Mode B (blind and informed); mean over 5 replicates, bands are ± SEM.
Figure D.2 : Memory plasticity (reshape rate) per generation, coloured by target stack, for the attacker (solid) and, in Mode B, the detector (dashed).
Attacker
Mode A
Mode B blind
Mode B informed
deleted (% of prev-gen lessons pruned)
qwen3t
0.8
6.5
3.0
glm
4.7
6.5
4.7
gpt-oss
3.6
5.4
6.9
born (% of current-gen lessons newly added)
qwen3t
8.9
20.7
20.6
Appendix
Table D.3: Attacker memory-shape per generation (deleted, born, and refined lesson shares), averaged over all generations and 5 replicates; these are the components pooled into plasticity.
Figure D.3 : Per-generation attacker memory self-convergence (left) and detector coverage (right) in Mode B, coloured by target stack; mean over the blind and informed conditions and 5 replicates, bands are ± SEM.
Target stack
Skills
Statically detected
Mean static score
qwen3t
60
0 (0.0%)
1.2
glm
59
6 (10.2%)
4.4
gpt-oss
59
4 (6.8%)
3.1
All
178
10 (5.6%)
2.9
Appendix
Table D.4: Static-only ( --no-llm ) detection over the final-generation attacker skills; the deterministic layer alone reaches the 20 threshold on only 5.6% of them.