As LLM-based agents perform increasingly complex tasks, Agent Skills have emerged as a flexible mechanism for extending their capabilities. An Agent Skill packages task-specific instructions with executable components and auxiliary resources to provide specialized functionalities. However, the growing adoption of third-party Skills introduces a new supply-chain attack surface. Malicious Skills can embed harmful behaviors that abuse agent privileges and compromise the agent execution environment or accessible resources. Although recent LLM-based malicious Skill auditing approaches have achieved promising performance, they often rely on capable commercial LLMs. How to achieve effective auditing with compact, locally deployable LLMs in security-sensitive and resource-constrained settings remains largely unexplored. Our investigation reveals that compact LLMs struggle to identify malicious behaviors hidden in complex Skill packages. This difficulty arises from both the implicit nature of such behaviors and the limited reasoning capacity of compact LLMs. To address these challenges, we propose SKILLLITE, an evidence-guided agentic framework for malicious Skill detection. SKILLLITE effectively extracts security-relevant behaviors and infers the intended functionality from complex Skill packages. It then employs a compact LLM to assess the maliciousness of the Skill based on the observed behaviors and their functional context. Experiments show that SKILLLITE improves malicious Skill detection across different compact LLM backbones and outperforms existing representative auditing baselines. Its effectiveness generalizes to behaviorally confirmed in-the-wild malicious Skills. Meanwhile, SKILLLITE maintains a low inference latency, supporting its practical deployment.
Figures & tables
Figure 1: Motivation and design of SkillLite for malicious Skill auditing with compact LLMs.
Figure 2: Overview of SkillLite .
Dataset
Method
Acc
Prec
Recall
F1
FPR ↓
FNR ↓
Latency (s) ↓
MalSkillBench
Cisco SkillScanner (Static)
0.571
0.679
0.257
0.373
0.120
0.743
0.1
Cisco SkillScanner (LLM)
0.722
0.661
0.907
0.764
0.460
0.093
26.6
NVIDIA SkillSpector (Static)
0.637
0.695
0.478
0.567
0.207
0.522
0.3
NVIDIA SkillSpector (LLM)
0.683
0.698
0.638
0.667
0.272
0.362
12.4
SkillWard
0.801
0.988
0.607
0.752
0.008
0.393
33.2
Skill-Vetter
0.684
0.630
0.882
0.735
0.511
0.118
35.2
Table 1: Main results on MalSkillBench and SkillTrustBench. Best and second-best detection results are shown in bold and underlined, respectively.
Method
Acc
Prec
Recall
F1
FPR ↓
FNR ↓
Latency (s) ↓
Cisco SkillScanner (Static)
0.651
0.458
0.072
0.122
0.044
0.928
0.1
Cisco SkillScanner (LLM)
0.763
0.618
0.815
0.703
0.264
0.185
20.1
NVIDIA SkillSpector (Static)
0.621
0.379
0.159
0.224
0.137
0.841
0.4
NVIDIA SkillSpector (LLM)
0.592
0.347
0.209
0.262
0.207
0.791
6.8
SkillWard
0.678
0.917
0.071
0.131
0.003
0.929
18.6
Skill-Vetter
0.807
0.664
0.892
0.761
0.238
0.108
31.7
Table 2: Generalization results on MaliciousAgentSkillsBench. Best and second-best detection results are shown in bold and underlined, respectively.
Figure 3: Effectiveness–efficiency comparison between direct zero-shot auditing and SkillLite across compact LLM backbones. Each connected pair represents the same underlying model.
Variant
Acc
Prec
Recall
F1
FPR ↓
FNR ↓
Full SkillLite
0.957
0.971
0.961
0.966
0.050
0.039
w/o Security Evidence Extraction
0.560
0.944
0.326
0.485
0.034
0.674
w/o Intent Analysis
0.549
0.996
0.292
0.451
0.002
0.708
w/o Evidence Synthesis
0.731
0.945
0.612
0.743
0.062
0.388
Table 3: Ablation study on the full SkillTrustBench binary benchmark.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Analyzer
Evidence Target
Representative Signals
General Security Patterns
Common security-sensitive operations
Local file access, network activity, system command execution, and permission modification.
Language-Aware Code Analysis
Security-sensitive implementation behavior
Process invocation, network API usage, environment-variable access, and dynamic execution.
Approval bypass, sandbox bypass, control-flow hijacking, and autonomous confirmation bypass.
Appendix
Table 4: Security evidence coverage of the deterministic analyzers in SkillLite .
Component
Configuration
OS
Ubuntu 22.04
CPU
2 × AMD EPYC 7543
RAM
250 GB
GPU
NVIDIA RTX A6000 48 GB
Python
3.11.4
Ollama
0.20.7
Appendix
Table 5: Experimental environment.
Figure 4: Recall across fine-grained attack categories on the three benchmarks. Rows denote the native malicious behavior, columns correspond to the evaluated detection methods. Each cell reports category-level recall.
Figure 5: Detection performance across Skill package sizes.
Dataset
Group
C-S
C-L
SS-S
SS-L
Ward
Vetter
AIG
SkillLite
MalSkillBench
No detected code
.171
.657
.310
.467
.679
.591
.802
.819
One language
.246
.717
.428
.544
.670
.664
.823
.911
Multiple languages
.484
.830
.683
.771
.816
.827
.880
.919
SkillTrustBench
No detected code
.364
.372
.571
.750
.545
.279
.571
.800
One language
.692
.773
.736
.716
.848
.769
.827
.952
Multiple languages
.819
.892
.884
.883
.912
.882
.901
.970
Appendix
Table 6: F1-score across packages with different programming-language complexity.
Dataset
Language
C-S
C-L
SS-S
SS-L
Ward
Vetter
AIG
SkillLite
MalSkillBench
Go
.483
.905
.629
.744
.750
.756
.850
.919
JavaScript
.331
.551
.448
.514
.769
.528
.665
.845
PowerShell
.754
.755
.810
.857
.789
.750
.839
.874
Python
.476
.883
.723
.810
.843
.880
.916
.934
Rust
.378
.852
.609
.821
.857
.828
.877
.885
SQL
.611
.911
.570
.689
.688
.900
.902
.939
Appendix
Table 7: F1-score across programming languages. Language groups are non-exclusive, as a Skill may contain multiple programming languages.
Model
Method
Acc
Prec
Recall
F1
FPR ↓
FNR ↓
Latency (s) ↓
Gemma4 E4B
Zero-shot
0.737
0.992
0.474
0.641
0.004
0.526
10.2
SkillLite
0.909
0.944
0.869
0.905
0.051
0.132
27.4
Qwen3.5 9B
Zero-shot
0.811
0.990
0.626
0.767
0.007
0.374
31.1
SkillLite
0.918
0.969
0.862
0.912
0.027
0.138
62.1
DeepSeek-R1 8B (Thinking)
Zero-shot
0.734
0.972
0.479
0.641
0.014
0.521
11.7
SkillLite
0.806
0.925
0.664
0.773
0.054
0.336
21.8
Appendix
Table 8: Complete cross-model results on MalSkillBench.
Model
Method
Acc
Prec
Recall
F1
FPR ↓
FNR ↓
Latency (s) ↓
Gemma4 E4B
Zero-shot
0.712
0.873
0.640
0.739
0.163
0.360
11.8
SkillLite
0.957
0.971
0.961
0.966
0.050
0.039
28.5
Qwen3.5 9B
Zero-shot
0.858
0.996
0.779
0.875
0.005
0.221
36.3
SkillLite
0.959
0.982
0.952
0.967
0.030
0.048
65.7
DeepSeek-R1 8B (Thinking)
Zero-shot
0.637
0.990
0.432
0.602
0.007
0.568
15.2
SkillLite
0.792
0.945
0.715
0.814
0.072
0.285
22.9
Appendix
Table 9: Complete cross-model results on SkillTrustBench.
Agent skills extend LLM agents with privileged third-party capabilities such as filesystem access, credentials, network calls, and shell execution. Existing safety work catches malicious prompts and risky runtime actions, but the skill artifact itself goes unverified. We formalize this as the behavioral integrity verification (BIV) problem: a typed set comparison between declared and actual capabilities over a shared taxonomy that bridges code, instructions, and metadata. The BIV framework instantiates this comparison by pairing deterministic code analysis with LLM-assisted capability extraction. The resulting structured evidence supports three downstream analyses: deviation taxonomy, root-cause classification, and malicious-skill detection. On 49,943 skills from the OpenClaw registry, the deviation taxonomy reveals a pervasive description-implementation gap: 80.0% of skills deviate from declared behavior, with four novel compound-threat categories surfaced. Root-cause classification finds that deviations are mostly oversight, not malice: 81.1% trace to developer oversight and 18.9% to adversarial intent, with 5.0% of skills carrying predicted multi-stage attack chains. On a 906-skill malicious-skill detection benchmark, BIV reaches an F1 of 0.946, outperforming state-of-the-art rule-based and single-pass LLM baselines. These results demonstrate behavioral integrity auditing for agent skills at scale.
LLM agents increasingly load skills, file-based packages of natural-language instructions written by third parties and distributed through marketplaces, that execute with the user's privileges. A single malicious skill can exfiltrate data, hijack the agent, or persist as a supply-chain foothold, which turns the skill marketplace into a new attack surface for agentic systems. Prompt-injection defenses do not carry over to this setting. They rely on a boundary between trusted instructions and untrusted data, whereas a skill is itself a body of instructions, so an injected command sits among many legitimate ones and inherits their authority. We present Locate-and-Judge, a two-stage detector designed for this regime. A lightweight locator scores the structural spans of a skill by the instruction-following attention each span draws and retains only the top-K. A judge then examines the retained spans in detail. Concentrating the costly judgment on a few high-attention spans lets the detector audit an entire marketplace instead of a sample. Compared to direct LLM-based scanning, this approach offers an order-of-magnitude cost reduction, dramatically increasing its scalability at a small cost to recall, and it dominates keyword and regex baselines at comparable expense. Deployed at marketplace scale and at negligible cost, Locate-and-Judge flags skills with high precision, the majority of which we manually confirmed as malicious, surfacing dozens of live malicious skills, including several disguised as benign functionality and many that SkillSpector and Cisco Skill Scanner fail to detect. We release the resulting labeled dataset.
Bacem Etteib, Daniele Lunghi, Tégawendé F. Bissyandé
Agent skills let LLM agents reuse instructions, resources, tools, and workflows, but they also create a new place for malicious behavior to hide. A skill may look benign in its documentation or code while becoming harmful only when it is invoked with particular user requests, local assets, persistent state, or multi-step tool interactions. This makes purely static vetting brittle. We present Runtime Skill Audit (RSA), a dynamic analysis method that audits skills by asking what the skill-mediated agent actually does under targeted runtime conditions. Instead of testing every skill with the same generic tasks, RSA profiles risk-relevant interfaces, prepares the execution context needed to exercise them, and assigns security labels from the resulting trace evidence. We instantiate RSA on OpenClaw and evaluate it on 100 skills against representative static baselines. RSA achieves 90.0% accuracy with an 88.0% true positive rate and an 8.0% false positive rate, improving accuracy by 13.0 percentage points over the best static baseline. Under self-evolving attacks, static detectors collapse after one or two rounds, while RSA continues to detect 19--20 out of 20 malicious skills across rounds.