As LLM-based agents perform increasingly complex tasks, Agent Skills have emerged as a flexible mechanism for extending their capabilities. An Agent Skill packages task-specific instructions with executable components and auxiliary resources to provide specialized functionalities. However, the growing adoption of third-party Skills introduces a new supply-chain attack surface. Malicious Skills can embed harmful behaviors that abuse agent privileges and compromise the agent execution environment or accessible resources. Although recent LLM-based malicious Skill auditing approaches have achieved promising performance, they often rely on capable commercial LLMs. How to achieve effective auditing with compact, locally deployable LLMs in security-sensitive and resource-constrained settings remains largely unexplored. Our investigation reveals that compact LLMs struggle to identify malicious behaviors hidden in complex Skill packages. This difficulty arises from both the implicit nature of such behaviors and the limited reasoning capacity of compact LLMs. To address these challenges, we propose SKILLLITE, an evidence-guided agentic framework for malicious Skill detection. SKILLLITE effectively extracts security-relevant behaviors and infers the intended functionality from complex Skill packages. It then employs a compact LLM to assess the maliciousness of the Skill based on the observed behaviors and their functional context. Experiments show that SKILLLITE improves malicious Skill detection across different compact LLM backbones and outperforms existing representative auditing baselines. Its effectiveness generalizes to behaviorally confirmed in-the-wild malicious Skills. Meanwhile, SKILLLITE maintains a low inference latency, supporting its practical deployment.
Figures & tables
Figure 1: Motivation and design of SkillLite for malicious Skill auditing with compact LLMs.
Figure 2: Overview of SkillLite .
Dataset
Method
Acc
Prec
Recall
F1
FPR ↓
FNR ↓
Latency (s) ↓
MalSkillBench
Cisco SkillScanner (Static)
0.571
0.679
0.257
0.373
0.120
0.743
0.1
Cisco SkillScanner (LLM)
0.722
0.661
0.907
0.764
0.460
0.093
26.6
NVIDIA SkillSpector (Static)
0.637
0.695
0.478
0.567
0.207
0.522
0.3
NVIDIA SkillSpector (LLM)
0.683
0.698
0.638
0.667
0.272
0.362
12.4
SkillWard
0.801
0.988
0.607
0.752
0.008
0.393
33.2
Skill-Vetter
0.684
0.630
0.882
0.735
0.511
0.118
35.2
Table 1: Main results on MalSkillBench and SkillTrustBench. Best and second-best detection results are shown in bold and underlined, respectively.
Method
Acc
Prec
Recall
F1
FPR ↓
FNR ↓
Latency (s) ↓
Cisco SkillScanner (Static)
0.651
0.458
0.072
0.122
0.044
0.928
0.1
Cisco SkillScanner (LLM)
0.763
0.618
0.815
0.703
0.264
0.185
20.1
NVIDIA SkillSpector (Static)
0.621
0.379
0.159
0.224
0.137
0.841
0.4
NVIDIA SkillSpector (LLM)
0.592
0.347
0.209
0.262
0.207
0.791
6.8
SkillWard
0.678
0.917
0.071
0.131
0.003
0.929
18.6
Skill-Vetter
0.807
0.664
0.892
0.761
0.238
0.108
31.7
Table 2: Generalization results on MaliciousAgentSkillsBench. Best and second-best detection results are shown in bold and underlined, respectively.
Figure 3: Effectiveness–efficiency comparison between direct zero-shot auditing and SkillLite across compact LLM backbones. Each connected pair represents the same underlying model.
Variant
Acc
Prec
Recall
F1
FPR ↓
FNR ↓
Full SkillLite
0.957
0.971
0.961
0.966
0.050
0.039
w/o Security Evidence Extraction
0.560
0.944
0.326
0.485
0.034
0.674
w/o Intent Analysis
0.549
0.996
0.292
0.451
0.002
0.708
w/o Evidence Synthesis
0.731
0.945
0.612
0.743
0.062
0.388
Table 3: Ablation study on the full SkillTrustBench binary benchmark.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Analyzer
Evidence Target
Representative Signals
General Security Patterns
Common security-sensitive operations
Local file access, network activity, system command execution, and permission modification.
Language-Aware Code Analysis
Security-sensitive implementation behavior
Process invocation, network API usage, environment-variable access, and dynamic execution.
Approval bypass, sandbox bypass, control-flow hijacking, and autonomous confirmation bypass.
Appendix
Table 4: Security evidence coverage of the deterministic analyzers in SkillLite .
Component
Configuration
OS
Ubuntu 22.04
CPU
2 × AMD EPYC 7543
RAM
250 GB
GPU
NVIDIA RTX A6000 48 GB
Python
3.11.4
Ollama
0.20.7
Appendix
Table 5: Experimental environment.
Figure 4: Recall across fine-grained attack categories on the three benchmarks. Rows denote the native malicious behavior, columns correspond to the evaluated detection methods. Each cell reports category-level recall.
Figure 5: Detection performance across Skill package sizes.
Dataset
Group
C-S
C-L
SS-S
SS-L
Ward
Vetter
AIG
SkillLite
MalSkillBench
No detected code
.171
.657
.310
.467
.679
.591
.802
.819
One language
.246
.717
.428
.544
.670
.664
.823
.911
Multiple languages
.484
.830
.683
.771
.816
.827
.880
.919
SkillTrustBench
No detected code
.364
.372
.571
.750
.545
.279
.571
.800
One language
.692
.773
.736
.716
.848
.769
.827
.952
Multiple languages
.819
.892
.884
.883
.912
.882
.901
.970
Appendix
Table 6: F1-score across packages with different programming-language complexity.
Dataset
Language
C-S
C-L
SS-S
SS-L
Ward
Vetter
AIG
SkillLite
MalSkillBench
Go
.483
.905
.629
.744
.750
.756
.850
.919
JavaScript
.331
.551
.448
.514
.769
.528
.665
.845
PowerShell
.754
.755
.810
.857
.789
.750
.839
.874
Python
.476
.883
.723
.810
.843
.880
.916
.934
Rust
.378
.852
.609
.821
.857
.828
.877
.885
SQL
.611
.911
.570
.689
.688
.900
.902
.939
Appendix
Table 7: F1-score across programming languages. Language groups are non-exclusive, as a Skill may contain multiple programming languages.
Model
Method
Acc
Prec
Recall
F1
FPR ↓
FNR ↓
Latency (s) ↓
Gemma4 E4B
Zero-shot
0.737
0.992
0.474
0.641
0.004
0.526
10.2
SkillLite
0.909
0.944
0.869
0.905
0.051
0.132
27.4
Qwen3.5 9B
Zero-shot
0.811
0.990
0.626
0.767
0.007
0.374
31.1
SkillLite
0.918
0.969
0.862
0.912
0.027
0.138
62.1
DeepSeek-R1 8B (Thinking)
Zero-shot
0.734
0.972
0.479
0.641
0.014
0.521
11.7
SkillLite
0.806
0.925
0.664
0.773
0.054
0.336
21.8
Appendix
Table 8: Complete cross-model results on MalSkillBench.
Model
Method
Acc
Prec
Recall
F1
FPR ↓
FNR ↓
Latency (s) ↓
Gemma4 E4B
Zero-shot
0.712
0.873
0.640
0.739
0.163
0.360
11.8
SkillLite
0.957
0.971
0.961
0.966
0.050
0.039
28.5
Qwen3.5 9B
Zero-shot
0.858
0.996
0.779
0.875
0.005
0.221
36.3
SkillLite
0.959
0.982
0.952
0.967
0.030
0.048
65.7
DeepSeek-R1 8B (Thinking)
Zero-shot
0.637
0.990
0.432
0.602
0.007
0.568
15.2
SkillLite
0.792
0.945
0.715
0.814
0.072
0.285
22.9
Appendix
Table 9: Complete cross-model results on SkillTrustBench.