Can Agents Trust Their Skills? Uncovering Unsafe Chains of Trust in Skill-Based LLM Agents
Authors: Yan Wang, Zhihao Zhang, Ke Chen, Kai Chen, Yaqin Zhang, Duohe Ma, Jun Dai, Xiaoyan Sun
Organizations: Institute of Information Engineering, Chinese Academy of Sciences · School of Cyberspace Security, University of Chinese Academy of Sciences · Worcester Polytechnic Institute
LLM agents increasingly rely on installable skills, which are packages of instructions, code, and resources that equip them with task-specific capabilities and, once installed, can be automatically invoked across subsequent user tasks. This creates a chain of trust in which users delegate authority to agents, while agent frameworks admit skill-provided content into the agents' context with insufficient validation, allowing malicious skills to influence agent behavior under that delegated authority. Yet, little is known about whether this trust model adequately constrains untrusted skill content before it reaches security-sensitive operations, or how frequently such trust violations arise in real-world agents. We present TrustProbe, a framework for uncovering unsafe chains of trust in skill-based LLM agents. First, TrustProbe analyzes agent source code to identify source-to-sink call paths from skill-controlled inputs to security-sensitive operations. Second, it generates semantically realistic SKILL.md seeds with injected canaries and evolves them through feedback-guided scheduling and mutation. Finally, it validates vulnerabilities using an oracle that confirms attacker-controlled flows and verifies observable harm. Across 11 open-source agents, eight with more than 10,000 GitHub stars, TrustProbe identifies 104 taint-style vulnerabilities. Validation on a large corpus of real-world skills collected from public hubs such as ClawHub further shows that 25.1% of skill-agent trials exercise the identified vulnerable paths, with payload injection successfully weaponizing 15 of the vulnerabilities. These results reveal a systematic trust failure in skill-based LLM agents: untrusted skill content can reach security-sensitive operations and exercise authority delegated by users to their agents.
Figures & tables
Figure 1
OpenClaw
OpenCode
Hermes Agent
Pochi
Kimi Code CLI
Qwen Code
Cline
Pi Coding Agent
Mistral Vibe
DB-GPT
Agent Zero
Total
Stars (K)
388.9
204.5
241.7
0.1
7.3
27.7
67.5
102.0
4.9
19.9
19.1
1.08M
Verified Vulns.
24
17
16
13
12
8
7
3
2
1
1
104
Time Cost (h)
3.79
2.64
13.17
1.15
4.04
5.38
4.97
0.93
0.88
2.67
5.57
45.20
TTE (min)
10.32
42.78
12.70
11.07
70.62
23.65
92.27
23.73
26.70
2.00
255.25
23.73
Table 1 : Verified vulnerabilities, execution time, and time to first exposure across agents.
OpenClaw
OpenCode
Hermes Agent
Pochi
Kimi Code CLI
Qwen Code
Cline
Pi Coding Agent
Mistral Vibe
DB-GPT
Agent Zero
Total
Skill discovery
24
17
16
13
12
8
7
3
2
1
1
104
Prompt replay
1
7
6
9
1
4
2
2
0
0
1
33
Reproduction rate ( ρa )
4.2%
41.2%
37.5%
69.2%
8.3%
50.0%
28.6%
66.7%
0.0%
0.0%
100.0%
31.7%
Table 2 : Direct-prompt replay of vulnerabilities verified through installed skills during discovery.
Method
TP
FP
FN
Prec(%)
Recall(%)
LLMSmith
5
328
15
1.50
25.0
AgentFuzz
0
0
20
N/A
0.0
TrustProbe
20
0
0
100
100
Table 3 : Baseline comparison using the 20 verified vulnerabilities as the reference set.
Configuration
Relative yield
GenericSeed
44.2%
NoSigma
37.9%
RandomSched
31.0%
Table 4 : Relative yield of ablations within their evaluated stages.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3 : The seed-generation prompt.
Figure 4 : The semantic-scoring ( σ ) prompt.
Figure 5 : The Semantic Mutator prompt.
Figure 6 : The Metadata Mutator prompt.
Figure 7 : The Body Mutator prompt.
Figure 8 : The Sink-Argument Mutator prompt.
Agent
GitHub repository
Version
Commit
OpenClaw
https://github.com/openclaw/openclaw
2026.5.6
b70a2451f8c9
Hermes Agent
https://github.com/NousResearch/hermes-agent
0.19.0
8fc278207b0f
OpenCode
https://github.com/sst/opencode
1.17.18
8a03fc265b6d
Pi Coding Agent
https://github.com/earendil-works/pi
0.82.1
b4f293684bba
Cline
https://github.com/cline/cline
3.0.47
7d63376d9824
Qwen Code
https://github.com/QwenLM/qwen-code
0.21.0
58fa6cf85b99
Appendix
Table 5 : GitHub repositories, versions, and tested revisions of the 11 evaluated agents.
Agent
Skill-loading mechanism
Execution configuration
OpenClaw
Workspace skills/<name>/SKILL.md
Fresh configuration; no approval override
Hermes Agent
Isolated HERMES_HOME/skills
Automatic approval; headless one-shot
OpenCode
Per-run skills.paths directory
--dangerously-skip-permissions
Pi Coding Agent
Explicit --skill loading
Print/headless mode; no approval gate
Cline
Workspace .cline/skills
--auto-approve true
Qwen Code
Workspace .qwen/skills ; explicit invocation
--yolo
Appendix
Table 6 : Native skill-loading mechanisms and execution configurations in the main campaign.
Parameter
Main-experiment setting
Framework and agent backend
deepseek-v4-flash
Initialization
One generated seed per audited source-to-sink call path
Pooling and budget
One pool per sink function with eight mutation rounds or a 15-minute cap, whichever comes first
Scheduler weights
α=β=γ=η=1 and distance exponent k=1
Score ranges
Semantic score σ∈[0,10] and distance score δ∈[0,10]
Mutation thresholds
σ<6 triggers semantic mutation and δ<8 triggers stage-directed mutation after the semantic rule
Appendix
Table 7: Campaign parameters shared by all subjects in the main campaign.
Package
Class
Methods
Type
subprocess , os
–
run , Popen , exec* , spawn*
CMDi
asyncio
–
create_subprocess_shell , create_subprocess_exec
CMDi
builtins
–
eval , exec , compile
CODEi
shutil , pathlib
Path
open , copy , move
PATHi
requests , httpx , aiohttp , urllib
–, Session
get , post , request , urlopen
SSRF
yaml , pickle
–
load , loads
Deser.
Appendix
Table 8 : Python sink inventory used by the source-to-sink analysis, grouped by package.
Package or runtime
Class
Methods
Type
Node.js
child_process
exec , spawn , fork
CMDi
JavaScript
builtins, vm
eval , Function , runInThisContext
CODEi
Node.js
fs
readFile , writeFile , copyFile
PATHi
Web and Node.js
fetch , axios
get , post , request
SSRF
js-yaml , v8
–
load , deserialize
Deser.
Object , lodash
–
assign , merge
Proto.
Appendix
Table 9 : TypeScript sink inventory used by the source-to-sink analysis, grouped by package or runtime.
Agent
Restrictive configuration
Verified outcomes
OpenClaw
tools.exec with security=allowlist and ask=always , empty allowlist
22/24
Hermes Agent
Isolated configuration with approvals.deny=["*"]
1/16
OpenCode
Remove permission-skipping flag and set Bash and Edit to ask
0/17
Pi Coding Agent
Official-style extension blocks Bash, Write, and Edit in headless mode
1/3
Cline
--auto-approve false
0/7
Qwen Code
--approval-mode default
1/8
Appendix
Table 10 : Per-agent restrictive approval configurations used in the RQ1 approval replay.
OpenClaw
OpenCode
Hermes Agent
Kimi Code CLI
Qwen Code
Cline
Pi Coding Agent
Mistral Vibe
Total
Permissive
24
17
16
12
8
7
3
2
89
Restrictive
22
0
1
6
1
0
1
0
31
Residual Exploitability
91.7%
0.0%
6.3%
50.0%
12.5%
0.0%
33.3%
0.0%
34.8%
Appendix
Table 11 : Verified counts per agent under the permissive and restrictive approval configurations.
Skill ID
Agent
Technique
Outcome
S01
OpenClaw
Command
Downloads and executes an attacker script, remote code execution
S02
OpenClaw
Inline domain
The full credential pair reaches the attacker
S03
Cline
Inline URL
Agent initiates an OAuth flow using an attacker-controlled endpoint
S04
OpenClaw
Inline domain
Agent creates an API key and mailbox through the attacker
S05
OpenClaw
Inline domain
All bot authentication traffic, 20 requests, passes through the attacker
S06
Cline
Inline URL
Agent connects to the attacker over MCP and calls its tools
Appendix
Table 12 : The 15 weaponized twins, reported using anonymized skill identifiers, with their agents, techniques, and observed outcomes.
Third-party skills are becoming the package ecosystem for LLM agents. They package natural-language instructions, helper scripts, templates, documents, and service configuration into reusable workflows. This makes skills useful, but it also introduces a new security problem: a malicious skill does not need to ask the model to perform an obviously harmful action. Instead, it can disguise the harmful behavior as part of a routine workflow, relying on the agent to execute that workflow with high-value permissions and limited human supervision. We introduce AgentTrap, a dynamic benchmark for evaluating whether LLM agents can use third-party skills while resisting malicious runtime behavior. AgentTrap contains 141 tasks: 91 malicious tasks and 50 benign utility tasks, covering 16 security-impact dimensions grounded in agent-skill supply-chain threats. In each task, the agent receives an ordinary user request, runs with installed skills that may contain malicious workflow elements, and is executed in a sandboxed environment. AgentTrap then judges complete trajectories for attack success, blocked or refused behavior, attack-not-triggered cases, and no-attack-evidence outcomes. Our central finding is that the most informative failures are not simple jailbreaks. Models often complete the visible user task while treating unsafe side effects introduced by the skill as part of the normal workflow. This motivates runtime evaluation of the concrete model--framework--workspace environment in which users actually delegate work. Code and data are available at https://github.com/zhmzm/AgentTrap and https://huggingface.co/datasets/zhmzm/AgentTrap.
Haomin Zhuang, Hanwen Xing, Yujun Zhou +5
University of Notre Dame · University of Southern California · LMU Munich +1
Skills are becoming the capability layer through which LLM agents turn plans into actions, but their use introduces security risks such as data leakage, unauthorized operations, and tool misuse. Existing vetting usually evaluates each skill in isolation, while real agent tasks often invoke multiple skills in a shared execution context. This creates Skill Composition Risk (SCR): a skill that appears benign alone can become harmful when its outputs, trust signals, authorization cues, or side effects influence later invocations along an activated path. We introduce SCR-Bench to evaluate this risk in controlled, sandboxed skill environments. Rather than relying only on textual intent or surface behavior, SCR-Bench records downstream state changes and path-level outcomes across composed skill executions. It contains three sub-benchmarks: SCR-CapFlow for capability-flow composition, SCR-TrustLift for trust-transfer composition, and SCR-AuthBlur for authorization-confusion composition. Across SCR-Bench, composed paths expose risks that are largely absent under isolated evaluation. In SCR-CapFlow, attack success rate reaches 33.6 percent under composition, compared with near-zero isolated baselines. In SCR-TrustLift, attack success rate exceeds 96.5 percent on four of five backends. In SCR-AuthBlur, the risky-approval rate increases by 71.8 percent relative to the L0 isolated baseline under the L1 context setting. These results show that agent skill security should be assessed at the level of activated paths rather than isolated artifacts. SCR and SCR-Bench provide a foundation for path-aware risk evaluation and defense in LLM agent skill ecosystems. Benchmark: https://github.com/saint-viperx/SCR_Bench.
Yi Xie, Jiawei Du, Yu Cheng +2
East China Normal University, Shanghai, China · Centre for Frontier AI Research A*STAR, Singapore · Shanghai Innovation Institute, Shanghai, China
Agent skills extend LLM agents with privileged third-party capabilities such as filesystem access, credentials, network calls, and shell execution. Existing safety work catches malicious prompts and risky runtime actions, but the skill artifact itself goes unverified. We formalize this as the behavioral integrity verification (BIV) problem: a typed set comparison between declared and actual capabilities over a shared taxonomy that bridges code, instructions, and metadata. The BIV framework instantiates this comparison by pairing deterministic code analysis with LLM-assisted capability extraction. The resulting structured evidence supports three downstream analyses: deviation taxonomy, root-cause classification, and malicious-skill detection. On 49,943 skills from the OpenClaw registry, the deviation taxonomy reveals a pervasive description-implementation gap: 80.0% of skills deviate from declared behavior, with four novel compound-threat categories surfaced. Root-cause classification finds that deviations are mostly oversight, not malice: 81.1% trace to developer oversight and 18.9% to adversarial intent, with 5.0% of skills carrying predicted multi-stage attack chains. On a 906-skill malicious-skill detection benchmark, BIV reaches an F1 of 0.946, outperforming state-of-the-art rule-based and single-pass LLM baselines. These results demonstrate behavioral integrity auditing for agent skills at scale.