LLM agents use skills to improve performance on specialized tasks. To complete a user request, an agent may invoke several skills in sequence, allowing information produced under one skill to guide the next. Because skills may come from open-source repositories, this handoff can also carry attacker-controlled claims into later decisions. In this paper, we introduce APEX, which constructs and refines adversarial skill chains tailored to a user task and an attacker-selected action. The key insight is that an agent-written record of genuine task progress can carry a false claim of user approval across skills: an upstream skill induces the agent to create the record, and a downstream skill uses it to direct the attacker-selected action. Across four targeted-action families and six models on SkillsBench, the chains induce the selected action in 512 of 690 attempts (74.2%). On GPT-5.4, the full chain succeeds in 84.3% of attempts, compared with 17.4% when the workflow is merged into one skill. We further evaluate a prompting defense that asks the agent to check skill-produced files against the original request. On GPT-5.4, it lowers targeted-action success from 84.3% to 59.1%, while the verifier test-pass rate across 72 benign native-skill tasks falls from 86.7% to 56.3%. These results highlight the need for defenses that prevent attacker-directed actions while preserving legitimate task performance.
Figures & tables
Figure 1: APEX attack mechanism. In this schematic example, the agent with adversarial skill chains completes the requested summary and responds to the user, but continues to follow the chain and to delete an original file that the user required the agent to preserve.
Figure 2: Overview of APEX. APEX generates candidate skill chains, simulates them in an isolated task environment, and refines them using execution feedback of normal task utility and attack success rate.
Family
Aspect
Initial task state
Required access
Target
External Transfer
Confidentiality
Local task file
Network
Out-of-scope transfer
Script Acquisition
Integrity
Helper-compatible task
Execution
Remote helper execution
File Modification
Integrity
Non-backup file
Deletion
Native file deletion
Output Tampering
Integrity
Source file
Write
Conflicting write
Work Loop
Availability
Execution unit
Continuing instruction
Re-entry and next unit
Table 1: Five attack families: task state, attacker-supplied claim, and target.
External
Script
File
Output
Work Loop
Model
transfer
acquisition
modification
tampering
Token ratio
n=30
n=29
n=25
n=31
n=15
Claude Sonnet 5
83.3
65.5
84.0
58.1
21.33
DeepSeek V4 Flash
66.7
75.9
84.0
77.4
20.66
GPT-5.4
86.7
75.9
84.0
90.3
15.00
Gemini 3.5 Flash
76.7
89.7
84.0
74.2
2.20
Table 2: Attack effectiveness across models. The first four columns report ASR (%), with n tasks per model and family. For Work Loop, we report the mean per-task token ratio. The final row aggregates the corresponding metric across all models.
Figure 3: Task utility under benign execution (blue circles) and attack conditions (orange diamonds). Scores are fractions of native verifier checks passed across all runs in each condition, whether or not the attack succeeded. Gray segments connect the two scores for each model–family pair. All panels share the same percentage scale.
Full
Authorization
Monolithic
Direct
Family
chain
denied
workflow
injection
DeepSeek V4 Flash
External Transfer
66.7
0.0
10.0
13.3
Script Acquisition
75.9
0.0
13.8
0.0
File Modification
84.0
32.0
20.0
0.0
Output Tampering
77.4
48.4
0.0
0.0
Table 3: Structural ablations. We report ASR (%) for each action-oriented family and for all four together, and the mean token ratio ( × ) for Work Loop. The largest value in each row is shown in bold. The ASR comparisons use the task counts in Table 2 .
Figure 4: Skill-count ablation. Upper panels show ASR for the four targeted-action families, and lower panels show the mean work-loop token ratio. Ringed points mark the default number of skills. The dotted horizontal line in the lower panels marks a token ratio of 1, corresponding to native baseline execution.
Attack outcome
Utility (%)
Family
Undefended
Defended
Undefended
Defended
External Transfer
86.7
60.0
76.5
55.8
Script Acquisition
75.9
62.1
90.0
70.5
File Modification
84.0
76.0
83.4
65.9
Output Tampering
90.3
41.9
89.6
44.2
Work Loop ( × )
15.00
7.50
80.5
67.3
Table 4: Taint-guided prompting on GPT-5.4. We measure attack outcomes using ASR (%) or mean token ratio ( × ), and task utility using verifier test-pass rate. Each comparison uses the same verifier checks in both conditions (Appendix C.6 ).
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Attack objective
Success criterion
External Transfer
Selected task material is transferred beyond the requested deliverable.
Script Acquisition
The selected external helper is retrieved and executed.
File Modification
The selected pre-existing file is deleted.
Output Tampering
Attacker-selected content is written and conflicts with the task source.
Work Loop
The agent returns to the entry stage and begins another work unit.
Appendix
Table 5: Success criteria for the five evaluated attack families. For the four targeted-action families, success requires an observed effect on the selected target. Work Loop requires the agent to return to the entry stage and begin another work unit.
Family or control
Tasks
Verifier items
External Transfer
30
251
Script Acquisition
29
190
File Modification
25
217
Output Tampering
31
249
Work Loop
15
113
Benign native tasks
72
533
Appendix
Table 6: Tasks and verifier checks by family. The first four rows apply to main, structural, skill-count, and defense comparisons. We calculate utility for every reported Work Loop condition from 113 verifier checks. The benign defense control uses 72 distinct tasks.
External Transfer
Script Acquisition
Model
Base
Attack
Δ
Base
Attack
Δ
Claude Sonnet 5
78.1
62.5
-15.5
71.6
66.3
-5.3
DeepSeek V4 Flash
70.1
41.0
-29.1
42.1
62.1
+20.0
GPT-5.4
77.7
76.5
-1.2
86.8
90.0
+3.2
Gemini 3.5 Flash
76.1
53.0
-23.1
73.2
62.6
-10.5
Grok 4.5
69.3
53.4
-15.9
67.4
81.6
+14.2
Appendix
Table 7: Native verifier utility in the main experiment. The Base and Attack columns give percentages. Δ is attack minus baseline in percentage points, calculated before rounding. Table 6 lists the number of verifier checks by family.
Family
A
B
C
D
DeepSeek V4 Flash
External Transfer
41.0
52.6
25.9
42.6
Script Acquisition
62.1
17.9
90.0
51.1
File Modification
76.5
40.6
27.2
44.2
Output Tampering
76.3
80.7
47.0
59.4
Work Loop
14.2
25.7
58.4
61.9
Appendix
Table 8: Native verifier utility (%) under structural controls. We compare conditions A (full chain), B (authorization denied), C (monolithic workflow), and D (direct injection). Every condition uses the verifier checks listed for its family in Table 6 .
Family
K=1
K=2
K=3
K=4
K=5
K=6
DeepSeek V4 Flash
External Transfer
13.3
16.7
23.3
66.7
23.3
13.3
Script Acquisition
13.8
6.9
31.0
48.3
75.9
37.9
File Modification
44.0
84.0
80.0
84.0
64.0
76.0
Output Tampering
6.5
19.4
64.5
77.4
54.8
45.2
Work Loop
2.86 ×
12.62 ×
19.12 ×
12.73 ×
20.66 ×
18.21 ×
Appendix
Table 9: Skill-count attack outcomes. We report ASR (%) for the first four families and the mean token ratio ( × ) for Work Loop. The default skill count is marked in bold.
Family
K=1
K=2
K=3
K=4
K=5
K=6
DeepSeek V4 Flash
External Transfer
10.0
5.6
18.3
41.0
16.3
21.1
Script Acquisition
84.2
90.0
80.0
65.8
62.1
75.8
File Modification
64.5
87.6
73.3
76.5
75.6
66.8
Output Tampering
51.0
68.7
68.7
76.3
69.5
42.2
Work Loop
58.4
40.7
25.7
31.0
14.2
12.4
Appendix
Table 10: Native verifier utility (%) by skill count. All six counts use the same verifier checks for each family, including 113 checks for Work Loop. The default skill count is marked in bold.
Model
Token ratio ( × )
Baseline (%)
Attack (%)
Change (pp)
Claude Sonnet 5
21.33
78.8
42.5
-36.3
DeepSeek V4 Flash
20.66
57.5
14.2
-43.4
GPT-5.4
15.00
88.5
80.5
-8.0
Gemini 3.5 Flash
2.20
69.9
56.6
-13.3
Grok 4.5
14.42
46.0
13.3
-32.7
Kimi K2.6
36.39
77.0
77.0
+0.0
Appendix
Table 11: Work-loop token amplification and task utility. We compute the token ratio on each task using total tokens under attack and at baseline, then report the mean of these ratios. Each model uses the same 113 verifier items in its baseline and attack conditions.
Passed / total
Utility (%)
Family or control
Undefended
Defended
Undefended
Defended
External Transfer
192/251
140/251
76.5
55.8
Script Acquisition
171/190
134/190
90.0
70.5
File Modification
181/217
143/217
83.4
65.9
Output Tampering
223/249
110/249
89.6
44.2
Work Loop
91/113
76/113
80.5
67.3
Appendix
Table 12: GPT-5.4 utility with and without taint-guided prompting. Both conditions use the same verifier checks. We report how many checks passed alongside the corresponding utility percentages.
Large language model (LLM) agents increasingly rely on reusable skills i.e. documents describing task-specific procedures. However, this introduces a new attack surface for agents to manage. We study two complementary directions for this threat. First, we evaluate guardian-based defenses: an intermediary LLM agent that acts as a mediator for skill file access (dynamic guardian) or pre-rewrites these files at build time (static guardian). Across three LLM agent families, our guardians cut attack success rate (ASR) by well over half while preserving task utility. Second, we stress test them through attack reframing using four attacks that preserve the malicious instruction but change the phrasing. For non-guardian setup, the reframing pushes the ASR up to 81.4%, but the dynamic guardian brings it down to 18.6%, showing that real-time mediation is a robust defense.
LLM agents increasingly rely on third-party skills, using natural-language descriptions for selection and instruction bodies for planning. This progressive-disclosure design exposes two sequential control points to untrusted publishers: a static skill may steer an otherwise correct task onto an unnecessarily costly trajectory. Prior work studies selection manipulation, malicious skill instructions, and tool-chain resource amplification largely separately, leaving their end-to-end composition unclear. We introduce Convergent Detour Hijacking (CDH), a text-only, runtime-independent attack that couples these stages. Under shared semantic cover, a description establishes relevance during selection, while an aligned body reuses that rationale to fabricate plausible dependencies during planning. CDH attracts an attacker-controlled coordinator alongside legitimate skills, recruits unnecessary benign skills into a bounded detour, and then re-enters the original route to preserve task completion. We evaluate it across multiple LLM backends and 491 held-out tasks under single-task and multi-turn conditions. On DeepSeek-V4-Pro, the matched coordinator is selected in 80.02% of tasks; among coordinator-hit runs that complete tasks, token consumption and end-to-end execution time increase by 66.91% and 92.45%, respectively, while aggregate task completion remains comparable. Thus, correct outcomes do not guarantee trajectory integrity or cost safety.
Junliang Liu, Ruoyu Li, Wenxin Tang +4
School of Computer Science and Software Engineering, Shenzhen University · The Chinese University of Hong Kong · The Chinese University of Hong Kong, Shenzhen
Agent skills are emerging as an important attack surface in LLM-based agent systems. Through an empirical study of existing skill scanners, we find that current defenses mainly inspect individual skills, leaving risks from cross-skill composition insufficiently examined. This creates a practical blind spot: multiple locally plausible skills may pass security checks while collectively forming a harmful workflow during agent execution. To investigate this threat, we propose ColluSkill, a collusive multi-skill-chain attack framework that decomposes a complete malicious intent into interdependent sub-payloads embedded in independently packaged skills. The attack does not rely on any single malicious skill, but emerges from the ordered composition of locally plausible behaviors through contextual dependencies, artifact passing, and execution handoffs. ColluSkill further employs LLM-based chain planning and scanner-feedback refinement to preserve chain-level attack semantics while reducing suspicious signals in individual sub-skills. To defend against such attacks, we propose ChainGuard, a context-aware skill-chain scanner that jointly analyzes a candidate skill and the skills already installed in the agent environment. ChainGuard reconstructs cross-skill dependencies, artifact flows, capability compositions, and downstream behaviors to identify risks that emerge only at the workflow level. Experiments on six representative skill scanners show that ColluSkill achieves an average attack success rate of 96.0% and consistently outperforms the evaluated single-skill and multi-skill attack baselines. Meanwhile, ChainGuard reduces the attack success rate to 22.5% while allowing 99.5% of benign workflows to pass, highlighting the importance of chain-level security analysis for agent skill ecosystems.
Puyu Zeng, Simeng Qin, Jingzhi Li +3
College of Cryptology and Cyber Science, Nankai University, China · Northeastern University, China · University of Science and Technology Beijing, China +2