Organizations: Max Planck Institute for Intelligent Systems; Tel Aviv University · Sapienza University of Rome; University of Cagliari · ELLIS Institute Tübingen
AI agents increasingly rely on modular third-party "skills" that are dynamically selected by skill routers to execute complex tasks. While recent studies highlight the threat of prompt injections embedded in these skills, existing evaluations often assume settings where the malicious skill is already selected for execution. We show that this assumption can substantially overestimate attack success. In realistic multi-skill environments, injected skills must first compete for retrieval, reducing the effective attack success rate (ASR) of existing injections by 87-97%. To address this limitation, we introduce CORSA (Cluster Optimization for Router-Aware Skill Attacks), a router-aware attack that optimizes skill injections for both retrieval and execution across clusters of related tasks. We evaluate skill injection attacks under router-managed multi-skill settings by extending the benchmark introduced by SkillRouter with eight malicious payload categories. CORSA uses successive optimization stages to first improve retrieval and then optimize end-to-end attack success, while we evaluate user utility and injection naturalism separately. Our experiments show that CORSA substantially improves both retrieval and end-to-end attack success over existing skill injections while preserving user utility, and that the resulting attacks transfer across different router architectures and LLM backbones.
Figures & tables
Figure 1: Skill injection under realistic routing. Prior evaluations assume that the poisoned skill is always executed by the agent. In a realistic setting, the skill must first be retrieved before its payload can execute. Our setting considers both retrieval and execution.
Figure 2: Overview of CORSA . Given related tasks, Stage A optimizes an injected skill for retrieval, while Stage B optimizes end-to-end success, requiring both retrieval and payload execution.
Figure 3: Main results across eight attack payloads and eight task clusters. We compare benign skills, the router-agnostic SkillJect baseline, and CORSA with GPT-5.4 on retrieval (Hit@1), user utility, end-to-end attack success (ASR), and naturalism.
Test Set
Method
Hit@1
ASR
Util
In-domain
SkillJect
11.0
10.0
38.3
CORSA (ours)
32.3
22.0
38.8
Paraphrase
SkillJect
4.0
1.0
25.0
CORSA (ours)
30.0
11.5
15.8
Synthetic
SkillJect
0.2
0.2
0.0
CORSA (ours)
13.9
5.4
11.8
Table 1: Generalization to task variations. In-domain is the original optimization tasks; Paraphrase and Synthetic use the frozen optimized skill without re-optimization. All numbers are averaged over the 8 payloads and reported in %.
Attacker
Method
Hit@1
GPT-5.4
DS-V4-Pro
GLM-5.3
GLM-5.3-Flash
Qwen3.8-27B
GPT-5.4
SkillJect
11.0
9.8 / 38.1
9.1 / 38.3
6.4 / 54.9
7.6 / 71.2
4.1 / 14.3
CORSA
32.3
22.5 / 39.0
19.4 / 49.4
13.4 / 58.0
17.1 / 71.6
6.2 / 10.4
DS-V4-Pro
SkillJect
11.5
8.9 / 36.3
8.6 / 38.7
5.1 / 61.1
6.7 / 65.0
4.6 / 15.1
CORSA
30.3
17.6 / 39.2
17.3 / 49.9
13.3 / 63.2
15.7 / 66.7
6.7 / 17.1
Table 2: Cross-model transfer of optimized skill-injection attacks. Rows identify the attacker and method; victim-model cells report ASR / Util. Skills are evaluated without further optimization. Hit@1 is shared across victims because the router is fixed. Values are percentages averaged over eight payloads.
Scaffold
Method
Hit@1
ASR
Util
Codex (source)
SkillJect
11.0
10.0
38.3
CORSA (ours)
32.3
22.0
38.8
OpenHands (transfer)
SkillJect
11.1
2.5
33.3
CORSA (ours)
32.0
16.2
34.6
Table 3: Scaffold transferability from Codex to OpenHands. Skills optimized using Codex are held fixed and evaluated on OpenHands without further optimization. We report Hit@1, ASR, and Utility (Util) in %, averaged across all eight payloads.
Optimized on
Metric
SR-0.6B
BM25
OAI-Emb-3L
R3-0.6B
Qwen3-8B
SR-0.6B
Hit@1
38.3
37.8
16.4
33.3
9.8
ASR
23.3
26.4
13.0
15.9
4.9
User utility
17.9
21.8
26.2
30.2
37.5
BM25
Hit@1
37.8
51.8
20.3
38.3
14.8
ASR
15.0
31.4
15.1
27.2
13.7
User utility
44.2
50.0
29.2
30.6
16.7
Table 4: Cross-router transfer of optimized skill-injection attacks using one fixed indirect helper-script payload per optimization router. Rows indicate optimization routers and columns evaluation routers. All values are percentages; bold indicates matched-router results.
Defense
Configuration
CORSA (ours) recall
SkillJect recall
FPR
SR
BM25
OAI
Cisco Skill Scanner
Static
0/8
0/8
0/8
0/8
2.35%
Qwen3.8-27B †
6/8
7/8
5/8
6/8
2.40%
GPT-5.4
0/8
1/8
1/8
1/8
3.15%
NVIDIA SkillSpector
Static
0/8
0/8
0/8
0/8
2.80%
Qwen3.8-27B
6/8
7/8
6/8
7/8
4.20%
Table 5: Pre-retrieval detection for a single payload. Recall is reported as packages flagged out of eight (one per task cluster) for CORSA under SkillRouter (SR), BM25, and OAI-Emb-3L (OAI), and for SkillJect . Scanners inspect complete packages; FPR uses the same 2,000 benign packages across attack methods. † Five invalid outputs treated as not flagged.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Hit@1 (%)
Cluster
Method
Exfil.
Ransom.
Destr.
Phish.
DoS
Backdoor
Poison.
Bias
Avg.
coding
SkillJect
4.2
8.1
11.8
12.4
15.1
15.2
16.8
19.2
12.9
CORSA (ours)
44.4
33.3
44.4
44.4
44.4
33.3
22.2
44.4
38.9
data-sci
SkillJect
1.3
4.0
7.1
7.0
9.2
9.4
11.0
13.0
7.8
CORSA (ours)
7.7
15.4
30.8
15.4
23.1
15.4
7.7
23.1
17.3
document
SkillJect
6.8
12.3
16.0
15.8
18.4
19.1
20.2
22.4
16.4
Appendix
Table 6: Per-cluster comparison of our method and SkillJect on GPT-5.4 across the eight payloads. We report Hit@1 and ASR in %. The final column ( Avg. ) reports the average across the eight payloads for each cluster, while the final block ( Weighted Avg. ) reports the task-weighted average across the eight clusters for each payload. For each cluster and payload, the higher value between the two methods is shown in bold .
Hit@1 (%)
Cluster
Method
Exfil.
Ransom.
Destr.
Phish.
DoS
Backdoor
Poison.
Bias
Avg.
coding
SkillJect
4.8
8.4
12.6
13.0
15.7
16.2
17.8
19.0
13.4
CORSA (ours)
33.3
33.3
22.2
33.3
66.7
44.4
44.4
44.4
40.2
data-sci
SkillJect
1.5
4.2
7.5
7.2
9.6
9.8
11.5
12.5
8.0
CORSA (ours)
15.4
15.4
23.1
15.4
30.8
23.1
23.1
15.4
20.2
document
SkillJect
7.0
12.7
16.7
16.2
19.1
19.7
21.0
22.0
16.8
Appendix
Table 7: Per-cluster comparison of our method and SkillJect on DeepSeek-V4-Pro across the eight payloads. We report Hit@1 and ASR in %. The final column ( Avg. ) reports the average across the eight payloads for each cluster, while the final block ( Weighted Avg. ) reports the task-weighted average across the eight clusters for each payload. For each cluster and payload, the higher value between the two methods is shown in bold .
Setting
Method
Hit@1
ASR
Util
Original skill set
SkillJect
11.0
10.0
38.3
CORSA (ours)
32.3
22.0
38.8
Distractor skill set
SkillJect
16.0
15.0
32.0
CORSA (ours)
33.3
19.9
33.0
Appendix
Table 8: Attack performance under normal and distractor skill routing. We report Hit@1, ASR, and Utility (Util) in %, averaged across all eight payloads.
Figure 4: Effect of optimization granularity and staging. We compare Task-Wise Optimization, Joint Optimization, and CORSA across the eight payloads. Results are reported for Hit@1, ASR, and user utility using GPT-5.4.
Figure 5: Effect of optimizing for stealthiness. We compare CORSA with a variant that includes stealthiness as an additional optimization reward. Adding the reward improves Hit@1 but reduces ASR and user utility. Results are averaged across the eight payloads using GPT-5.4.
Agent skills introduce a new and more severe form of indirect injection for LLM agents: unlike traditional indirect prompt injection, attackers can hide malicious instructions inside a dense, action-oriented skill that already functions as a legitimate instruction source. We study pre-execution skill-poison detection and show that successful skill poisoning induces a structured internal effect, attention hijacking, in which response-time attention shifts from trusted context to malicious skill spans and drives harmful behavior. Motivated by this mechanism, we propose RouteGuard, a frozen-backbone detector that combines response-conditioned attention and hidden-state alignment through reliability-gated late fusion. Across both real and synthetic open-source skill benchmarks, RouteGuard is consistently the strongest or most robust detector; on the critical Skill-Inject channel slice, it reaches 0.8834 F1 and recovers 90.51% of description attacks missed by lexical screening, showing that defending against skill poisoning requires internal-signal detection rather than text-only filtering
Wenjie Xiao, Xuehai Tang, Biyu Zhou +2
University of Chinese Academy of Sciences · Institute of Information Engineering, Chinese Academy of Sciences
Agent skills provide a lightweight mechanism for extending general-purpose agents, but their open format exposes them to skill-poisoning attacks. A practically dangerous injection must stay invisible: if executing the payload derails the user's legitimate task, the resulting failure signal invites inspection of the skill. We therefore evaluate attacks by Attack Success Rate, which requires the injected payload to execute and the user's task to still pass its verifier in the same trial. Prior skill-poisoning attacks face a reliability-stealth trade-off under this lens: YAML-header injections are reliably loaded but easily inspected, whereas stealthier body injections that place explicit malicious commands in the skill prose are less reliable because out-of-context commands invite the agent's own suspicion. We introduce POISE, a position-aware attack that compresses the trigger into a single, benign-looking body instruction, placing it at a feasible position and using a context-aware generator to blend it with nearby setup or prerequisite steps. On Skill-Inject with codex+gpt-5.2, POISE achieves an 89.3% ASR, 28.0 points above a random-placement body baseline and 2.6 points above a YAML-only baseline, while retaining the stealth advantage of body placement. That stealth is the decisive margin: because legitimate skill bodies naturally require privileged tool operations, LLM scanners are hyper-sensitive, falsely flagging 74.6% of clean skills on average across four judges and both benchmarks. Blending into these false alarms, POISE causes only 5.6% of poisoned variants to gain a new high-risk alert over their clean baselines, rendering current static defenses ineffective.
Haochang Hao, Dehai Min, Zhifang Zhang +4
University of Illinois at Chicago · University of Queensland · 3Tulane University +1
LLM agents now draw on growing skill libraries to handle complex tasks. However, injecting more skills does not always improve task completion and can even degrade it. Existing methods still treat skill injection as a static step, selecting skills with fixed criteria, fixing the budget in advance, and leaving descriptions unchanged. We argue that this static treatment can undermine the utility of skills, because which skills are exposed, how many are included, and how they are presented all affect downstream performance. We propose SkillsInjector, a two-stage adaptive method that jointly addresses these decisions. First, a context planner learns execution-grounded skill preferences and admits an adaptive number of skills for each task. A set-aware renderer then tailors how selected descriptions are presented relative to their co-injected neighbors. Across tau2-bench, SkillsBench, and ALFWorld, SkillsInjector achieves the highest score, improving over the strongest baseline by 3.9, 6.1, and 7.3 percentage points, respectively. Ablation studies show that skill selection, adaptive budgeting, and set-aware rendering each contribute to the gain. These results show that skill-augmented agents benefit from optimizing the injected context itself. Code will be released upon publication