Organizations: Max Planck Institute for Intelligent Systems; Tel Aviv University · Sapienza University of Rome; University of Cagliari · ELLIS Institute Tübingen
AI agents increasingly rely on modular third-party "skills" that are dynamically selected by skill routers to execute complex tasks. While recent studies highlight the threat of prompt injections embedded in these skills, existing evaluations often assume settings where the malicious skill is already selected for execution. We show that this assumption can substantially overestimate attack success. In realistic multi-skill environments, injected skills must first compete for retrieval, reducing the effective attack success rate (ASR) of existing injections by 87-97%. To address this limitation, we introduce CORSA (Cluster Optimization for Router-Aware Skill Attacks), a router-aware attack that optimizes skill injections for both retrieval and execution across clusters of related tasks. We evaluate skill injection attacks under router-managed multi-skill settings by extending the benchmark introduced by SkillRouter with eight malicious payload categories. CORSA uses successive optimization stages to first improve retrieval and then optimize end-to-end attack success, while we evaluate user utility and injection naturalism separately. Our experiments show that CORSA substantially improves both retrieval and end-to-end attack success over existing skill injections while preserving user utility, and that the resulting attacks transfer across different router architectures and LLM backbones.
Figures & tables
Figure 1: Skill injection under realistic routing. Prior evaluations assume that the poisoned skill is always executed by the agent. In a realistic setting, the skill must first be retrieved before its payload can execute. Our setting considers both retrieval and execution.
Figure 2: Overview of CORSA . Given related tasks, Stage A optimizes an injected skill for retrieval, while Stage B optimizes end-to-end success, requiring both retrieval and payload execution.
Figure 3: Main results across eight attack payloads and eight task clusters. We compare benign skills, the router-agnostic SkillJect baseline, and CORSA with GPT-5.4 on retrieval (Hit@1), user utility, end-to-end attack success (ASR), and naturalism.
Test Set
Method
Hit@1
ASR
Util
In-domain
SkillJect
11.0
10.0
38.3
CORSA (ours)
32.3
22.0
38.8
Paraphrase
SkillJect
4.0
1.0
25.0
CORSA (ours)
30.0
11.5
15.8
Synthetic
SkillJect
0.2
0.2
0.0
CORSA (ours)
13.9
5.4
11.8
Table 1: Generalization to task variations. In-domain is the original optimization tasks; Paraphrase and Synthetic use the frozen optimized skill without re-optimization. All numbers are averaged over the 8 payloads and reported in %.
Attacker
Method
Hit@1
GPT-5.4
DS-V4-Pro
GLM-5.3
GLM-5.3-Flash
Qwen3.8-27B
GPT-5.4
SkillJect
11.0
9.8 / 38.1
9.1 / 38.3
6.4 / 54.9
7.6 / 71.2
4.1 / 14.3
CORSA
32.3
22.5 / 39.0
19.4 / 49.4
13.4 / 58.0
17.1 / 71.6
6.2 / 10.4
DS-V4-Pro
SkillJect
11.5
8.9 / 36.3
8.6 / 38.7
5.1 / 61.1
6.7 / 65.0
4.6 / 15.1
CORSA
30.3
17.6 / 39.2
17.3 / 49.9
13.3 / 63.2
15.7 / 66.7
6.7 / 17.1
Table 2: Cross-model transfer of optimized skill-injection attacks. Rows identify the attacker and method; victim-model cells report ASR / Util. Skills are evaluated without further optimization. Hit@1 is shared across victims because the router is fixed. Values are percentages averaged over eight payloads.
Scaffold
Method
Hit@1
ASR
Util
Codex (source)
SkillJect
11.0
10.0
38.3
CORSA (ours)
32.3
22.0
38.8
OpenHands (transfer)
SkillJect
11.1
2.5
33.3
CORSA (ours)
32.0
16.2
34.6
Table 3: Scaffold transferability from Codex to OpenHands. Skills optimized using Codex are held fixed and evaluated on OpenHands without further optimization. We report Hit@1, ASR, and Utility (Util) in %, averaged across all eight payloads.
Optimized on
Metric
SR-0.6B
BM25
OAI-Emb-3L
R3-0.6B
Qwen3-8B
SR-0.6B
Hit@1
38.3
37.8
16.4
33.3
9.8
ASR
23.3
26.4
13.0
15.9
4.9
User utility
17.9
21.8
26.2
30.2
37.5
BM25
Hit@1
37.8
51.8
20.3
38.3
14.8
ASR
15.0
31.4
15.1
27.2
13.7
User utility
44.2
50.0
29.2
30.6
16.7
Table 4: Cross-router transfer of optimized skill-injection attacks using one fixed indirect helper-script payload per optimization router. Rows indicate optimization routers and columns evaluation routers. All values are percentages; bold indicates matched-router results.
Defense
Configuration
CORSA (ours) recall
SkillJect recall
FPR
SR
BM25
OAI
Cisco Skill Scanner
Static
0/8
0/8
0/8
0/8
2.35%
Qwen3.8-27B †
6/8
7/8
5/8
6/8
2.40%
GPT-5.4
0/8
1/8
1/8
1/8
3.15%
NVIDIA SkillSpector
Static
0/8
0/8
0/8
0/8
2.80%
Qwen3.8-27B
6/8
7/8
6/8
7/8
4.20%
Table 5: Pre-retrieval detection for a single payload. Recall is reported as packages flagged out of eight (one per task cluster) for CORSA under SkillRouter (SR), BM25, and OAI-Emb-3L (OAI), and for SkillJect . Scanners inspect complete packages; FPR uses the same 2,000 benign packages across attack methods. † Five invalid outputs treated as not flagged.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Hit@1 (%)
Cluster
Method
Exfil.
Ransom.
Destr.
Phish.
DoS
Backdoor
Poison.
Bias
Avg.
coding
SkillJect
4.2
8.1
11.8
12.4
15.1
15.2
16.8
19.2
12.9
CORSA (ours)
44.4
33.3
44.4
44.4
44.4
33.3
22.2
44.4
38.9
data-sci
SkillJect
1.3
4.0
7.1
7.0
9.2
9.4
11.0
13.0
7.8
CORSA (ours)
7.7
15.4
30.8
15.4
23.1
15.4
7.7
23.1
17.3
document
SkillJect
6.8
12.3
16.0
15.8
18.4
19.1
20.2
22.4
16.4
Appendix
Table 6: Per-cluster comparison of our method and SkillJect on GPT-5.4 across the eight payloads. We report Hit@1 and ASR in %. The final column ( Avg. ) reports the average across the eight payloads for each cluster, while the final block ( Weighted Avg. ) reports the task-weighted average across the eight clusters for each payload. For each cluster and payload, the higher value between the two methods is shown in bold .
Hit@1 (%)
Cluster
Method
Exfil.
Ransom.
Destr.
Phish.
DoS
Backdoor
Poison.
Bias
Avg.
coding
SkillJect
4.8
8.4
12.6
13.0
15.7
16.2
17.8
19.0
13.4
CORSA (ours)
33.3
33.3
22.2
33.3
66.7
44.4
44.4
44.4
40.2
data-sci
SkillJect
1.5
4.2
7.5
7.2
9.6
9.8
11.5
12.5
8.0
CORSA (ours)
15.4
15.4
23.1
15.4
30.8
23.1
23.1
15.4
20.2
document
SkillJect
7.0
12.7
16.7
16.2
19.1
19.7
21.0
22.0
16.8
Appendix
Table 7: Per-cluster comparison of our method and SkillJect on DeepSeek-V4-Pro across the eight payloads. We report Hit@1 and ASR in %. The final column ( Avg. ) reports the average across the eight payloads for each cluster, while the final block ( Weighted Avg. ) reports the task-weighted average across the eight clusters for each payload. For each cluster and payload, the higher value between the two methods is shown in bold .
Setting
Method
Hit@1
ASR
Util
Original skill set
SkillJect
11.0
10.0
38.3
CORSA (ours)
32.3
22.0
38.8
Distractor skill set
SkillJect
16.0
15.0
32.0
CORSA (ours)
33.3
19.9
33.0
Appendix
Table 8: Attack performance under normal and distractor skill routing. We report Hit@1, ASR, and Utility (Util) in %, averaged across all eight payloads.
Figure 4: Effect of optimization granularity and staging. We compare Task-Wise Optimization, Joint Optimization, and CORSA across the eight payloads. Results are reported for Hit@1, ASR, and user utility using GPT-5.4.
Figure 5: Effect of optimizing for stealthiness. We compare CORSA with a variant that includes stealthiness as an additional optimization reward. Adding the reward improves Hit@1 but reduces ASR and user utility. Results are averaged across the eight payloads using GPT-5.4.