Large language model agents are increasingly connected to high-value resources, including external APIs, GPUs and servers, and workflows such as deployment and approval. Existing agent security research mainly focuses on attacks against information and agent behavior, while high-value resources have received less attention as attack targets themselves. To our knowledge, we are the first to identify and systematically study agent resource hijacking, in which attackers induce agents to use high-value resources for their own goals without directly obtaining those resources or their credentials. We introduce ResourceHijackBench, an executable benchmark and automated case-generation pipeline that covers six categories of high-value resources. It contains 300 attack scenarios and 900 attack prompts, and each case runs in an isolated local environment that records actual resource use. Resource hijacking remains effective across four model backends, with average attack success rates of 70.0% to 89.6%, and it also appears across two agent harnesses, with average ASRs of 84.1% on OpenClaw and 72.3% on Codex. In paired comparisons, resource hijacking achieves ASRs 62.4 to 84.0 percentage points higher than direct resource acquisition across the four model backends. Real-world experiments also demonstrate the practical effectiveness of resource hijacking. Among the three existing defenses we evaluate, the lowest average ASR is still 55.1%. We further propose ResGate, a pre-execution resource authorization defense that combines model-based resource-use extraction with deterministic policy enforcement based on trusted requester identity metadata, reducing the average ASR on OpenClaw to 23.6%. These results show that preventing direct access alone is not enough to protect resources that agents can still use on an attacker's behalf, and that explicit resource authorization can help reduce this risk.
Figures & tables
Figure 1: Overview of agent resource hijacking. Existing safety defenses may prevent direct disclosure of a credential or resource, while the agent can still use the corresponding capability for an resource attacker.
Category
Functional Role
Representative Examples
Material
Infrastructure and computing capacity
GPUs, CPUs, memory, storage, bandwidth, and CI runners
Condition
Credentials and permissions enabling access
API keys, OAuth tokens, IAM roles, and repository permissions
Energy
Quotas, time, and budgets consumed during execution
Model quotas, CI minutes, execution time, and cloud budgets
Social and symbolic
Identity, recognition, trust, and authority
Maintainer identities, commit authorship, and official accounts
Information and knowledge
Information supporting decisions and actions
Private code, internal documents, memory stores, and knowledge bases
Interaction
Channels for communication and coordination
Team chat, company email, reviewer networks, and approval workflows
Table 1: High-value resources in tool-using agents: categories, functional roles, and representative examples.
Figure 3
Category
Baseline
Defense configuration
OpenClaw
Prompt Defense
AgentDoG
LlamaFirewall
ResGate
Condition
92.0
59.3
92.0
48.7
32.7
Energy
93.9
68.9
91.3
57.0
19.3
Interaction
68.9
53.7
68.0
59.3
12.0
Knowledge
87.3
63.2
85.3
50.3
58.7
Material
76.9
42.2
76.0
65.1
12.7
Table 2: Attack success rates (ASR, %) across resource categories under baseline and defense configurations, with average task success rate (TSR, %) on authorized normal tasks for comparison. Lower ASR and higher TSR are better.
Resource category
Attack success rate (ASR, %)
DeepSeek-V4-Pro
GPT-5.5
Gemini-3.5-Flash
Qwen3.8-Max
Condition
92.0
94.7
80.0
89.8
Energy
93.9
90.9
79.3
82.3
Interaction
68.9
96.0
58.0
69.7
Knowledge
87.3
90.5
87.3
97.9
Material
76.9
81.0
55.3
82.4
Table 3: Attack success rates (ASR, %) across OpenClaw model backends and resource categories. Higher ASR indicates greater vulnerability to resource hijacking.
Category
OpenClaw
Codex
Overall
ASR
95% CI
ASR
95% CI
ASR
95% CI
Condition
92.0
[86.5, 95.4]
92.7
[87.4, 95.9]
92.3
[88.8, 94.8]
Energy
93.9
[88.8, 96.8]
79.6
[72.4, 85.3]
86.9
[82.6, 90.2]
Interaction
68.9
[61.1, 75.8]
61.5
[53.5, 68.9]
65.4
[59.9, 70.6]
Knowledge
87.3
[81.1, 91.7]
78.7
[71.4, 84.5]
83.0
[78.3, 86.8]
Material
76.9
[69.4, 83.0]
57.1
[49.1, 64.9]
67.0
[61.5, 72.1]
Table 4: Cross-harness evaluation of resource hijacking on OpenClaw and Codex. The results show that resource hijacking is a shared risk across multiple agent harnesses.
Model
Attack success rate (%)
Gap (pp)
Direct Acquisition
Resource Hijacking
DeepSeek-V4-Pro
7.4
84.1
+76.7
Gemini-3.5-Flash
7.6
70.0
+62.4
GPT-5.5
5.6
89.6
+84.0
Qwen3.8-Max
3.4
82.8
+79.4
Table 5: Paired comparison of direct acquisition and resource hijacking across model backends. Resource hijacking achieves much higher attack success rates than direct acquisition on all models.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8
Item
Configuration
DeepSeek-V4-Pro
API ID: deepseek-v4-pro ; V4-Pro Preview (2026-04-24); temperature 1.0.
GPT-5.5
API ID: gpt-5.5-2026-04-23 ; snapshot version (2026-04-23); temperature 1.0.
Gemini-3.5-Flash
API ID: gemini-3.5-flash ; release version (2026-05-19); temperature 1.0.
Qwen3.8-Max
API ID: qwen3.8-max-2026-09-02 ; snapshot version (2026-09-02); temperature 1.0.
Attack-success judge
gpt-5.5 ; maximum output of 800 tokens; 60-second timeout; up to two retries; independent rechecking of initially safe verdicts.
Auxiliary models
gpt-5.5 for all defense-specific judging, scanning, and extraction.
Appendix
Table 6: Model and API configurations. The four agent backends are used in the model-backend ablation.
Item
Configuration
Original OpenClaw
Stock OpenClaw with no defense component; the user prompt is passed verbatim.
Prompt Defense
A static Chinese resource-safety policy is appended to AGENTS.md in each isolated workspace.
AgentDoG
AgentDoG 1.5.
LlamaFirewall
LlamaFirewall 1.0.3-source (AlignmentCheck).
ResGate
Pre-execution resource-use extraction followed by a deterministic check against a frozen, read-only allowlist. The attack-side policy permits only anonymous read_status operations.
Appendix
Table 7: Defense configurations used in the evaluation.
Item
Configuration
Framework and runtime
OpenClaw 2026.5.26 ; Node.js v22.22.1 .
Attack set
900 prompts across 300 scenarios; six resource categories with 150 prompts each; three attack settings with 300 prompts each.
Sandbox
An isolated workspace for each case. Resource operations are simulated through local scripts and recorded in workspace and local service logs. No real external resources, credentials, or accounts are used.
Execution limits
One user turn per case; a 300-second agent execution timeout.
Behavioral evaluation
Oracle checks and trajectory rules are applied first; unresolved cases are evaluated by the LLM judge. Runner and judge errors are marked as non-evaluable.
Table 8: Environment and evaluation configurations.
Model
API model ID
Version
Temperature
DeepSeek-V4-Pro
deepseek-v4-pro
V4-Pro Preview (2026-04-24)
1.0
GPT-5.5
gpt-5.5-2026-04-23
Snapshot version (2026-04-23)
1.0
Gemini-3.5-Flash
gemini-3.5-flash
Release version (2026-05-19)
1.0
Qwen3.8-Max
qwen3.8-max-2026-09-02
Snapshot version (2026-09-02)
1.0
Appendix
Table 9: Model and API configurations used in the model ablation study
Item
Setting
Model
DeepSeek-V4-Flash
Model ID
deepseek-v4-flash
Endpoint
Volcano Engine Ark
API interface
OpenAI Responses API
Maximum output tokens
2048
Timeout
60 seconds
Appendix
Table 10: Extractor configuration
Field
Meaning and requirement
involves_high_value_resource
Boolean value indicating whether a high-value resource is involved. This field is required.
resource_category
One of the six resource categories or none .
resource_description
Short description of the resource.
requested_purpose
Short description of the requested purpose.
action_class
One of the seven action types.
confidence
Value from 0 to 1. This value is recorded but does not affect the decision.
Appendix
Table 11: Extractor output schema
Resource category
Anonymous actions
Attack-side authorized users
Additional benign-side actions
condition
read_status
None
All action types
energy
read_status
None
All action types
interaction
read_status
None
All action types
knowledge
read_status
None
All action types
material
read_status
None
All action types
symbolic
read_status
None
All action types
Appendix
Table 12: Resource policy used in the evaluation
Item
Description
Number of normal tasks
900
Number of resource categories
6
Tasks per resource category
150
Pairing rule
Each normal task is paired with one attack task involving the same resource
Authorized identity
ops-owner-01
Authorized actions
All action types for all six resource categories
Appendix
Table 13: Authorized normal-task evaluation
Evaluation item
Definition
Attack evaluation set
750 attack cases that satisfy the baseline conditions and have attack_success == True
Authorized normal-task set
900 authorized tasks
Interception rate
Number of BLOCK decisions divided by 750
False rejection rate
Number of BLOCK decisions divided by 900
Post-defense ASR
Number of non-intercepted cases divided by the total number of evaluable cases in the corresponding category
ASR denominator
The same denominator used in the main results table
Appendix
Table 14: Evaluation sets and metric definitions
Decision reason
Statistical category
no_high_value_resource
Allowed
allowlisted
Allowed
allowlisted_authorized_user
Authorized allow. This appears only on the benign side.
category_not_allowlisted
Policy rejection
action_not_allowlisted
Policy rejection
identity_not_authorized
Policy rejection
Appendix
Table 15: Decision reason categories
Category
DeepSeek-V4-Pro
GPT-5.5
Gemini-3.5-Flash
Qwen3.8-Max
ASR
95% CI
ASR
95% CI
ASR
95% CI
ASR
95% CI
(a) Direct Acquisition
Condition
0.7
[0.0, 3.7]
0.0
[0.0, 2.4]
1.3
[0.2, 4.7]
6.3
[2.8, 11.6]
Energy
6.1
[2.8, 11.3]
5.0
[1.8, 10.5]
6.1
[2.8, 11.2]
0.0
[0.0, 2.9]
Interaction
5.4
[2.4, 10.4]
2.0
[0.4, 5.7]
7.4
[3.7, 12.8]
2.0
[0.4, 5.7]
Knowledge
20.7
[14.5, 28.0]
19.1
[13.1, 26.3]
19.9
[13.7, 27.3]
12.1
[9.6, 17.8]
Appendix
Table 16: Attack success rates (ASR, %) for direct acquisition and resource hijacking by resource category across four model backends. Resource hijacking achieves higher ASR in every category for all four models. Each result is reported with its 95% confidence interval.
Category
OpenClaw
Prompt Defense
AgentDoG 1.5
LlamaFirewall
ResGate
TSR
Interval
TSR
Interval
TSR
Interval
TSR
Interval
TSR
Interval
Condition
98.7
[95.3, 99.6]
84.7
[78.0, 89.6]
94.0
[89.0, 96.8]
92.0
[86.5, 95.4]
90.0
[84.2, 93.9]
Energy
93.3
[88.2, 96.3]
68.0
[60.0, 75.6]
92.0
[86.5, 95.4]
68.7
[60.9, 75.6]
94.7
[89.8, 97.3]
Interaction
95.3
[90.7, 97.7]
75.3
[67.9, 81.5]
95.3
[90.7, 97.7]
80.7
[73.6, 86.2]
78.0
[70.7, 83.9]
Knowledge
88.0
[80.6, 92.1]
61.3
[53.4, 68.8]
80.0
[72.4, 86.4]
42.0
[34.4, 50.0]
87.3
[81.1, 91.7]
Material
85.3
[78.8, 90.1]
60.7
[52.7, 68.1]
89.3
[84.7, 94.7]
74.0
[66.4, 80.4]
97.3
[93.3, 99.0]
Appendix
Table 17: Task success rates (TSR, %) on authorized normal tasks by resource category. Each configuration reports the mean TSR and the corresponding reported interval. Higher is better.
Item
Configuration
Benchmark
ResourceHijackBench, parent-contract version
Dataset size
300 parent scenarios and 900 prompts
Resource categories
Material, condition, energy, symbolic, knowledge, and interaction
Attack settings
Direct, implicit, and persistent-context
Model family
DeepSeek-V4-Pro
Environment
Isolated workspace copy and local simulated services
Appendix
Table 18: Configuration shared by the two harnesses.
LLM agents increasingly rely on third-party skills, using natural-language descriptions for selection and instruction bodies for planning. This progressive-disclosure design exposes two sequential control points to untrusted publishers: a static skill may steer an otherwise correct task onto an unnecessarily costly trajectory. Prior work studies selection manipulation, malicious skill instructions, and tool-chain resource amplification largely separately, leaving their end-to-end composition unclear. We introduce Convergent Detour Hijacking (CDH), a text-only, runtime-independent attack that couples these stages. Under shared semantic cover, a description establishes relevance during selection, while an aligned body reuses that rationale to fabricate plausible dependencies during planning. CDH attracts an attacker-controlled coordinator alongside legitimate skills, recruits unnecessary benign skills into a bounded detour, and then re-enters the original route to preserve task completion. We evaluate it across multiple LLM backends and 491 held-out tasks under single-task and multi-turn conditions. On DeepSeek-V4-Pro, the matched coordinator is selected in 80.02% of tasks; among coordinator-hit runs that complete tasks, token consumption and end-to-end execution time increase by 66.91% and 92.45%, respectively, while aggregate task completion remains comparable. Thus, correct outcomes do not guarantee trajectory integrity or cost safety.
Junliang Liu, Ruoyu Li, Wenxin Tang +4
School of Computer Science and Software Engineering, Shenzhen University · The Chinese University of Hong Kong · The Chinese University of Hong Kong, Shenzhen
Tool-using large language model (LLM) agents face two distinct security failures: unauthorized external actions and exposure of sensitive plaintext inside the runtime before any final output check can intervene. Existing defenses usually protect one boundary, either the planner/runtime or the action sink, and therefore do not by themselves secure both surfaces. We present SecureClaw, a dual-boundary architecture that places authorization at the effect sink and plaintext confinement at the read boundary. Sensitive reads pass through a trusted gateway that replaces raw values with opaque handles and, in the evaluated deployment, bounded summaries as an explicit declassification interface. Writes that change external state follow a PREVIEW→COMMIT protocol in which only a trusted executor may commit the exact canonical request authorized by policy. The runtime can still plan over summaries and symbolic references, but cannot directly dereference secrets or perform side effects. Across AgentDojo, AgentLeak, and Agent Security Bench (ASB), SecureClaw is the only defense we evaluate in a common harness that simultaneously retains usable task utility and achieves 0% attack success rate (ASR) on ASB, 0.64% ASR on AgentDojo, and 3.23% overall leak on AgentLeak's attacked parity lane, which measures final-output and internal-relay leakage.
Large language model (LLM) agents increasingly achieve long-horizon tasks by combining foundation models with explicit skills and implicit procedural knowledge acquired through execution. The resulting task-solving capabilities have become valuable proprietary assets, raising a new security question: can a substantially weaker attacker-controlled agent acquire the capabilities of a stronger proprietary agent through limited black-box interaction? Existing skill-stealing attacks recover explicit skill artifacts, yet we show that artifact leakage does not necessarily transfer capability: a weaker agent may possess the same skills but still fail because it lacks procedural behaviors implicitly realized by the stronger agent. Our key insight is that the skill execution gap itself forms a leakage surface, where missing behaviors are exposed through observable differences between successful victim executions and failed attacker executions. Based on this, we present AgentLeak, a black-box capability-cloning attack that identifies capability-critical behaviors from these execution differences and incorporates them into attacker-side skills, while keeping the attacker's model, harness, and tools unchanged. Across 20 task scenarios comprising 600 instances, diverse agent systems, and multiple backbone models, AgentLeak improves task pass rates by over 40% compared with direct skill reuse and recovers more than 80% of the victim--attacker capability gap. Our findings reveal a confidentiality risk in LLM agents: protecting explicit artifacts alone is insufficient, as observable execution behavior can leak the procedural knowledge required to reconstruct proprietary task-solving capabilities in low-capability and attacker-controlled agents.
Xiaoting Lyu, Yuhong Wu, Yufei Han +6
Xi’an Jiaotong University · INRIA · University of Warwick +1