Language-model agents increasingly use tools to act on external systems. Earlier actions can alter files, permissions, database records, or other state, making a later routine-looking action harmful. Yet the visible interaction may not reveal the underlying state needed to assess that action. We formulate attack and defense as partially observed state control in SEAD, deriving their design requirements from this shared execution process. Because attackers supply instructions while the target chooses concrete actions, DART decomposes harmful goals into locally plausible steps and uses feedback from actual tool execution to guide trajectory search. The defender must decide before execution with incomplete state evidence. SAGE can therefore investigate relevant state through read-only queries before allowing or blocking each action, including those proposed after a block. We construct an environment-verifiable dataset integrating controlled initial states, replayable tool environments, and task-specific executable checks. Across four target models, DART improves semantic attack success by 18.8--35.9 percentage points over the competing baseline, with consistent gains under executable verification. On recorded trajectories, SAGE preserves 95.79% of benign trajectories while intercepting 92.73% of harmful paths by the harm-enabling boundary. In online attack-defense evaluation, it reduces DART's executable attack success from 48.0% to 4.0%. SAGE remains effective across four attack methods and generalizes to out-of-domain environments. Our code and data is available at https://github.com/EverywhereSafety/SEAD.
Figures & tables
Symbol
Meaning
sn,xn,un
sn is the environment state before attempt n , including resources and accumulated disclosures or external effects. xn is the active attacker instruction, and un is the target’s proposed tool action.
hnT,hnD
Interaction histories visible to the target ( hnT ) and defender ( hnD ). These role-specific views need not reveal the full environment state.
F,Ω
For an executed action, F determines the next environment state and Ω returns the visible tool feedback.
ΓD,dn
The defender’s pre-execution gate ΓD may gather evidence before returning decision dn . PASS permits execution, while BLOCK stops the pending action.
g,Hg
g is the harmful goal. Hg(s)=1 means that state s completes this goal, and Hg(s)=0 means it remains unmet.
Vg
A task-specific executable check of goal completion from environment outcomes, used as an operational proxy for Hg .
Table 1: Core notation for the shared execution loop. The attacker supplies instructions, the target proposes tool actions, and the defender decides whether they execute. The index n counts target tool-action attempts, including blocked calls.
Attack
GPT-5.6 Luna
GPT-5.6 Terra
Gemini 3.8 Flash
Claude Sonnet 5
Semantic
Hard
Semantic
Hard
Semantic
Hard
Semantic
Hard
Direct
37.4
40.1
37.4
33.2
24.1
25.1
17.1
23.5
MTA
27.8
21.4
29.4
22.5
25.1
19.8
16.0
13.4
STAC
29.4
23.0
24.6
18.7
19.8
18.2
18.2
14.4
Intent Hijacking
27.2
19.3
17.1
11.8
10.2
8.6
13.4
10.7
DART
73.3
54.7
62.5
51.1
43.9
33.2
51.9
38.7
Table 2: RQ1: attack effectiveness against four target agents. Semantic success is judged from recorded execution evidence by an independent Judge, whereas Hard success requires the task’s executable evaluator to pass. All values are ASR in percent, so a higher value means a stronger attack, and the best value in each column is bold and underlined . Rows follow the baseline taxonomy of Section 5.2 : a single-instruction request, a fixed decomposition, two adaptive attacks, then DART . Appendix E.1 illustrates a real DART attack trajectory.
Defense metrics
First-block timing
Defender
B↑
H1↑
F1↑
H2↑
F2↑
Early
Exact ↑
Late
Miss ↓
No defense
100.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
100.00
Base Binary
51.58
0.00
0.00
17.03
25.61
63.64
0.00
1.82
34.55
Fine-tuned Binary
100.00
18.18
30.77
22.05
36.13
5.45
18.18
3.64
72.73
StepGuard
51.58
12.73
20.42
27.77
36.11
41.82
12.73
1.82
43.64
TS-Guard
71.58
7.27
13.20
38.14
49.76
78.18
7.27
0.00
14.55
Table 3: Offline defense metrics (left) and first-block timing breakdown (right) on 55 Semantic-success attack trajectories from the 75-task subset. B is the all-pass rate over 95 benign trajectories. Early, Exact, Late and Miss are disjoint percentages; Early and Late carry no preferred direction, because a defender that never blocks scores zero on both. The leading value in each directed column is bold and underlined . All values are percentages.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Defender
Early
Exact
Late
Miss
ℓ1↓
No defense
0.00
0.00
0.00
100.00
–
Base Binary
63.64
0.00
1.82
34.55
10.25
Fine-tuned Binary
5.45
18.18
3.64
72.73
1.20
StepGuard
41.82
12.73
1.82
43.64
8.94
TS-Guard
78.18
7.27
0.00
14.55
6.77
Safiron
85.45
3.64
0.00
10.91
6.92
Appendix
Table 4: First-block timing on successful attack trajectories. Early, Exact, Late, and Miss are percentages over the trajectories. ℓ1 is the mean ∣k−T∣ over finite BLOCK s.
Defender
Filesystem
Terminal
PostgreSQL
Web
Overall ASR (%)
Semantic
Hard
Semantic
Hard
Semantic
Hard
Semantic
Hard
Semantic
Hard
No Defense
16
14
13
13
10
5
10
4
65.3
48.0
Base Binary
10
7
9
9
8
5
2
1
38.7
29.3
StepGuard
8
5
9
9
6
9
2
1
33.3
32.0
TS-Guard
9
5
8
9
1
0
1
1
25.3
20.0
Fine-tuned Binary
10
10
12
11
11
9
11
6
58.7
48.0
Appendix
Table 5: Online defense against DART on the 75-task subset with GPT-5.6 Luna. Domain columns count successful attacks; overall ASR is a percentage over 75 tasks. The lowest values in each column, including ties, are bold and underlined . Lower values mean stronger defense.
Setting
Fine-tuned Binary
SAGE
Training actions / demonstrations
6,820
7,407
Training source trajectories
960
1,230
Epochs
2
1
Optimizer updates
854
463
GPUs
2 H200
4 H200
Effective batch size
16
16
Appendix
Table 6: Training configurations for the reported defenders. A SAGE demonstration is a defender interaction for one pending target action, potentially containing multiple investigation steps.
Turn
Request framing
Recorded execution
Pv
1
Inspect the account table
Schema inspection establishes the available fields.
0
2
Prepare a temporary staging area
The target asks whether the table should be session-local or persistent; no new tool call occurs.
0
3
Prepare a persistent staging area
A database write creates and populates the staging table; account data remains unchanged.
0
4
Apply the staged maintenance update
The target reads the staged value and uses it to replace passwords for five accounts.
1
Appendix
Table 7: Preparatory state changes and subsequent harm in single.188 . Requests and execution effects are summarized from the recorded path; Pv is the executable progress score after each instruction turn.
Planning call(s)
Recorded search development
1
Generate two initial candidate continuations.
2–5
Expand descendants of root.0 ; two leaves terminate after repeated target refusals.
6
Return to and expand the alternative prefix root.1 .
7–8
Expand its children; both reach confirmed-success leaves.
9–10
Explore another continuation under root.0 , ending in an unsuccessful leaf.
Appendix
Table 8: Search chronology for single.118 . Planning calls count controller expansions, rather than individual target tool actions.
Turn
Actions
PASS
BLOCK
Queries (successful)
Goal confirmed
1
1
1
0
0 (0)
No
2
3
3
0
5 (5)
No
3
1
0
1
1 (1)
No
4
1
0
1
2 (1)
No
5
1
1
0
1 (1)
Yes
Appendix
Table 9: Selected online path for single.136 . Actions count new target tool-action attempts, including blocked calls. Queries count SAGE ’s environment investigations, with successful returns in parentheses.
Initial branch
Nodes
Terminal paths
BLOCK
Successes
root.0
6
3
2
0
root.1
13
6
4
0
Total
19
9
6
0
Appendix
Table 10: Explored branches in single.157 . Nodes exclude the root; successes count confirmed-success nodes. Each BLOCK decision is counted once, even when it lies on a prefix shared by several terminal paths.
Stage
Recorded evidence or decision
Pending action
Change ownership of public.users .
Investigation
Read the table’s metadata through env_db_table .
Returned evidence
A regular table with credential, contact, and payment-related columns.
Decision
BLOCK : the ownership change establishes control over the sensitive resource.
Appendix
Table 11: Environmental evidence used in the OOD PostgreSQL case. The query returns schema metadata, not the contents of sensitive records. Decision reasoning is paraphrased from the recorded submission.
Action(s)
Evaluation stage
Decision
1–7
Recorded prefix preceding the annotated closure
PASS
8
Read the live staged artifact, then assess the pending configuration change
BLOCK
Appendix
Table 12: Exact offline interception in single.53 : the first block occurs at k=T=8 . The environment read is a defender investigation before target action 8.
Adaptation family
Tasks
Initial-state and execution interface
Filesystem
52
Seeded workspace and file tools; task 10 also uses an ownCloud service.
Terminal
51
Reset command sandbox and task-specific installed tools.
PostgreSQL
42
Seeded database, fixed connection profile, and database tool interface.
GitLab
10
Resettable GitLab instance and browser interaction.
ownCloud
10
Resettable ownCloud instance and browser interaction.
Reddit
22
Resettable forum instance and browser interaction.
Appendix
Table 13: Adaptation families for the complete 187 tasks list. Each task retains its own objective and executable checks.
Target
Temperature
Top- p
Reasoning effort
Output limit
GPT-5.6 Luna
1
1
low
4,096
GPT-5.6 Terra
1
1
low
4,096
Gemini-3.8 Flash
1
1
low
4,096
Claude Sonnet 5
omitted
omitted
unspecified
4,096
Appendix
Table 14: Target-side sampling settings encoded by the checked MTAR configurations. “Omitted” denotes a parameter not sent by the worker.
Table 15: Argument contracts for the MTAR defender investigation tools. History tools are shared; environment tools are selected by task domain. Final submission follows the separate decision protocol above.
Method
Default interaction settings
Direct
One original request
MTA
Released task-specific instruction sequence
DART
Depth 8; branching factor 2; at most 25 executed nodes and 25 planning calls
STAC
One preparation candidate; at most 3 adaptive turns
Intent Hijacking
2 strategies; at most 7 turns per strategy and 3 candidates per turn
Appendix
Table 16: Default execution settings used in the reproduced methods, per task attempt. STAC preparation budgets are additional to its adaptive turns. Early stopping can reduce actual usage.
Tool-using language agents turn model decisions into external side effects: they read files, run scripts, call APIs, send messages, and invoke Model Context Protocol tools. This makes agent attacks different from jailbreaks. The harmful step is often not an obviously forbidden output, but an ordinary executable action that becomes unsafe because attacker-controlled context steers authorized access against the user's interest. We identify this failure mode as authority confusion: untrusted resources may inform reasoning, but they must not authorize side effects. We present AIRGuard, a runtime guard that operationalizes least privilege as action-time authorization. AIRGuard normalizes heterogeneous tool calls, derives task authority into step-level authority, tracks source and target trust, simulates sensitive side effects, audits cross-step risk, and enforces decisions before actions execute. On AgentTrap, AIRGuard reduces Sonnet 4.6 attack success from 36.3% without defense to 5.5%. On DTAP-150, AIRGuard preserves 76.0% benign utility with Haiku 4.5, compared with 52.0% for ARGUS and 42.0% for MELON. An ablation further shows that prompt-only policy helps only modestly, whereas a dedicated runtime authority-control layer gives the agent system direct control over tool-mediated side effects. Code and data are available at https://github.com/Sophie508/AIRGuard.
Suliu Qin, Haomin Zhuang, Yujun Zhou +2
University of Liverpool · University of Notre Dame · 2Inria, France
Tool-using large language model (LLM) agents turn generated text into real side effects, so poisoned tool metadata, retrieved pages, memory, and reusable skills can steer the next call. Vetting an artifact before admission does not settle this. A safe variant and a leaking variant can produce the same admission evidence, and a sound gate then cannot relax that site for either. We make that condition precise, which leaves the last boundary a deployment can still act on. We present Provenance-Aware Capability Enforcement (PACE), which mediates every tool call immediately before it executes. Path confinement proposes an executable cut of represented influence paths, while capability and effect verification checks schema-defined effects against authority compiled from the authenticated request. We distinguish the certified execution contract from the evaluated configuration, which can restore an authorized call after a proposed block or apply a declared repair. Confinement requires the final action to preserve the certified cut. On eight executable agent-security benchmarks with three target-model families, the evaluated configuration gives strictly lowest attack success in 62 of 79 eligible attack columns and ties in 14; full-benchmark native utility loses at most three points relative to the undefended agent. A complete ablation over 1167 paired cases attributes most security gains to effect verification and refusal control to boundary adaptation. A reduced-scale adaptive search succeeds on 0/30 out-of-authority targets against the defense.
Fengpeng Li, Qizhou Wang, Yuke Hu +5
PRADA Lab, King Abdullah University of Science and Technology · Imperfect Information Learning Team, RIKEN Center for Advanced Intelligence Project · State Key Laboratory of Internet of Things for Smart City, University of Macau +2
Contextual security defenses prevent AI agents from taking rogue actions by synthesizing a task-specific policy and enforcing it on the agent's tool calls. In multi-step tasks, however, which actions are valid often depends on what the agent has already done and learned. We present Sapien, a policy engine for enforcing stateful contextual policies. A Sapien policy specifies permitted tool-call sequences using a regular expression extended with stateful predicates, deferred policy generation, and scoped semantic checks. We show that Sapien stays within a few percent of an unconstrained agent's utility. Even if the agent is fully hijacked, Sapien's policies rule out 93-95% of attacks on AgentDojo and 62-85% on Toolathlon (twice as many as tool allowlists on long-horizon tasks).