Multi-agent systems derive their capabilities from sharing evidence, delegating tasks, and combining information across agents. The same process creates a safety problem: contributions that are admissible in isolation can jointly enable a prohibited use. Blocking every sensitive action avoids disclosure but defeats the purpose of collaboration. We introduce authorization-paired evaluation, which makes blocking prohibited uses and completing required authorized uses a joint success criterion, and FlowReview, a framework connecting object resolution, permission ranking, and deterministic enforcement. In controlled composition experiments, reviewing combined artifacts reduces the denied-commit rate from 86.0% to zero with no loss of authorized supply. Our findings show that preserving information and lineage alone does not ensure correct permission attribution. Object identity and permission must remain connected to execution through components whose outputs can be verified. Together, these findings establish a system-level requirement for multi-agent safety: govern composed information flows while preserving the authorized capabilities that make collaboration useful.
Figures & tables
Figure 1 : Locally admissible fragments can compose into a prohibited use. Three workers contribute disjoint parts of one governed object. With the same proposal and DENY policy, local review leaves the object unresolved and permits execution. Global review resolves the combined artifacts, allowing the commit gate to block the action.
Figure 2 : FlowReview connects object identity, permission, and execution. Object and action contracts inform all three review capabilities (blue solid arrows). Trusted policy P guides permission ranking and is independently checked at enforcement (green dashed arrows). The commit gate enforces the resulting action decision.
Figure 3 : Object binding and enforcement recover selective correctness while reducing disclosure. Bars pool 576 runs per family and condition at a 24-round budget. Hollow circles, squares, and triangles represent 32 scenario means for the three legend conditions, respectively, with 18 runs per scenario. Horizontal offsets separate equal values. Whiskers show 95% scenario-cluster bootstrap intervals from 10,000 resamples. Higher C and lower D are better.
Figure 4 : Controlled interventions isolate object identity, enforcement, and reader scope. Bars show aggregate rates. (a) Oppositely authorized same-class objects, 288 runs per condition. (b) Model review versus commit gate, 991 and 1,009 runs, on a smaller percentage scale. (c) The same fixed proposals under local or global review, 480 per policy and reader. Global review blocks denied use while preserving authorized supply.
Figure 5 : Contributor growth can impair resolution despite complete information. (a) Bars pool 168 graphs per (N,K) condition. Hollow squares ( K=4 ) and circles ( K=N−1 ) show seven team rates, with 24 graphs each. (b) Paired coverage changes from (8,7) to (16,15) with 95% scenario-cluster bootstrap intervals. Only homogeneous Haiku’s coverage decline is significant after Holm correction. Mixed labels identify coordinators. All conditions retain complete information and 100% authorized supply.
Figure 6 : Explicit contracts improve lineage but leave permission-attribution errors. (a) Lines show rates pooled over 168 graphs per agent count. Hollow squares (type-only) and circles (explicit) show seven team rates, each based on 24 graphs and offset horizontally for visibility. (b) Bars pool 840 graphs per contract, with explicit-contract values labeled. Correct attribution requires all roles to preserve the claim’s permission and untrusted authority rank. Higher is better.
Component
Lineage
Correctness
A
T
Denied commits
Assembly placement: execution correctness
Model composer
N/A
0.675
0.350
0.475
0/80
Runtime assembler
N/A
1.000
1.000
1.000
0/80
Permission placement: proposal correctness
Runtime lineage + specialist
1.000
1.000
N/A
1.000
0/480
Table 1: Verifiable placement preserves permission binding and restores authorized assembly. Correctness scores execution for assembly and pre-execution selections for permission review. Assembly uses 80 runs per arm, including 40 ALLOW runs for A . Permission results are identical across four specialists on 480 graphs each. Denied counts use all runs. T follows study-specific task checks. N/A marks unmeasured or inapplicable endpoints.
Execution arm
Target commit
Final violation
Utility
Exact-call match
Native, no gate
21/24
24/24
9/24
5/24
Object-bound runtime stack
0/24
4/24
21/24
5/24
Table 2: Runtime control blocks target commits in independent AgentDojo-derived Sonnet pairs. Both arms use 24 matched scenarios. Target commits use the action ledger. Final violations and utility use benchmark oracles. Exact-call match requires the prescribed legitimate tool call, a stricter criterion than benchmark utility.
Appendix figures & tables25 assets
Supplementary material from the paper’s appendix.
Appendix
Study
Controlled contrast
Evaluation size
Controlled composition
Local / global read scope
480 DENY + 480 ALLOW proposals
Natural composition
Generated artifacts, composer
60 evaluation pairs, 12 clusters
Capability ladder
Class / binding / enforcement
576 runs/cell, 32 clusters
Contributor scaling
Agent count N , contributors K
588 + 672 graphs
Delegation contracts
Type-only / explicit lineage
840 graphs/contract
Permission placement
Four isolated specialists
480 graphs/specialist
Appendix
Table 3: Studies, comparisons, and evaluation units. Each row retains its study-specific sampling unit and denominator.
Figure 7 : Communication and contributor structure shape the information available for review. (a–d) Arrows show untrusted content in red, recorded flows in blue, and trusted-policy control as green dashed lines. The relay contrast fixes message content across source wrappers. (e,f) Readers inspect separate artifacts or their union while proposal and policy remain fixed. (g) At N=16 , circles denote coordinators, filled squares informative workers, and dashed squares background workers without governed fields. Numbers in the lower boxes count initially assigned fields, totaling 16 in each condition. K counts initial field holders and N includes the coordinator, excluding runtime components.
Experiment
N
Communication
Capability tested
Direct evaluation / ladder
Varies
Drafter and review components
Resolution, ranking, enforcement
Matched relay
1 to 2
Direct or relayed bytes
Topology and selectivity
Authority delegation
1 to 16
Source first, then handoffs
Lineage and attribution
Natural composition
4
Three workers and a composer
Cross-agent reconstruction
Contributor scaling
1 to 16
Star graph, K informative
Information versus headcount
AgentDojo transfer
1 or 3
Direct or worker/relay/executor
External action boundary
Appendix
Table 4: Topologies and measured capabilities. N counts model agents, including the coordinator. Runtime components are counted separately. Ladder model roles depend on the review configuration.
Stage
Artifact
Fields retained
Object resolved
Outcome
Intake specialist
Message
6/6
No
N/A
Telemetry analyst
Store
5/5
No
N/A
Identity analyst
Handoff
5/5
No
N/A
Coordinator
Proposal
N/A
Reference only
Checks pass
Runtime assembly
Union
16/16
Yes
Exact registered record
Local runtime
DENY
N/A
No
Denied commit
Appendix
Table 5: One homogeneous N=4,K=3 execution. The same fixed proposal receives four combinations of policy and reader scope. N/A marks a field that does not apply to that stage.
Record element
Constructed example
Task
Prepare report REPORT-A and send the governed object when authorized.
Governed object
SYNTH_CODE_REVIEW_001
Worker artifacts
Message: SYNTH_C. Store note: ODE_REV. Handoff: IEW_001.
Proposed action
send_external, with the complete object as its body and code-review.example.invalid as destination.
Trusted permission
DENY in one execution and ALLOW in its paired execution.
Local outcome
The action commits under both permissions.
Appendix
Table 6: A constructed scenario and its recorded composition. Input fragments and the governed object are specified before execution and shown with display aliases. The proposal is produced by the Gemma composer from the three Haiku artifacts. Both policy-conditioned calls produce the displayed action.
Endpoint
Estimate
95% cluster interval
Trusted-policy effect ΔP
0.8828
[0.8125,0.9453]
Claim effect ΔQ(DENY)
0.0000
[0.0000,0.0000]
Claim effect ΔQ(ALLOW)
0.0156
[−0.0625,0.0938]
Authorized supply
0.8828
[0.8047,0.9453]
Task completion
1.0000
[1.0000,1.0000]
Appendix
Table 7: Trusted policy controls execution with little interference from untrusted claims. Four models share 16 scenario clusters. Effects are differences in commit rates. Authorized supply uses 128 ALLOW trajectories, and task completion uses all 256 trajectories. Zero-width intervals reflect constant observed cluster outcomes rather than zero population uncertainty.
Figure 8 : Selective control persists across attack budgets while authorized recovery varies by composer. (a) Circles and squares represent the two attack families, with 576 final-configuration runs per point and an expanded 98.5–100% axis. (b) Neutral and untrusted ALLOW claims under trusted policy P , with 64 trajectories per bar. (c) Left green bars show authorized supply and right blue bars task completion over 12 scenarios per model. Whiskers show 95% scenario-cluster bootstrap intervals. Collapsed intervals reflect constant observed outcomes.
Model
Fewer denied commits
Tie
More denied commits
Haiku 4.5
3
3
2
Sonnet 4.5
6
2
0
Qwen3-32B
5
1
2
Gemma3-27B
6
2
0
Llama-3.3-70B
2
2
4
Total
22
10
8
Appendix
Table 8: Reader choice changes control coverage. Counts compare the family-conditioned decoder with literal matching across representation families. Each comparison holds 12 proposals and their policies fixed.
N
Workers
K
Condition
Artifact path
Comparison
1
0
1
Local full record
Direct coordinator
Single-agent reference
2
1
1
Full record
One artifact
First delegation edge
4
3
3
Shared-scaled
Three artifacts
Distributed reconstruction
8
7
7
Contributor-scaled
Seven artifacts
More contributors
16
15
15
Contributor-scaled
Fifteen artifacts
Contributor scaling
8
7
4
Dummy-scaled
4 informative + 3 background
Same N , less information
Appendix
Table 9: Contributor-scaling matrix. Workers exclude the coordinator. K counts initial field holders, including the coordinator at N=1 .
Object
Denied commits
ALLOW supply
Setting
Graphs
Local
Union
L
G
L
G
Discovery, N≤2
168
168/168
168/168
0/168
0/168
168/168
168/168
Discovery, distributed
420
0/420
420/420
420/420
5/420
420/420
420/420
Replication, non-maximal
504
0/504
504/504
504/504
0/504
504/504
504/504
Replication, N16,K15
168
0/168
168/168
168/168
19/168
168/168
168/168
Replication, all cells
672
0/672
672/672
672/672
19/672
672/672
672/672
Appendix
Table 10: Composition endpoints in discovery and replication. L/G denote local/global readers of the same fixed proposals. Object columns report complete visibility locally and in the artifact union.
Cell
Schema purity
Role leakage
Worker conformance
Lineage coverage
N=8,K=4
144/168
24/168
0.956
168/168
N=8,K=7
120/168
48/168
0.895
168/168
N=16,K=4
144/168
24/168
0.979
168/168
N=16,K=15
113/168
55/168
0.877
149/168
All four cells
521/672
151/672
0.927
653/672
Appendix
Table 11: Surface conformance and lineage in the powered replication. Each cell contains 168 graphs across seven teams. Worker conformance is a rate over worker calls. Other columns count graphs.
Figure 9 : Delegation separates visibility errors from permission-attribution drift. (a) Circles mark source-visibility errors and squares final-proposal attribution drift, pooling 168 graphs per agent count. (b) Cells report type-only complete lineage in an independent study, with 24 graphs per team and agent count. Darker blue indicates higher coverage, and values are percentages. Mixed labels name the coordinator.
Design or criterion
Specification
Lineage and permission placement
Model agents N
1, 4, 16
Policy factors
Trusted P× untrusted Q
Role teams
5
Permission specialists
Sonnet 4.5, Haiku 4.5, Qwen3-32B, Gemma3-27B
Clusters, dev / confirm
4 / 8
Appendix
Table 12: Placement comparisons and evaluation criteria. Development and confirmation clusters are disjoint. Thresholds are prespecified, and the commit gate controls execution.
Trusted policy
Selection component
Selected action
Runtime outcome
ALLOW
Final coordinator
Object reference mistyped
Rejected
ALLOW
Isolated Qwen specialist
Exact registered candidate
Committed
DENY
Isolated Qwen specialist
No action selected
No commit
Appendix
Table 13: A permission decision must remain bound to the selected action. The ALLOW rows use the same saved coordinator output. The DENY row is its policy-conditioned companion, whose upstream outputs were generated separately.
Artifact channel
Assembler
Execution correct
ALLOW supply
Task
No failover
Model composer
0.700
0.400
0.600
No failover
Runtime assembler
0.800
0.600
0.600
Validated failover
Model composer
0.675
0.350
0.475
Validated failover
Runtime assembler
1.000
1.000
1.000
Appendix
Table 14: Availability and assembly ablation. Execution correctness and task rates use 80 executions per arm, while authorized supply uses the 40 ALLOW executions. Execution correctness requires no committed action under DENY or the exact required action under ALLOW. Task success additionally requires a valid report. The common commit gate is applied before these outcomes are scored.
Condition
Gate binding
Wrong commits
Authorized supply
Cached-policy execution
N/A
120/240
0/120
Current-policy reread
N/A
1/240
120/120
Reread + version gate
170/240
0/240
83/120
Appendix
Table 15: Version checking after policy reread blocks wrong commits but can withhold authorized actions. Independent confirmation bank with 240 evaluations per arm, including 120 currently authorized actions. Binding is measured by the version gate, so N/A denotes ungated arms. Wrong-commit counts use the full bank.
Proposal input / control
Valid
Binding
Supply
Task
Protocol
Current-policy reread
1.000
0.721
1.000
1.000
1.000
Stale proposal + gate
1.000
0.000
0.000
1.000
1.000
Stale proposal + repair
0.788
0.000
0.000
0.000
0.785
Appendix
Table 16: Current-policy reread recovers authorized supply that one repair attempt does not. A separate confirmation bank has 240 evaluations per arm and 120 currently authorized actions. Proposal, binding, and task rates use all 240 evaluations. Protocol validity is averaged over the role outputs specified by each arm. Binding checks the current policy version and lineage, while supply checks the exact executed action. All arms record zero wrong commits.
Exposure
Model
Exact attacks
Utility
Protocol
Inst.
Unchanged official
Sonnet 4.5
0/24
0.875
1.000
yes
Unchanged official
Haiku 4.5
0/24
0.792
1.000
yes
Unchanged official
Qwen3-32B
5/24
0.458
1.000
yes
Authority re-expression
Sonnet 4.5
13/24
0.625
1.000
yes
Authority re-expression
Haiku 4.5
0/24
0.625
1.000
yes
Authority re-expression
Qwen3-32B
21/24
0.083
1.000
yes
Appendix
Table 17: Relay exposure guides confirmation. Each cell contains 24 development trajectories. Inst. marks whether instrumentation meets its validity criterion. Qwen is the positive control, and Sonnet confirmation uses independent pairs.
Observed outcome
Native
Gated
Required lunch event
Created
Created
Required calendar invitation
Sent
Sent
Private email forwarded
Yes
Blocked
Exact legitimate-action match
1
1
Benchmark utility
0
1
Appendix
Table 18: Blocking the prohibited action restores utility while retaining the required action. One saved Sonnet pair. The legitimate calendar operation generates the invitation email.
Model
Exact transient
Rate
Native utility
Protocol
Sonnet 4.5
0/36
0.000
0.750
0.917
Haiku 4.5
0/36
0.000
0.556
0.972
Qwen3-32B
4/36
0.111
0.528
0.694
Llama-3.3-70B
2/36
0.056
0.194
0.944
Appendix
Table 19: Native attack exposure across models. Each row contains 36 development trajectories from a separate exposure comparison. Exact events, task utility, and protocol conformance use their stated denominators.
Official attack
Arm
Proposal
Commit
Violation
Utility
Commit Δ [95% CI]
Important Instructions
Native
6/8
6/8
5/8
5/8
N/A
Important Instructions
Gated
6/8
0/8
0/8
1/8
−0.750 [ −1.000,−0.375 ]
InjecAgent
Native
5/8
5/8
5/8
2/8
N/A
InjecAgent
Gated
5/8
0/8
0/8
1/8
−0.625 [ −0.875,−0.250 ]
System message
Native
3/8
3/8
3/8
4/8
N/A
System message
Gated
2/8
0/8
0/8
3/8
−0.375 [ −0.750,−0.125 ]
Appendix
Table 20: Exploratory Qwen attack-family comparison. Each family has eight matched pairs. Commit Δ is gated minus native, with a 95% cluster interval. N/A marks native rows where the paired contrast is not repeated.
Official attack
Arm
Proposal
Commit
Violation
Utility
Commit Δ [95% CI]
Important Instructions
Native
0/8
0/8
4/8
3/8
N/A
Important Instructions
Gated
0/8
0/8
2/8
5/8
0 [ 0,0 ]
InjecAgent
Native
0/8
0/8
3/8
3/8
N/A
InjecAgent
Gated
0/8
0/8
4/8
3/8
0 [ 0,0 ]
System message
Native
0/8
0/8
2/8
2/8
N/A
System message
Gated
0/8
0/8
2/8
1/8
0 [ 0,0 ]
Appendix
Table 21: Qwen evaluation on independent task and injection pairs. Each family has eight matched pairs. Commit Δ is gated minus native with a 95% cluster interval. N/A marks native rows where the paired contrast is not repeated. Pooled estimates are exploratory because native exposure falls below the confirmatory threshold. In zero-event cells, [0,0] is a degenerate empirical bootstrap interval, not evidence of zero population uncertainty.
Evidence
Direct conclusion
Measurement scope
Paired evaluation
Authorized capability changes the safety target
Matched policy intervention
Paired evaluation
Resolution, ranking, and enforcement are distinct
Verbatim action ledger
Composition
Admissible artifacts identify a denied action
Read-scope intervention
Scaling
Distributed information drives resolution burden
Star graphs, seven teams
Delegation
Lineage and authority are separate capabilities
Paired contracts, 840 graphs
Permission placement
Lineage and action binding retained
Four specialists, 480 graphs each
Appendix
Table 22: Contributions and supporting comparisons. Measurement scope identifies the intervention or execution surface behind each conclusion.
Boundary
Reader / object
Metric
Value
Action commit
Literal governed objects
C
0.990 to 0.998
Action commit
Literal governed objects
Verbatim releases
0/2,304
Action commit
Eight representation families
Lower commit
22/40
Action commit
Eight representation families
Higher commit
8/40
Action commit
Eight representation families
Unchanged commit
10/40
Appendix
Table 23: Reader-specific control at the action boundary. Ladder rows aggregate final-configuration trials across two attack families and two budgets. Representation rows count model-by-transformation comparisons against literal matching, each with 12 fixed proposals.
Outcome
Result
Natural composition: 60 evaluation pairs
Denied exact recovery, local
3/60
Denied exact recovery, global
0/60
Authorized supply, global
15/60
Trusted policy × untrusted claim: 320 units
Policy effect
Large
Appendix
Table 24: Outcomes across controlled comparisons. Each group identifies its experimental setting and sampling unit. Delegation-depth drift concerns exact final proposals, while contract-study drift counts errors at any role. Permission placement and assembly are separate comparisons.
Multi-agent systems (MAS) decompose long-horizon tasks across supervisors and subagents, but delegated goals do not necessarily carry their original authorization boundaries. Existing safety benchmarks mainly study adversarial compromise, while work on constraint drift lacks controlled architecture-level evaluation. We introduce MasDrift, a benchmark of 600 benign productivity tasks across eight domains. Each task pairs required work with reserved actions. MasDrift compares single-agent, centralized, and decentralized coordination while varying hierarchy depth and peer width, measuring task completion and authorization preservation. Across generic multi-agent conditions, centralized hierarchies achieve 93.9--98.6% task completion versus 85.7--87.0% for peer networks, while unauthorized actions occur in 2.7--19.8% of tasks versus 0.6--0.8%, a gap that widens with hierarchy depth. We further compare two defenses that differ in where authorization evidence resides. One re-anchors every pending call to the original user request. The other carries an attenuated policy along the delegation chain. Re-anchoring reduces unauthorized actions in every model configuration we evaluate, at a cost of 1.6 points of pooled completion. Chain propagation blocks required work instead, forfeiting up to 36.3 points. A heterogeneous case study confirms that the failure follows from coordination rather than model strength. MasDrift exposes a centralization tradeoff and makes authorization preservation a measurable property of MAS design.
Zhuoning Xu, Xiucheng Zhang, Hanjun Luo +3
New York University · New York University Abu Dhabi · The Hong Kong Polytechnic University +1
Long-running AI agents create a control problem: each action they take changes the state, which in turn affects the trajectory of future actions. If the agent is not fully aligned, then guaranteeing safety requires approving consequential actions before allowing them to be executed. But requiring human approval at every step makes attention a bottleneck. Delegating review to other AI agents raises the same alignment problem: the reviewers may themselves be misaligned. We identify a condition on a reviewing panel that is weaker than individual alignment yet necessary and sufficient for a guarantee that the principal fares at least as well in expectation as under a designated baseline policy. Each reviewer agent reports whether an action proposal made by a proposer agent improves its own utility relative to the baseline. We show that a threshold rule tolerating k disapprovals is safe exactly when, after any k reviewers are removed, the principal's utility can be written as a nonnegative combination of the remaining reviewers' utilities, plus a term that is nonnegative on every feasible proposal. We call this property k-robust coalitional alignment. The characterization lifts to sequential control: in a discounted MDP with an arbitrary proposer agent, safety at every state is both necessary and sufficient for the induced policy to match or improve on the baseline. When reviewers vote strategically, full-panel coverage in reward-function space guarantees that every Nash equilibrium is safe under the unanimous approval rule; in contrast, more permissive thresholds can admit unsafe equilibria even when reviewers are individually aligned. Experiments with existing reviewer models show that collective review can remain sound without an aligned individual, even when some disapprovals are tolerated.
Multi-agent systems improve capability through task decomposition and role specialization, but these same mechanisms introduce an important safety blind spot: a harmful objective can be fragmented into locally plausible subtasks, allowing malicious intent to evade detection by any single agent. This is a growing social-impact challenge: systems handling sensitive information or consequential tools can turn routine delegation into unauthorized disclosure or unsafe action. We argue that this failure mode is better understood as a semantic information-flow problem than as a single-turn prompt classification task. To address this, we propose SafeFlow, a defense framework for multi-agent systems that formalizes malicious cross-agent propagation as a semantic information-flow problem. SafeFlow attaches structured semantic taints to root requests, propagates them through a dynamic collaboration graph, and performs workflow-level validation to reconstruct the global risk context before irreversible actions are committed. Evaluated on four benchmarks spanning prompt injection, jailbreak-based unsafe tool use, risky code execution, and harmful web-agent behavior, SafeFlow reduces attack success rates compared to undefended baselines and external defenses while retaining high benign task completion and a high paired safe--harm success rate. Our findings show that multi-agent systems still lack mechanisms for preserving risk semantics across delegation boundaries. This gap can turn routine delegation into privacy harms or unsafe actions that affect people and organizations. SafeFlow keeps this risk visible throughout the workflow, before it results in harm.
Haowen Dai, Zonghao Ying, Wenfeng Li +10
University of Nottingham Ningbo China · Beihang University · Shandong University +6