When agents share a reward for completed tasks, reporting unsafe work can reduce the reporter's reward by stopping a task. Audits can make reporting optimal without ensuring that further training teaches a silent team to report. We study this learning problem in a game where any witness can stop a task by reporting. With k witnesses per task sharing a policy and drawing independently, the expected-reward derivative with respect to their shared silence probability counts each task's benefit k times at universal silence. The comparison with universal reporting counts it once. For arbitrary policy groups, we give an audit condition sufficient for exact policy-gradient updates to reach universal reporting and, apart from boundary cases, necessary near universal silence. In a balanced family, the cheapest audits meeting the condition with prescribed positive margins cost exactly k times as much for full sharing as for one policy per role. We train PPO policies on 24 witness graphs from learned silence. Separating co-witnesses reduces unsafe completion by 33.59 percentage points compared with shuffled groups of the same sizes under the same audits (95% graph-bootstrap interval: 21.03-45.13). Only 9 of 48 witness-group runs achieve below 1% unsafe completion while retaining at least 90% legitimate completion. At the same audit budget, a fully shared network meets both thresholds in none of 48 runs with independent action draws and all 48 with a common draw.
Figures & tables
Figure 1. Exact expected reward J(p) of a fully shared policy with B=1 , D=1.5 and k=4 . Universal reporting ( p=0 ) earns 0 and universal silence ( p=1 ) earns −0.5 . With independent draws the curve falls to its minimum at p∗≈0.72 and then rises to p=1 , so silence is a local maximum that exact gradients reach from every p above p∗ . With a common draw, Jcommon(p)=(B−D)p falls linearly to −0.5 . A plot of expected team reward against the silence probability p. The independent-draw curve starts at 0 when p is 0, falls to about minus 0.81 at p equal to 0.72, which a dotted vertical line marks as p star, and rises to minus 0.5 at p equal to 1, so both endpoints are local maxima. The common-draw line falls linearly from 0 to minus 0.5.
Arm
Networks and audits
Witness groups
One network per coloring group
Shuffled groups
Witness-group sizes, agents permuted
Shuffled, redesigned allocation †
Shuffled groups, own allocation
Witness + uniform audits
Witness groups, uniform audits
Witness + coalition audits
Witness groups, coalition audits
Fully shared
One network
Table 1. Arms of the experiments. Unless stated, arms use the reference allocation and independent draws. Daggers mark controls designed after the main study.
Reference budget
Eight times the reference budget
Arm
Condition
Met
Unsafe completion
Joint recovery
Met
Joint recovery
Witness groups
Certificate
24
51.20
9
24
44
Shuffled groups
Certificate
1
84.79
0
24
45
Shuffled, redesigned allocation †
Certificate
14
77.23
0
–
–
Witness + uniform audits
Certificate
0
99.95
0
19
41
Witness + coalition audits
Certificate
1
88.20
1
24
48
Table 2. Outcomes on the 24 general graphs. The Condition column gives a sufficient condition for exact updates to reach universal reporting, evaluated at the benefit upper bound, and Met gives the number of graphs that satisfy it. Joint recovery is out of 48 runs. Legitimate completion is at least 99.65% in every row. A dash marks an arm that does not form a grouping or a budget that was not run.
(a) Joint recoveries
At the bound
In each unsafe context
Budget
Certified
Uncertified
Certified
Uncertified
1
10 of 52
0 of 188
10 of 78
0 of 162
2
32 of 106
3 of 134
35 of 126
0 of 114
4
113 of 160
16 of 80
129 of 186
0 of 54
8
200 of 208
13 of 32
206 of 224
7 of 16
Table 3. Post hoc splits by certification on the 24 general graphs. (a) Joint recoveries in the witness, shuffled, fully shared, uniform-audit and coalition-audit arms, by whether the allocation certifies the arm’s grouping at the benefit upper bound and in each of the run’s four unsafe contexts. Half the reference budget certifies nothing and gives no joint recovery. (b) Witness advantage in points by whether the allocation certifies shuffled groups at the bound, with graph counts in parentheses.
Figure 2. Responses to the audit budget on the 24 general graphs, with 48 runs per point. Coalition audits keep witness groups and change only the allocation. Two panels with the audit budget on a logarithmic axis from 0.5 to 8 times the reference budget. The left panel shows mean unsafe completion. At the reference budget, witness groups are near 51 percent while shuffled groups, the fully shared network, the shared network with agent IDs and coalition audits stay above 84 percent. At eight times the reference budget, all methods except the fully shared network and the shared network with agent IDs are below 2 percent. The right panel shows joint recoveries out of 48. Witness groups rise from 0 to 9, 18, 29 and 44, coalition audits from 0 to 1, 11, 37 and 48, shuffled groups from 0 to 0, 6, 23 and 45, the shared network with agent IDs from 0 to 0, 10, 27 and 41, and the fully shared network from 0 to 0, 0, 19 and 35.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Experiment
Records
Common pretraining
72
Seven methods at five audit budgets
2,520
Independent-network architecture
72
Five architectures at two head resets
720
Two groupings at two optimism levels
192
Three methods under two witness-error levels
432
Appendix
Table S1. The fixed main-study schedule. The 4,080 records include pretraining. Continuations extend selected base runs and are not independent replicates. All records completed.
Family
k
Unsafe W
Unsafe S
S−W
Joint recovery W
Joint recovery S
Uniform
4
43.32
78.43
35.11
0/6
0/6
Uniform
8
86.39
88.02
1.62
0/6
0/6
Ring
4
31.36
63.09
31.73
4/6
0/6
Ring
8
17.19
99.99
82.80
4/6
0/6
Modular
4
31.48
81.77
50.29
0/6
0/6
Modular
8
84.50
98.37
13.86
0/6
0/6
Appendix
Table S2. Reference-budget results by family and witness width. Each row contains three graph seeds and two training replicates. W denotes witness groups and S shuffled groups. Completion columns are percentages, and differences are percentage points. Balanced and rewired rows are secondary and are excluded from the primary estimate.
Figure S1. The 24 paired general-graph effects at the reference budget. Each marker is the mean of a graph’s two training replicates, and the ticks at the ends of the vertical line through it are the effects in the replicates themselves. Within a replicate, both arms start from the same pretrained network. Filled markers have witness width four and open markers width eight. Horizontal black segments are family means, and positive values favor witness groups. The graph means are the units of the primary bootstrap. A dot plot of shuffled-group minus witness-group unsafe completion in percentage points for 24 graphs in four families. Each graph has a marker at the mean of its two training replicates and a vertical line whose end ticks mark the two replicate effects. Family means are about 18 for uniform, 57 for ring, 32 for modular and 27 for hub graphs. Most graph means are positive. Four lie at or below zero, the lowest near minus 60 in the ring family and another near minus 32 in the uniform family. Replicate effects range from about minus 66 to 100, and the two replicates of a graph differ by more than 50 points on eight graphs.
Budget
Rounds for C
Graphs
Joint recovery R
Joint recovery C
Unsafe R
Unsafe C
2
Fewer
10
3/20
2/20
61.36
48.78
2
Equal
6
6/12
6/12
17.45
11.08
2
More
3
3/6
2/6
7.65
30.21
2
Uncertified
5
6/10
1/10
33.36
87.22
4
Fewer
6
0/12
5/12
74.03
12.24
4
Equal
18
29/36
32/36
9.09
1.48
Appendix
Table S3. Coalition audits against the reference allocation for witness groups on the 24 general graphs. Graphs are split by whether coalition audits need fewer, equal or more elimination rounds than the reference allocation, or do not certify witness groups, which the reference allocation certifies on every graph at these budgets. R denotes the reference allocation and C coalition audits. Each row covers both training replicates of its graphs, and the unsafe columns are mean percentages. The split is post hoc and coincides with the graph subsets named in the text.
Method
Budget
Certified
Unsafe
Legitimate
Joint recovery
Witness groups
0.5
0/24
99.99
100.00
0/48
Witness groups
1
24/24
51.20
99.99
9/48
Witness groups
2
24/24
37.84
99.98
18/48
Witness groups
4
24/24
25.33
99.98
29/48
Witness groups
8
24/24
1.86
99.97
44/48
Shuffled groups
0.5
0/24
98.51
100.00
0/48
Appendix
Table S4. Complete budget aggregates on the general families. Each row contains 48 runs. Budget is a multiple of the reference budget, subject to the full-audit cap. Uniform and coalition rows keep witness groups. Certified counts the graphs on which the audits certify the architecture’s groups, and it does not apply to scalar offsets or agent IDs, which do not form a grouping. Completion columns are percentages, and all action draws are independent.
Budget
Grouping and allocation
Unsafe
Legitimate
Joint recovery
1
Witness, reference
51.20
99.99
9/48
1
Shuffled, reference
84.79
99.99
0/48
1
Shuffled, redesigned
77.23
99.99
0/48
2
Witness, reference
37.84
99.98
18/48
2
Shuffled, reference
56.75
99.99
6/48
2
Shuffled, redesigned
42.62
99.98
13/48
Appendix
Table S5. Allocation control with same-host repetitions. Every row contains the same 24 general graphs and both training replicates. Completion columns are percentages. The horizon, endpoint and starting policy are unchanged.
Architecture
p0
Unsafe
Legitimate
Joint recovery
Shared + agent IDs
0.75
26.96
55.46
2/72
Shared + agent IDs
0.95
77.30
99.92
1/72
Shared + scalar offsets
0.75
50.02
55.55
3/72
Shared + scalar offsets
0.95
93.05
100.00
5/72
Fully shared
0.75
55.55
55.55
0/72
Fully shared
0.95
100.00
100.00
0/72
Appendix
Table S6. Output-head resets at the reference budget across all six families, with 72 runs per row. p0 is the reset silence probability, and the resets keep the hidden layer but change the starting policy. Completion columns are percentages.
Architecture
p0
k
Start
Unsafe
Legitimate
Legitimate below 90%
Joint recovery
Witness groups
0.75
4
31.64
2.27
86.10
5/36
25/36
Witness groups
0.75
8
10.01
2.37
11.11
32/36
1/36
Witness groups
0.95
4
81.45
4.65
99.99
0/36
22/36
Witness groups
0.95
8
66.34
43.54
99.99
0/36
10/36
Shuffled groups
0.75
4
31.64
60.58
91.66
3/36
4/36
Shuffled groups
0.75
8
10.01
9.14
13.89
31/36
0/36
Appendix
Table S7. Output-head resets by witness width, with 36 runs per row across all six families. The start column is the unsafe and legitimate completion 100p0k at the reset. Completion columns are percentages.
Method
Error
Unsafe
Legitimate
Joint recovery
Shared + scalar offsets
10%
91.66
100.00
6/72
Shared + scalar offsets
25%
95.94
100.00
1/72
Witness + uniform audits
10%
99.99
99.99
0/72
Witness + uniform audits
25%
99.99
99.99
0/72
Witness groups
10%
68.58
99.99
7/72
Witness groups
25%
70.01
99.99
5/72
Appendix
Table S8. Witness error at the reference budget across all six families, with 72 runs per row. Error is the nominal fraction of tasks whose witness set is replaced, rounded to an integer task count. The true graph supplies the budget and the environment. Completion columns are percentages.
Grouping
ω
Unsafe
Legitimate
Joint recovery
Shuffled groups
0
96.63
99.73
0/48
Shuffled groups
0.5
98.96
100.00
0/48
Witness groups
0
82.63
99.62
1/48
Witness groups
0.5
87.65
99.99
2/48
Appendix
Table S9. Optimistic advantage transformations at the reference budget. Each row has 48 runs over six families, two widths, the first two graph seeds and two training replicates. The transformation is max{ωA,A} . Completion columns are percentages. This tests one component and does not reproduce the full method of Zhao et al. (2024) .
Architecture
Rollouts
Unsafe
Legitimate
Joint recovery
Shared + scalar offsets
4,096
95.83
100.00
1/24
Shared + scalar offsets
16,384
91.66
100.00
2/24
Shuffled groups
4,096
92.34
99.99
0/24
Shuffled groups
16,384
78.67
100.00
0/24
Witness groups
4,096
42.44
99.99
8/24
Witness groups
16,384
26.97
100.00
14/24
Appendix
Table S10. Fixed continuation subset at the reference budget. Each row has 24 runs over six families, two widths, the first graph seed and two training replicates. The 16,384-rollout run continues the 4,096-rollout run with optimizer and random state preserved, and each endpoint uses the final quarter of its own horizon. Completion columns are percentages.
Figure S2. Learning in the common-draw control. Lines are medians across the 48 general-graph runs and shading gives the interquartile range. Unsafe completion uses a logarithmic scale. The legitimate panel shows that the common draw does not reach low unsafe completion by reporting everywhere. The tabulated endpoint is a final-quarter mean. Two panels against recovery rollouts from 0 to 4096. In the left panel, on a logarithmic scale, unsafe completion under independent draws stays near 100 percent, while under the common draw it falls within a few hundred rollouts to below 0.01 percent and ends near 0.0007 percent. In the right panel, legitimate completion stays near 100 percent for both arms throughout training.
Fraud can pose a challenge in many resource allocation domains, including social service delivery and credit provision. For example, agents may misreport private information in order to gain benefits or access to credit. To mitigate this, a principal can design strategic audits to verify claims and penalize misreporting. In this paper, we introduce a general model of audit policy design as a principal-agent game with multiple agents, where the principal commits to an audit policy, and agents collectively choose an equilibrium that minimizes the principal's utility. We examine both adaptive and non-adaptive settings, depending on whether the principal's policy can be responsive to the distribution of agent reports. Our work provides efficient algorithms for computing optimal audit policies in both settings and extends these results to a setting with limited audit budgets.
Reinforcement-learning (RL) policies are often distributed as opaque neural checkpoints, while training logs show that a run occurred without explaining what the policy learned. We study whether independently trained policies can be represented and composed through auditable discrete behavioral rules. We define auditability as six separately testable predicates: trace integrity, lossless coding, rule coverage, behavioral agreement, composition quality, and value-model reliability. Our protocol uses a shared frozen symbolizer, passive rule extraction, an append-only hash-bound ledger, exact environment replay, and offline confidence-ranked arbitration with an explicit blind-spot fallback. The results place strict limits on this description layer. Rule-set overlap does not imply behavioral agreement: policies may share symbolic rules while choosing near-chance-matching actions on fresh states. The fused policy therefore selects among existing rules rather than generating a new skill. On a conflict-dominated task, an apparent fusion failure is traced to an induction/deployment mismatch: rules induced from sampled actions were evaluated under argmax actions, and deployment-consistent re-induction reverses the arbitration ordering. A fitted-Q generalized-policy-improvement diagnostic also fails in both environments, limiting claims that rule fusion is superior to value-based composition. One exploratory comparison favors rule fusion, but its comparator is post hoc, the task is partly saturated, and the fused policy remains below the strongest held-out actor. We contribute an evidence-bounded audit and composition protocol, not a claim of universal interpretability or autonomous skill generation. Future work must add temporally extended skills, cross-skill interfaces, composition search, and independent novelty audits.
Long-running AI agents create a control problem: each action they take changes the state, which in turn affects the trajectory of future actions. If the agent is not fully aligned, then guaranteeing safety requires approving consequential actions before allowing them to be executed. But requiring human approval at every step makes attention a bottleneck. Delegating review to other AI agents raises the same alignment problem: the reviewers may themselves be misaligned. We identify a condition on a reviewing panel that is weaker than individual alignment yet necessary and sufficient for a guarantee that the principal fares at least as well in expectation as under a designated baseline policy. Each reviewer agent reports whether an action proposal made by a proposer agent improves its own utility relative to the baseline. We show that a threshold rule tolerating k disapprovals is safe exactly when, after any k reviewers are removed, the principal's utility can be written as a nonnegative combination of the remaining reviewers' utilities, plus a term that is nonnegative on every feasible proposal. We call this property k-robust coalitional alignment. The characterization lifts to sequential control: in a discounted MDP with an arbitrary proposer agent, safety at every state is both necessary and sufficient for the induced policy to match or improve on the baseline. When reviewers vote strategically, full-panel coverage in reward-function space guarantees that every Nash equilibrium is safe under the unanimous approval rule; in contrast, more permissive thresholds can admit unsafe equilibria even when reviewers are individually aligned. Experiments with existing reviewer models show that collective review can remain sound without an aligned individual, even when some disapprovals are tolerated.