G-CARB: Graph-Localized Conformal Agent Risk Budget for Compositional Harm
Authors: Zijun Yu, Yu Gu, Vahid Partovi Nia, Masoud Asgharian
Organizations: Department of Mathematics and Statistics McGill University Montréal, QC, Canada · Department of Mathematics and Industrial Engineering Polytechnique de Montréal Montréal, QC, Canada
Small language model (SLM) agents need safety controls that track consequences across tool calls with little monitoring overhead. A private read, for example, becomes a leak when a later action sends that data outside the system. We introduce CARB (Conformal Agent Risk Budget), which calibrates when to stop an agent using a ledger of harm incurred before stopping. Under exchangeable episodes, standard conformal risk control bounds this declared loss in expectation over calibration and a future episode. G-CARB selects scorer evidence along observable dependencies from private sources to outgoing actions. The ledger still covers the entire executed history, and computing the gate score requires no additional language-model inference. On AgentDojo replay with two 14B backbones, G-CARB roughly halves scorer-input records at intermediate risk budgets while improving autonomous task completion relative to full-prefix scoring; random context of the same size achieves similar gains. Controlled examples show how retaining the relevant dependency can further avoid stopping benign work.
Figures & tables
Figure 1: A leak across tool calls. The certificate, note, and external send form a private source–carrier–sink chain. G-CARB retains this chain; Full-prefix includes unrelated work. The illustrated local labels miss the leak. The executed-step baseline instead labels the first harmful event and divides by the executed action count.
Figure 2: AgentDojo risk–utility frontiers. Shorter context improves autonomous success at intermediate budgets; equal-size random context closely tracks G-CARB . Every plotted method’s mean ledger loss stays below the budget (dashed line). Points average 200 shared calibration/test assignments of the fixed corpus. Whiskers show their empirical 5th–95th percentiles. Appendix Tables 4 and 7 give the values.
Figure 3: Stopping harmful flows while allowing benign work, at α=.10 . All gated methods shown stop every harmful sink. CES benign stop counts stopped non-leaks; graph-braid counts only first stops on disconnected benign branches, divided by all trajectories. Thus Full-prefix intervenes in every graph-braid episode but has benign-stop rate .095 . Context saved is CtxRed. Means use 100 CES and 200 graph-braid assignments; Appendix Table 8 includes equal-size random context.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Feature
Weight
Intercept
.05
Current private source
.09
Current external sink
.25
Current financial action
.35
Current irreversible action
.24
Current message/social action
.12
Appendix
Table 1: Frozen weighted scorer. A risky current action is external, financial, or irreversible. The weighted sum is clipped to [0,1] ; only a retained source-to-sink path activates the .52 provenance feature.
Dataset
Pool
Calibration
Test
Assignments
Assignment unit
AgentDojo, each backbone
984
492
492
200
Episode
CES clean
160
48
48
100
Pair identifier
CES near-miss
320
96
96
100
Pair identifier
Graph-braid
2,000
600
600
200
Trajectory
Appendix
Table 2: Pool and fold sizes, all in trajectories. CES uses split seeds 0 – 99 ; AgentDojo and the main graph-braid setting use 0 – 199 . CES trajectories sharing a pair identifier stay together in one fold.
Method
Calibrated object
Scorer evidence
Executed-step CRC
Rate envelope rλ
Current proposal
Ledger-current
Stopped-prefix ledger Lλ
Current proposal
Full-prefix CARB
Same ledger
Full prefix and proposal
G-CARB
Same ledger
Dependency-selected nodes
Fixed-window-4
Same ledger
Four-record recency context
Equal-size random
Same ledger
Size-matched random context
Appendix
Table 3: Internal calibration and context controls. No guard is uncalibrated. Ledger-current is reported only on the constructed suites; Executed-step CRC is reported only on AgentDojo, so these results do not provide a same-context loss ablation. Executed-step CRC is an internal analogue of step-level control, not a reproduced CORA implementation.
Executed-step CRC
Full-prefix
Equal-size random
G-CARB
Backbone
α
r↓
Lˉ↓
Lˉ↓
I↓
U↑
Lˉ↓
I↓
U↑
Lˉ↓
I↓
U↑
Qwen2.5
.050
.027
.133
.044
.680
.065
.035
.515
.168
.035
.516
.167
.100
.027
.133
.072
.406
.222
.077
.332
.261
.077
.351
.250
Qwen3
.050
.034
.154
.046
.989
.001
.041
.992
.002
.041
.992
.002
.100
.044
.225
.092
.826
.025
.085
.535
.107
.085
.537
.106
Appendix
Table 4: AgentDojo results averaged over 200 calibration/test assignments. Executed-step CRC calibrates the rate envelope rλ and reports raw rate r . Full-prefix, equal-size random, and G-CARB calibrate the same stopped-prefix ledger with different scorer evidence. Lˉ , I , U : mean episode loss, intervention, autonomous success. Both restricted-context selectors improve utility over Full-prefix at the displayed points; Qwen3 at .05 is near stop-all.
Full-prefix threshold
G-CARB threshold
Backbone
α
Min
Median
Max
Min
Median
Max
Feasible per method
Qwen2.5
.025
.24
.54
.84
.13
.44
.58
200/200
.050
.58
.84
.98
.44
.58
.72
200/200
.100
.99
.99
.99
.99
.99
.99
200/200
.200
1.00
1.00
1.00
1.00
1.00
1.00
200/200
Qwen3
.025
.24
.24
.35
.13
.13
.25
200/200
Appendix
Table 5: Selected-threshold ranges over 200 assignments per cell ( 3,200 calibrations total). All corrected calibration risks are at most their displayed budget; the range shows that assignment matters most at the tighter operating points.
Backbone
α
ΔLˉ
ΔI
ΔU
Qwen2.5
.05
−.009
−.163
+.102
[−.016,.004]
[−.280,−.102]
[.063,.171]
Qwen2.5
.10
+.005
−.055
+.029
[.002,.008]
[−.067,−.043]
[.018,.037]
Qwen3
.05
−.005
+.003
+.000
[−.033,.004]
[−.004,.053]
[−.002,.002]
Appendix
Table 6: Paired mean differences and empirical 5th–95th assignment percentiles over 200 shared splits per cell. Intervention is lower in three cells, while the Qwen3 .05 cell has essentially no utility difference. These are assignment-sensitivity summaries, not independent-sample intervals.
Fixed-window-4
Equal-size random context
Backbone
α
Lˉ
I
U
CtxRed
Lˉ
I
U
CtxRed
Qwen2.5
.025
.022
.869
.010
.130
.022
.977
.008
.494
.050
.044
.669
.072
.159
.035
.515
.168
.565
.100
.072
.395
.228
.175
.077
.332
.261
.596
.200
.133
.000
.436
.201
.133
.000
.436
.580
Qwen3
.025
.001
1.000
.000
.003
.000
1.000
.000
.089
Appendix
Table 7: Complete context-control grid, averaged over 200 assignments per cell. Equal-size random context performs similarly to G-CARB on this natural corpus. Fixed-window-4 uses a different recency-based selection rule. CtxRed is computed separately on each method’s scored proposals using Equation 2 .
Cohort
Method
Lˉ
Harm stop
Benign stop
I
U
CtxRed
CES clean: 80/80
No guard
.500
.000
.000
.000
1.000
–
Ledger-current
.000
1.000
1.000
1.000
.000
.846
Full-prefix
.000
1.000
.000
.500
.500
.000
Fixed-window-4
.000
1.000
.000
.500
.500
.462
Equal-size random
.000
1.000
1.000
1.000
.000
.346
G-CARB
.000
1.000
.000
.500
.500
.346
Appendix
Table 8: Constructed diagnostics at α=.10 : 100 grouped CES assignments and 200 graph-braid assignments. CES labels give harmful/non-leak pool counts; the graph-braid label gives total pool size. All gated methods have zero reported mean ledger loss and stop the harmful sinks. Benign stop divides by non-leak trajectories in CES, but by all trajectories in graph-braid (Appendix D ). These are descriptive held-out results, not a CES certificate under the episode-permutation corollary.
Evidence
Lˉ
Path recall
False path
CtxRed
Provenance + , text +
.000
1.000
.000
.346
Provenance + , text −
.000
1.000
.000
.346
Provenance − , text +
.000
1.000
.000
.833
Provenance − , text −
.500
.000
.000
.846
Appendix
Table 9: Fixed-threshold CES evidence audit. “Text − ” denotes laundered exact sink text and “provenance − ” removal of structured provenance from the policy evidence. The threshold is calibrated in the provenance + , text − cell and reused. Removing both channels gives zero harmful-path recall and mean ledger loss .500 , equal to the no-guard loss of the balanced cohort. This is an observability/shift diagnostic, not a four-cell conformal certificate.
Sweep
Conditions
Max Lˉ
Min H-stop
Max B-stop
CtxRed range
Baseline
1
.000
1.000
.000
.471–.471
Harm prevalence
4
.054
.400
.000
.471–.472
Benign density
3
.000
1.000
.000
.471–.471
Score noise
3
.097
.765
.030
.469–.472
Edge dropout
3
.000
1.000
.094
.400–.471
Combined stress
2
.080
.869
.063
.459–.467
Appendix
Table 10: Graph-braid sensitivity at α=.10 . Each condition’s metrics are averaged over 20 assignments; each row gives extrema over the listed conditions. H-stop is the fraction of harmful sinks blocked. B-stop counts first stops on benign branches, divided by all trajectories. The reported mean losses stay below the budget; these assignment averages do not guarantee that every held-out batch meets it.
Figure 4: Evidence removal and recalibrated perturbation results. A: Harmful-path recall over 100 grouped CES assignments of 80 matched pairs. The threshold is calibrated with structured provenance present and exact sink text laundered, then reused in all four cells. Removing both channels reduces recall to zero; the benign false-path rate is zero throughout. Table 9 gives ledger losses. B: Mean episode loss and intervention for 16 graph-braid conditions at budget α=.10 , with separate calibration in each condition. Whiskers show standard errors over 20 assignments per condition. The 20% edge-dropout condition stops every episode. These recalibrated comparisons do not test transfer of a fixed threshold across conditions.
Language-model agents act through structured tool calls whose arguments carry very different risks: untrusted content may legitimately shape an email body but should never set a recipient, account, command, or credential. Existing conformal risk control methods certify a tool call as a whole, so a failure in one rare high-risk field can be averaged away by the many benign arguments around it, leaving the argument that causes harm uncertified. We introduce role-stratified per-field conformal risk control, a calibration layer that wraps any per-field detector and assigns a separate threshold and risk budget to each semantic argument role. We show that aggregate certification pays a price of coarseness, tightening a rare role's effective budget in proportion to how often that role appears, whereas role-stratified calibration certifies each sufficiently sampled role directly with a finite-sample guarantee and pools the rarest roles. Across AgentDojo and InjecAgent with six language models, our method achieves the most consistent role-specific budget compliance among the methods we evaluate under model and attack transfer, detector noise, gradual drift, unseen tool suites, and adaptive attacks, providing formal per-role guarantees under exchangeability or after recalibration. These results suggest that structured tool calls should be certified at the semantic-role level, not the whole action.
Safety alignment for large language models (LLMs) in conversational settings is largely framed around whether to answer or refuse a request. In agentic settings, however, the same models must decide whether to act as permission-critical evidence emerges during execution. This creates a distinct challenge: apparent risk, action permissibility, and task competence are easily confounded, making agentic over-refusal difficult to distinguish from ordinary task failure. To address this, we introduce AgentBound, the first four-way counterfactual generation-and-evaluation framework for tool-using agent safety. AgentBound transforms the same executable workflow by independently varying apparent risk and action permissibility, enabling controlled comparisons of risky-looking but authorized tasks and routine-looking but unauthorized tasks. These comparisons jointly diagnose over-refusal and unsafe compliance while controlling for task competence. We instantiate AgentBound as a human-validated 4,000-task evaluation suite with trajectory-based and post-state-based judgments. Across 17 model and harness configurations, high safety frequently coexists with poor authorized-task completion: GPT-5.5 blocks 99.5% of routine-looking unauthorized actions yet completes only 28.7% of risky-looking authorized tasks. We further train a lightweight runtime calibration module that improves authorized-task completion by 18.2% on average across 10 evaluated configurations, while improving unsafe-action blocking by 5.4% on average. These show that effective agentic alignment requires action decisions to track permission-relevant execution evidence, rather than refusal strength alone.
Tianzhuo Yang, Zirui Mi, Yantao Huang +4
Peking University · Beijing Academy of Artificial Intelligence
Computer-use agents extend language models from text generation to persistent action over tools, files, and execution environments. Unlike chat systems, they maintain state across interactions and translate intermediate outputs into concrete actions. This creates a distinct safety challenge in that harmful behavior may emerge through sequences of individually plausible steps, including intermediate actions that appear locally acceptable but collectively lead to unauthorized actions. We present \textbf{AgentHazard}, a benchmark for evaluating harmful behavior in computer-use agents. AgentHazard contains \textbf{2,653} instances spanning diverse risk categories and attack strategies. Each instance pairs a harmful objective with a sequence of operational steps that are locally legitimate but jointly induce unsafe behavior. The benchmark evaluates whether agents can recognize and interrupt harm arising from accumulated context, repeated tool use, intermediate actions, and dependencies across steps. We evaluate AgentHazard on Claude Code, OpenClaw, and IFlow using mostly open or openly deployable models from the Qwen3, Kimi, GLM, and DeepSeek families. Our experimental results indicate that current systems remain highly vulnerable. In particular, when powered by Qwen3-Coder, Claude Code exhibits an attack success rate of \textbf{73.63%}, suggesting that model alignment alone does not reliably guarantee the safety of autonomous agents.
Yunhao Feng, Yifan Ding, Yingshui Tan +6
Alibaba Group · Fudan University · Hunan Institute of Advanced Technology +2