G-CARB: Graph-Localized Conformal Agent Risk Budget for Compositional Harm
Authors: Zijun Yu, Yu Gu, Vahid Partovi Nia, Masoud Asgharian
Organizations: Department of Mathematics and Statistics McGill University Montréal, QC, Canada · Department of Mathematics and Industrial Engineering Polytechnique de Montréal Montréal, QC, Canada
Small language model (SLM) agents need safety controls that track consequences across tool calls with little monitoring overhead. A private read, for example, becomes a leak when a later action sends that data outside the system. We introduce CARB (Conformal Agent Risk Budget), which calibrates when to stop an agent using a ledger of harm incurred before stopping. Under exchangeable episodes, standard conformal risk control bounds this declared loss in expectation over calibration and a future episode. G-CARB selects scorer evidence along observable dependencies from private sources to outgoing actions. The ledger still covers the entire executed history, and computing the gate score requires no additional language-model inference. On AgentDojo replay with two 14B backbones, G-CARB roughly halves scorer-input records at intermediate risk budgets while improving autonomous task completion relative to full-prefix scoring; random context of the same size achieves similar gains. Controlled examples show how retaining the relevant dependency can further avoid stopping benign work.
Figures & tables
Figure 1: A leak across tool calls. The certificate, note, and external send form a private source–carrier–sink chain. G-CARB retains this chain; Full-prefix includes unrelated work. The illustrated local labels miss the leak. The executed-step baseline instead labels the first harmful event and divides by the executed action count.
Figure 2: AgentDojo risk–utility frontiers. Shorter context improves autonomous success at intermediate budgets; equal-size random context closely tracks G-CARB . Every plotted method’s mean ledger loss stays below the budget (dashed line). Points average 200 shared calibration/test assignments of the fixed corpus. Whiskers show their empirical 5th–95th percentiles. Appendix Tables 4 and 7 give the values.
Figure 3: Stopping harmful flows while allowing benign work, at α=.10 . All gated methods shown stop every harmful sink. CES benign stop counts stopped non-leaks; graph-braid counts only first stops on disconnected benign branches, divided by all trajectories. Thus Full-prefix intervenes in every graph-braid episode but has benign-stop rate .095 . Context saved is CtxRed. Means use 100 CES and 200 graph-braid assignments; Appendix Table 8 includes equal-size random context.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Feature
Weight
Intercept
.05
Current private source
.09
Current external sink
.25
Current financial action
.35
Current irreversible action
.24
Current message/social action
.12
Appendix
Table 1: Frozen weighted scorer. A risky current action is external, financial, or irreversible. The weighted sum is clipped to [0,1] ; only a retained source-to-sink path activates the .52 provenance feature.
Dataset
Pool
Calibration
Test
Assignments
Assignment unit
AgentDojo, each backbone
984
492
492
200
Episode
CES clean
160
48
48
100
Pair identifier
CES near-miss
320
96
96
100
Pair identifier
Graph-braid
2,000
600
600
200
Trajectory
Appendix
Table 2: Pool and fold sizes, all in trajectories. CES uses split seeds 0 – 99 ; AgentDojo and the main graph-braid setting use 0 – 199 . CES trajectories sharing a pair identifier stay together in one fold.
Method
Calibrated object
Scorer evidence
Executed-step CRC
Rate envelope rλ
Current proposal
Ledger-current
Stopped-prefix ledger Lλ
Current proposal
Full-prefix CARB
Same ledger
Full prefix and proposal
G-CARB
Same ledger
Dependency-selected nodes
Fixed-window-4
Same ledger
Four-record recency context
Equal-size random
Same ledger
Size-matched random context
Appendix
Table 3: Internal calibration and context controls. No guard is uncalibrated. Ledger-current is reported only on the constructed suites; Executed-step CRC is reported only on AgentDojo, so these results do not provide a same-context loss ablation. Executed-step CRC is an internal analogue of step-level control, not a reproduced CORA implementation.
Executed-step CRC
Full-prefix
Equal-size random
G-CARB
Backbone
α
r↓
Lˉ↓
Lˉ↓
I↓
U↑
Lˉ↓
I↓
U↑
Lˉ↓
I↓
U↑
Qwen2.5
.050
.027
.133
.044
.680
.065
.035
.515
.168
.035
.516
.167
.100
.027
.133
.072
.406
.222
.077
.332
.261
.077
.351
.250
Qwen3
.050
.034
.154
.046
.989
.001
.041
.992
.002
.041
.992
.002
.100
.044
.225
.092
.826
.025
.085
.535
.107
.085
.537
.106
Appendix
Table 4: AgentDojo results averaged over 200 calibration/test assignments. Executed-step CRC calibrates the rate envelope rλ and reports raw rate r . Full-prefix, equal-size random, and G-CARB calibrate the same stopped-prefix ledger with different scorer evidence. Lˉ , I , U : mean episode loss, intervention, autonomous success. Both restricted-context selectors improve utility over Full-prefix at the displayed points; Qwen3 at .05 is near stop-all.
Full-prefix threshold
G-CARB threshold
Backbone
α
Min
Median
Max
Min
Median
Max
Feasible per method
Qwen2.5
.025
.24
.54
.84
.13
.44
.58
200/200
.050
.58
.84
.98
.44
.58
.72
200/200
.100
.99
.99
.99
.99
.99
.99
200/200
.200
1.00
1.00
1.00
1.00
1.00
1.00
200/200
Qwen3
.025
.24
.24
.35
.13
.13
.25
200/200
Appendix
Table 5: Selected-threshold ranges over 200 assignments per cell ( 3,200 calibrations total). All corrected calibration risks are at most their displayed budget; the range shows that assignment matters most at the tighter operating points.
Backbone
α
ΔLˉ
ΔI
ΔU
Qwen2.5
.05
−.009
−.163
+.102
[−.016,.004]
[−.280,−.102]
[.063,.171]
Qwen2.5
.10
+.005
−.055
+.029
[.002,.008]
[−.067,−.043]
[.018,.037]
Qwen3
.05
−.005
+.003
+.000
[−.033,.004]
[−.004,.053]
[−.002,.002]
Appendix
Table 6: Paired mean differences and empirical 5th–95th assignment percentiles over 200 shared splits per cell. Intervention is lower in three cells, while the Qwen3 .05 cell has essentially no utility difference. These are assignment-sensitivity summaries, not independent-sample intervals.
Fixed-window-4
Equal-size random context
Backbone
α
Lˉ
I
U
CtxRed
Lˉ
I
U
CtxRed
Qwen2.5
.025
.022
.869
.010
.130
.022
.977
.008
.494
.050
.044
.669
.072
.159
.035
.515
.168
.565
.100
.072
.395
.228
.175
.077
.332
.261
.596
.200
.133
.000
.436
.201
.133
.000
.436
.580
Qwen3
.025
.001
1.000
.000
.003
.000
1.000
.000
.089
Appendix
Table 7: Complete context-control grid, averaged over 200 assignments per cell. Equal-size random context performs similarly to G-CARB on this natural corpus. Fixed-window-4 uses a different recency-based selection rule. CtxRed is computed separately on each method’s scored proposals using Equation 2 .
Cohort
Method
Lˉ
Harm stop
Benign stop
I
U
CtxRed
CES clean: 80/80
No guard
.500
.000
.000
.000
1.000
–
Ledger-current
.000
1.000
1.000
1.000
.000
.846
Full-prefix
.000
1.000
.000
.500
.500
.000
Fixed-window-4
.000
1.000
.000
.500
.500
.462
Equal-size random
.000
1.000
1.000
1.000
.000
.346
G-CARB
.000
1.000
.000
.500
.500
.346
Appendix
Table 8: Constructed diagnostics at α=.10 : 100 grouped CES assignments and 200 graph-braid assignments. CES labels give harmful/non-leak pool counts; the graph-braid label gives total pool size. All gated methods have zero reported mean ledger loss and stop the harmful sinks. Benign stop divides by non-leak trajectories in CES, but by all trajectories in graph-braid (Appendix D ). These are descriptive held-out results, not a CES certificate under the episode-permutation corollary.
Evidence
Lˉ
Path recall
False path
CtxRed
Provenance + , text +
.000
1.000
.000
.346
Provenance + , text −
.000
1.000
.000
.346
Provenance − , text +
.000
1.000
.000
.833
Provenance − , text −
.500
.000
.000
.846
Appendix
Table 9: Fixed-threshold CES evidence audit. “Text − ” denotes laundered exact sink text and “provenance − ” removal of structured provenance from the policy evidence. The threshold is calibrated in the provenance + , text − cell and reused. Removing both channels gives zero harmful-path recall and mean ledger loss .500 , equal to the no-guard loss of the balanced cohort. This is an observability/shift diagnostic, not a four-cell conformal certificate.
Sweep
Conditions
Max Lˉ
Min H-stop
Max B-stop
CtxRed range
Baseline
1
.000
1.000
.000
.471–.471
Harm prevalence
4
.054
.400
.000
.471–.472
Benign density
3
.000
1.000
.000
.471–.471
Score noise
3
.097
.765
.030
.469–.472
Edge dropout
3
.000
1.000
.094
.400–.471
Combined stress
2
.080
.869
.063
.459–.467
Appendix
Table 10: Graph-braid sensitivity at α=.10 . Each condition’s metrics are averaged over 20 assignments; each row gives extrema over the listed conditions. H-stop is the fraction of harmful sinks blocked. B-stop counts first stops on benign branches, divided by all trajectories. The reported mean losses stay below the budget; these assignment averages do not guarantee that every held-out batch meets it.
Figure 4: Evidence removal and recalibrated perturbation results. A: Harmful-path recall over 100 grouped CES assignments of 80 matched pairs. The threshold is calibrated with structured provenance present and exact sink text laundered, then reused in all four cells. Removing both channels reduces recall to zero; the benign false-path rate is zero throughout. Table 9 gives ledger losses. B: Mean episode loss and intervention for 16 graph-braid conditions at budget α=.10 , with separate calibration in each condition. Whiskers show standard errors over 20 assignments per condition. The 20% edge-dropout condition stops every episode. These recalibrated comparisons do not test transfer of a fixed threshold across conditions.