stat.MLOct 4, 2026

G-CARB: Graph-Localized Conformal Agent Risk Budget for Compositional Harm

Authors: Zijun Yu, Yu Gu, Vahid Partovi Nia, Masoud Asgharian

Organizations: Department of Mathematics and Statistics McGill University Montréal, QC, Canada · Department of Mathematics and Industrial Engineering Polytechnique de Montréal Montréal, QC, Canada

Abstract

Small language model (SLM) agents need safety controls that track consequences across tool calls with little monitoring overhead. A private read, for example, becomes a leak when a later action sends that data outside the system. We introduce CARB (Conformal Agent Risk Budget), which calibrates when to stop an agent using a ledger of harm incurred before stopping. Under exchangeable episodes, standard conformal risk control bounds this declared loss in expectation over calibration and a future episode. G-CARB selects scorer evidence along observable dependencies from private sources to outgoing actions. The ledger still covers the entire executed history, and computing the gate score requires no additional language-model inference. On AgentDojo replay with two 14B backbones, G-CARB roughly halves scorer-input records at intermediate risk budgets while improving autonomous task completion relative to full-prefix scoring; random context of the same size achieves similar gains. Controlled examples show how retaining the relevant dependency can further avoid stopping benign work.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. AgentBoundary: Counterfactual Evaluation of Safety in Tool-Using LLM Agents

    Sep 27, 2026Tianzhuo Yang, Zirui Mi, Yantao Huang +4RefusalsEvaluation Benchmarks

  2. AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents

    Apr 3, 2026Yunhao Feng, Yifan Ding, Yingshui Tan +6Computer-Use AgentsHazard