When using an LLM agent in a consequential domain, making an informed decision about whether to trust its output or intervene requires calibrated confidence in the agent's success. Confidence estimation for agents is difficult because evidence about success is distributed across heterogeneous, interdependent steps of an agent's trajectory. Practical agentic deployments introduce further challenges: frontier LLMs often provide limited access to internal signals, agent roll-outs are costly, and training data may be unavailable or quickly become outdated. To address these challenges, we introduce Confidence Reasoning Graphs (CRGs), an inference-time framework that estimates the probability an agent accomplished its task from a single trajectory, without privileged model access or training data. Rather than compressing an execution into a single holistic judgment, a CRG begins with the claim that the agent accomplished its task, decomposes it into contextualized sub-claims grounded in trajectory evidence, estimates confidence for each terminal claim, and finally aggregates these into an overall confidence estimate. Across three agentic benchmarks, three backbone models, and three agent frameworks, CRGs yield better-calibrated confidence and stronger risk-aware decision making than verbalized, sampling-based, and white-box surrogate baselines. We further find that calibration error alone can be misleading: a white-box surrogate baseline appears well calibrated while providing near-chance discrimination. Ablations attribute CRG's improvements to claim-level confidence estimation and aggregation rather than graph construction alone. Finally, a CRG exposes the claims and trajectory evidence underlying each confidence estimate, enabling it to be audited at decision time.
Figures & tables
Figure 1: Overview of the CRG confidence estimation procedure ( § 4.1 - § 4.4 ). Given a single task and trajectory τ , a constructor LLM (1) decomposes and particularizes the task-success claim and then (2) attaches cited evidence from τ to the resulting leaf claims, using labels Supports , Unverified , and Undermines (two types shown). In (3), an estimator LLM assigns a confidence ψ(Gi) to each leaf given its evidence and the trajectory reduced according to the cited steps. In (4), confidence in each non-leaf claim is computed as the product of its children. CRGs support estimation and inspection: low-confidence paths can be traced to uncertain leaf claims and their associated evidence.
Benchmark
Num. Problems
Total Trajectories
GPT-5.5
Gemini-3.5 Flash
MiniMax-M3
Avg.
SWE-Bench Verified
306
900
73.9
74.3
71.7
73.3
EnterpriseOps-Gym
325
975
41.8
44.3
31.4
39.2
SkillsBench
81
219
50.7
50.7
45.2
48.9
Table 1: OpenHands success rates (%) by agent LLM. n is the number of problems.
SWE-Bench Verified
EnterpriseOps-Gym
SkillsBench
Estimator
LLM
ECE ↓
Brier ↓
AUROC ↑
BAS ↑
ECE ↓
Brier ↓
AUROC ↑
BAS ↑
ECE ↓
Brier ↓
AUROC ↑
BAS ↑
Basic Verbalizer
GPT-5.6 Sol
0.20
0.23
0.57
− 0.04
0.58
0.57
0.50
− 2.48
0.46
0.46
0.54
− 2.17
Qwen-3.8 27B
0.18
0.22
0.66
0.27
0.49
0.47
0.60
− 0.57
0.41
0.41
0.65
− 0.33
Reason-as-Graph
GPT-5.6 Sol
0.17
0.22
0.57
0.06
0.57
0.56
0.47
− 2.69
0.43
0.44
0.51
− 2.09
Qwen-3.8 27B
0.18
0.22
0.66
0.27
0.46
0.44
0.62
− 0.50
0.38
0.38
0.63
− 0.40
Verbal Consistency
Qwen-3.8 27B
0.24
0.25
0.54
− 7.58
0.45
0.45
0.62
− 13.72
0.45
0.44
0.60
− 13.87
Table 2: Confidence estimation results on OpenHands trajectories, averaged over agent LLMs. We report adaptive ECE (10 bins), Brier Score, AUROC, and BAS. Best per column in bold , second underlined . † White-box: uses token probabilities from a surrogate LLM.
Figure 2: Reliability diagrams (top) and empirical confidence CDFs (bottom) for CRG (Ours) , Reason-as-Graph , and Surrogate LNSP , using 10 equal-mass bins. Perfect calibration lies on the diagonal; under-confidence is above and overconfidence below (shaded). Dashed and solid vertical black lines indicate mean confidence and benchmark success rate, respectively. CRG is best calibrated on all three benchmarks, with estimates distributed across the confidence range.
Figure 3: Selective utility across cost thresholds t for the best method from each class (all Qwen-3.8 27B). The dashed zero line corresponds to abstaining from all solutions, so values below it are worse than abstention. All estimators coincide at low t , where accepting all solutions is optimal. As t increases, Reason-As-Graph incurs losses as large as −3.6 on EnterpriseOps-Gym, while CRG (Ours) remains above Surrogate LNSP through the mid range.
Codex (GPT 5.5)
Claude Code (Opus 4.7)
Estimator
ECE ↓
Brier ↓
AUROC ↑
BAS ↑
ECE ↓
Brier ↓
AUROC ↑
BAS ↑
Qwen 3.8 27B
Basic Verbalizer
0.19
0.24
0.63
0.22
0.19
0.23
0.64
0.25
Reason-as-Graph
0.19
0.23
0.63
0.23
0.19
0.23
0.67
0.26
CRG (ours)
0.05
0.20
0.61
0.38
0.11
0.22
0.54
0.34
Table 3: Transfer to other agent frameworks on SWE-Bench Verified. Codex (GPT-5.5) has n=305 trajectories and 72.5% success; Claude Code (Claude Opus 4.7) has n=304 and 72.0% success. All estimators use Qwen-3.8 27B. ECE is adaptive ECE with 10 bins.
Method
ECE ↓
Brier ↓
AUROC ↑
BAS ↑
Reason-as-Graph (no explicit graph)
0.38
0.37
0.73
− 0.23
Reason-with-Graph (no propagation)
0.35
0.36
0.61
− 0.16
CRG (graph + propagation)
0.10
0.23
0.68
0.17
Table 4: Ablating graph construction and confidence propagation on development data (Qwen-3.8 27B, n=495 ). Reason-as-Graph decomposes the task in the prompt and reports one confidence; Reason-with-Graph builds the graph but scores the root directly rather than propagating from leaves.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: CRGs for two successful EnterpriseOps-Gym trajectories. (a) A correct estimate: an HR-enrollment task decomposes into verifiable subtasks (case created, status set, survey linked, notification sent), each with high confidence. (b) An incorrect estimate, caused by incomplete graph construction: the agent’s first attempt used a status outside the allowed enum domain and was rejected, but no evidence item records the subsequent correction. The policy-compliance claim is therefore scored against the violation alone, receiving 0.10 and driving the root to 0.08 . Because the graph exposes which claim and which cited step produced the estimate, a user can inspect this reasoning and overrule it.
Figure 5: Cost–calibration tradeoff of each estimator: average cost per trajectory (USD) against adaptive ECE ( ↓ ), so lower-left is better. CRGs are the best-calibrated estimators on every benchmark, sharing the Pareto frontier with only the cheaper but worse-calibrated Surrogate LNSP . Importantly, CRGs are cost-efficient in absolute terms: 0.04–0.07pertrajectory,comparedtothecostofgeneratingthetrajectorybeingjudged,whichaverages0.15, 1.06,and0.95 on EnterpriseOps-Gym, SWE-Bench, and SkillsBench respectively.
SWE-Bench Verified
EnterpriseOps-Gym
SkillsBench
Estimator
LLM
5
8
10
15
20
5
8
10
15
20
5
8
10
15
20
Standard ECE (equal-width bins)
Basic Verbalizer
GPT-5.6 Sol
0.21
0.21
0.21
0.21
0.21
0.58
0.58
0.58
0.58
0.58
0.46
0.46
0.46
0.46
0.46
Qwen-3.8 27B
0.18
0.18
0.18
0.18
0.18
0.49
0.49
0.49
0.49
0.49
0.41
0.41
0.41
0.41
0.41
Reason-as-Graph
GPT-5.6 Sol
0.17
0.17
0.17
0.17
0.18
0.57
0.57
0.57
0.57
0.57
0.43
0.43
0.43
0.44
0.44
Qwen-3.8 27B
0.18
0.18
0.18
0.18
0.18
0.46
0.46
0.46
0.46
0.46
0.38
0.38
0.38
0.38
0.38
Appendix
Table 5: Sensitivity of confidence-estimation calibration error to the number of bins. ECE uses equal-width bins and Adaptive ECE uses equal-mass bins. Best per column in bold , second underlined . † White-box: uses token probabilities from a surrogate LLM.
Figure 6: Observed refinement depth vs. configured maximum refinement depth k on development data. Maximum and 95th-percentile observed depth stop growing at k=5 and k=4 , so we choose k=5 throughout our main experiments.
Effort
ECE
Brier
AUROC
Generated tokens
Low
0.08
0.19
0.78
5.3K
Medium
0.09
0.19
0.77
6.1K
Xhigh
0.08
0.19
0.79
10.4K
Appendix
Table 6: Sensitivity of the Qwen-3.8 27B CRG estimator to reasoning effort in both graph construction and leaf scoring.
SWE-Bench Verified
EnterpriseOps-Gym
SkillsBench
Estimator
LLM
ECE ↓
Brier ↓
AUROC ↑
BAS ↑
Mean Conf.
Std. Conf.
ECE ↓
Brier ↓
AUROC ↑
BAS ↑
Mean Conf.
Std. Conf.
ECE ↓
Brier ↓
AUROC ↑
BAS ↑
Mean Conf.
Std. Conf.
Action LNSP (first) †
Qwen-3.8 27B
0.30
0.30
0.48
0.28
0.43
0.08
0.22
0.31
0.49
− 0.02
0.61
0.15
0.14
0.27
0.52
0.13
0.41
0.11
Action LNSP (last) †
Qwen-3.8 27B
0.10
0.21
0.47
0.36
0.69
0.07
0.22
0.28
0.61
0.04
0.61
0.09
0.16
0.28
0.46
0.09
0.57
0.13
Action LNSP (mean) †
Qwen-3.8 27B
0.10
0.20
0.48
0.37
0.79
0.04
0.35
0.36
0.60
− 0.09
0.74
0.09
0.23
0.30
0.48
0.07
0.72
0.04
Action LNSP (min) †
Qwen-3.8 27B
0.34
0.32
0.51
0.26
0.39
0.08
0.13
0.26
0.56
0.07
0.53
0.12
0.15
0.26
0.58
0.14
0.37
0.09
Action SP (first) †
Qwen-3.8 27B
0.73
0.73
0.49
0.00
0.00
0.00
0.39
0.39
0.58
0.00
0.00
0.00
0.49
0.49
0.53
0.00
0.00
0.00
Appendix
Table 7: Surrogate log-probability aggregation results on OpenHands trajectories, averaged over agent LLMs. We report adaptive ECE (10 bins), Brier Score, AUROC, BAS, mean confidence, and confidence standard deviation. Best per ECE, Brier, AUROC, and BAS column in bold , second underlined . † White-box: uses token probabilities from a surrogate LLM.
EnterpriseOps-Gym ( n=192 )
SWE-Smith ( n=187 )
Aggregation rule
ECE ↓
Brier ↓
AUROC ↑
BAS ↑
ECE ↓
Brier ↓
AUROC ↑
BAS ↑
Arithmetic mean
0.51
0.47
0.71
− 0.43
0.25
0.28
0.57
0.08
Geometric mean (propagating)
0.46
0.42
0.71
− 0.35
0.25
0.28
0.57
0.08
Geometric mean (leaves)
0.45
0.41
0.72
− 0.33
0.25
0.28
0.57
0.08
Maximum
0.59
0.58
0.51
− 1.10
0.29
0.30
0.54
− 0.10
Upper Fréchet bound (minimum)
0.28
0.29
0.72
− 0.07
0.20
0.26
0.56
0.17
Appendix
Table 8: Comparing rules for combining leaf confidences into root confidence on development data (Qwen-3.8 27B). All rows use the same graphs, so differences isolate the aggregation rule. ECE denotes adaptive ECE with 10 bins.
SWE-Bench Verified
EnterpriseOps-Gym
SkillsBench
Estimator
LLM
ECE ↓
Brier ↓
AUROC ↑
BAS ↑
ECE ↓
Brier ↓
AUROC ↑
BAS ↑
ECE ↓
Brier ↓
AUROC ↑
BAS ↑
Basic Verbalizer + Temp. Scaling
Qwen-3.8 27B
0.19
0.22
0.66
0.34
0.16
0.26
0.60
0.07
0.14
0.25
0.65
0.15
Reason-as-Graph + Temp. Scaling
Qwen-3.8 27B
0.16
0.21
0.66
0.35
0.18
0.26
0.62
0.06
0.14
0.25
0.63
− 0.01
Surrogate LNSP † + Temp. Scaling
Qwen-3.8 27B
0.22
0.24
0.47
0.32
0.12
0.25
0.61
0.08
0.10
0.25
0.46
0.14
CRG (ours) + Temp. Scaling
Qwen-3.8 27B
0.13
0.21
0.57
0.36
0.07
0.23
0.62
0.10
0.06
0.23
0.65
0.16
Appendix
Table 9: Confidence estimation results on OpenHands trajectories, averaged over agent LLMs, after applying post-hoc temperature scaling. We tune the temperature parameter on the development set according to the procedure in Guo et al. (2017) . We report adaptive ECE (10 bins), Brier score, AUROC, and BAS. Best per column in bold , second underlined . † White-box: uses token probabilities from a surrogate LLM.
SWE-Bench Verified
EnterpriseOps-Gym
SkillsBench
CRG LLM
Judge
Suff.
Nec.
Partic.
Non-red.
Suff.
Nec.
Partic.
Non-red.
Suff.
Nec.
Partic.
Non-red.
GPT-5.6 Sol
GPT-5.6 Sol
0.98
0.79
0.75
0.82
0.96
0.77
0.65
0.95
0.90
0.86
0.69
0.84
Qwen-3.8 27B
0.92
0.71
0.41
0.88
0.85
0.65
0.37
0.96
0.85
0.71
0.53
0.86
Gemini-3.6 Flash
0.99
0.80
0.69
0.81
0.98
0.80
0.74
0.93
0.92
0.87
0.75
0.81
Qwen-3.8 27B
GPT-5.6 Sol
0.79
0.60
0.51
0.87
0.92
0.59
0.30
0.96
0.75
0.80
0.56
0.82
Qwen-3.8 27B
0.67
0.38
0.29
0.91
0.82
0.42
0.14
0.97
0.68
0.61
0.38
0.81
Appendix
Table 10: Macro-averaged graph-quality scores by benchmark, CRG LLM, and judge.
Benchmark
Baseline
ECE
Brier
AUROC
BAS
SWE-Bench Verified
Basic Verbalizer (GPT)
+0.115∗
+0.034∗
+0.006
+0.411∗
Basic Verbalizer (Qwen)
+0.095∗
+0.018
−0.084∗
+0.101∗
Reason-as-Graph (GPT)
+0.081∗
+0.019
−0.002
+0.318∗
Reason-as-Graph (Qwen)
+0.092∗
+0.017
−0.088∗
+0.102∗
Verbal Consistency
+0.154∗
+0.052∗
+0.030
+7.959∗
Surrogate LNSP
+0.009
+0.004
+0.098∗
+0.011
Appendix
Table 11: Paired task-bootstrap comparison of CRG (Qwen-3.8 27B) with every baseline in Table 2 . Entries are signed score differences: baseline minus CRG for ECE and Brier, and CRG minus baseline for AUROC and BAS. Positive values favor CRG. Exploratory mean rows average the six signed differences above them, with intervals recomputed on common task resamples. ∗ marks a difference whose nominal 95% bootstrap interval excludes zero, indicating statistical significance. No multiple-comparison adjustment is applied.