LLM-based multi-agent systems (MAS) have attracted growing attention for improving reasoning through interaction among multiple agents. In this work, we focus on parallel multi-agent reasoning systems, where several agents solve the same problem over multiple rounds and aggregate their outputs into a final answer. Despite their strong reasoning performance, uncertainty estimation for such systems remains underexplored: the reliability of a MAS depends not only on individual generations, but also on how agents interact and evolve across rounds. We propose SAUCE (Sequential Agent Uncertainty through Consensus Evolution), a lightweight, training-free uncertainty estimator that formulates MAS uncertainty as sequential inference over a latent system-level belief. SAUCE aggregates round-level agreement and generation-uncertainty signals through a filtering-style update. Across five backbones, five benchmarks, and two MAS protocols, SAUCE improves misclassification detection, selective prediction, and calibration over a broad set of uncertainty estimation baselines, including standard log-likelihood-based methods and MAS-specific estimators.
Figures & tables
Figure 1 : Overview. (A) Two MAS trajectories can reach the same final answer but differ greatly in reliability: a stable run whose outputs converge across rounds (top) is more trustworthy than an unstable run whose outputs oscillate (bottom). A static view that only considers the final output cannot distinguish them, motivating uncertainty estimation from interaction dynamics. (B) SAUCE models multi-round interaction as a sequence of round-level observations (Mt,Rt) , where agreement is weighted by generation uncertainty, and performs a filtering-style update over a latent system-level belief. The resulting trajectory summary defines the SAUCE uncertainty score u(x,τ) (Eqn. 11 ).
Figure 2 : Round-level μt on Llama-3.1-8B / MATH-500. Both classes grow more confident with each round, but correct trajectories converge to a consistently higher consensus. Shading shows ±1 standard deviation across trajectories.
Figure 3 : Distribution of (a) the baseline agreement signal Mˉt and (b) SAUCE’s posterior-mean estimate μt (each averaged over rounds) for correct vs. incorrect MAS outputs (MMLU-Pro, Gemini-2.5-FL, DyLAN). SAUCE sharpens the separation between classes (AUROC 0.579→0.805 ). Densities are normalized independently in each panel.
Figure 4 : Probabilistic view of SAUCE. A latent system belief st evolves across rounds and emits a noisy observation ot formed from the agents’ outputs. SAUCE’s per-round estimate is the posterior expectation E[st∣o1:t]=μt .
Method
MATH-500
MMLU-Pro
BBH
Avg
Debate
DyLAN
Debate
DyLAN
Debate
DyLAN
ROC
ARC
ROC
ARC
ROC
ARC
ROC
ARC
ROC
ARC
ROC
ARC
ROC
ARC
Qwen3-4B
PE
64.89
83.45
64.58
82.70
59.54
73.63
63.53
74.92
68.90
87.13
59.13
83.79
63.43
80.94
LL
65.93
83.14
66.36
84.08
68.39
78.49
69.88
79.27
69.56
87.62
60.61
83.82
66.79
82.74
UDPO
57.29
79.14
72.50
84.10
58.22
66.43
67.95
73.43
59.18
84.38
64.97
88.00
63.35
79.25
Table 1 : Misclassification detection (AUROC ↑ ) and selective prediction (AUARC ↑ ) on open-source models. Boldface and underlining denote the best and second-best performance.
Method
Olympiad
MMLU-Pro
BBEH-mini
Avg
Debate
DyLAN
Debate
DyLAN
Debate
DyLAN
ROC
ARC
ROC
ARC
ROC
ARC
ROC
ARC
ROC
ARC
ROC
ARC
ROC
ARC
Gemini-2.5-Flash-Lite
PE
65.51
66.86
66.35
65.23
48.71
73.35
70.22
89.09
49.98
19.13
41.73
16.06
57.08
54.95
LL
68.67
68.13
67.42
66.67
64.90
82.82
72.18
90.05
46.71
18.50
39.41
15.46
59.88
56.94
UDPO
63.08
48.89
81.58
71.05
54.59
77.52
57.83
82.71
59.23
25.12
75.41
27.29
65.29
55.43
Table 2 : Misclassification detection and selective prediction on closed-source models. Metrics and notation follow Table 1 . Boldface and underlining denote the best and second-best performance.
Figure 5 : Brier score after Platt scaling (5-fold cross-validation) on MMLU-Pro under the Debate protocol, across all five backbones. SAUCE attains the lowest Brier-Platt score on four of the five backbones.
Figure 6 : Per-round behavior on a 15-round Qwen3-4B Debate run on MMLU-Pro, grouped by final-answer correctness. Left : SAUCE posterior mean μt ; the correct/incorrect gap persists through round 15. Middle : Average per-agent log-likelihood at round t ; both classes become increasingly token-confident and the gap vanishes. Right : AUROC against final-consensus correctness as a function of debate length T ; SAUCE plateaus while LL degrades toward chance (dashed line).
Method
AUROC
Brier-Platt
Last round, 1−Mt
58.4
21.3
(a) Round-mean of Mt
68.6
20.0
(b) Static joint with Rt
68.4
20.9
(c) SAUCE (Ours)
77.1
19.2
Table 3 : Component-wise ablation on Qwen3-4B / MMLU-Pro / Debate. Row (b) is the static joint of Mt and Rt ; row (c) is the full sequential update with μˉ readout. AUROC ↑ and Brier-Platt ↓ .
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Format
Open-source size
Closed-source size
MATH-500
Free-form math
500
–
MMLU-Pro
Multiple-choice
12,032
5,000 (subset)
BBH
MCQ / short string
1,250
–
OlympiadBench
Free-form math
–
674
BBEH-mini
MCQ / short string
–
≈460
Appendix
Table 4 : Benchmark formats and evaluation sizes used in the main results. “–” denotes a benchmark not used for that model class.
MAS protocol
λ
q
σ02
μ0
Debate
0.95
0.10
0.25
0.0
DyLAN
0.70
0.001
0.25
0.0
Appendix
Table 5 : SAUCE hyperparameter defaults transferred across all cells.
Method
MATH-500
MMLU-Pro
mean (Rt) – no filter
77.00
64.93
Learned λθ(ht)
82.85
75.18
Const. λ∗ (ours)
82.94
74.76
Appendix
Table 6 : Learned vs. constant transition coefficient on Qwen3-4B / Debate (AUROC ↑ , ×100 ). The two filter variants achieve similar AUROC on both datasets; both improve substantially over the no-filter baseline.
Figure 7 : Per-axis sensitivity of SAUCE’s four hyperparameters at the calibration cell (Qwen3-4B / MMLU-Pro). Top row: Debate, selected values (λ,q,σ02,μ0)=(0.95,0.10,0.25,0.0) . Bottom row: DyLAN, selected values (0.70,0.001,0.25,0.0) . In each panel, one hyperparameter is swept while the other three are held at their per-protocol selected value (marked by ★ and a dotted vertical line).
Method
MATH-500
MMLU-Pro
BBH
Avg
Debate
DyLAN
Debate
DyLAN
Debate
DyLAN
Qwen3-4B
PE
16.92
17.06
22.44
21.64
15.23
16.48
18.30
LL
17.12
16.97
22.46
21.65
15.41
16.55
18.36
UDPO
17.16
15.13
21.40
19.84
15.54
15.85
17.49
MATU
17.86
17.94
22.38
22.34
16.32
16.16
18.83
Appendix
Table 7 : Calibration quality on open-source models, measured by Brier score after Platt scaling with 5-fold cross-validation ( ↓ , ×100 ). Bold denotes the lowest score per (model, protocol, dataset) cell, and per-backbone Avg.
Method
OlympiadBench
MMLU-Pro
BBEH-mini
Avg
Debate
DyLAN
Debate
DyLAN
Debate
DyLAN
Gemini-2.5-Flash-Lite
PE
22.35
21.87
18.70
14.84
17.21
15.05
18.34
LL
22.47
22.18
14.83
13.53
17.18
15.56
17.63
UDPO
21.60
16.63
18.47
15.93
16.80
14.21
17.27
MATU
22.02
23.45
18.62
15.93
17.06
14.56
18.61
Appendix
Table 8 : Calibration quality on closed-source models, measured by Brier score after Platt scaling with 5-fold cross-validation ( ↓ , ×100 ). Notation follows Table 7 .
Method
Qwen3-4B
GPT-5.4-mini
MATH-500
MMLU-Pro
BBH
OlympiadBench
MMLU-Pro
BBEH-mini
EWMA
66.80
68.43
69.56
81.63
66.95
53.46
Slope
37.79
38.76
40.75
32.36
39.56
57.25
Switches
61.76
60.59
61.47
71.08
57.58
52.43
TTC
66.47
68.29
68.29
79.97
66.31
56.94
Mean 1−Mt
66.60
68.56
69.72
82.33
67.08
55.95
Appendix
Table 9 : Simpler trajectory summaries under Debate (AUROC ↑ , ×100 ). Boldface and underlining denote the best and second-best performance. Evaluation sizes and score definitions follow Appendix C.4 .
Method
Qwen3-4B
GPT-5.4-mini
MATH-500
MMLU-Pro
BBH
OlympiadBench
MMLU-Pro
BBEH-mini
Original
68.59
76.42
72.86
83.83
69.53
49.96
Reversed
68.30
71.68
74.26
82.29
65.21
49.72
Shuffled
68.41
72.45
73.84
82.99
66.22
49.12
Δ
+0.19
+3.97
−1.41
+0.84
+3.31
+0.24
Appendix
Table 10 : Temporal-order ablation under Debate (AUROC ↑ , ×100 ). Shuffled averages three seeds; Δ is original minus the better alternative. Best and second-best scores are bold and underlined.
Statistic
Qwen3-4B
GPT-5.4-mini
MATH-500
MMLU-Pro
BBH
OlympiadBench
MMLU-Pro
BBEH-mini
SAUCE
68.59
76.42
72.86
83.83
69.53
49.96
95% CI
62.6–74.7
75.6–77.3
69.4–76.3
80.7–87.1
67.4–71.5
43.8–56.3
vs. PE
0.011
<0.001
0.019
<0.001
<0.001
0.121
vs. LL
0.020
<0.001
0.034
<0.001
<0.001
0.002
vs. UDPO
<0.001
<0.001
<0.001
<0.001
<0.001
1.000
Appendix
Table 11 : Paired-bootstrap comparisons under Debate. The first two rows report SAUCE’s AUROC ( ×100 ) and 95% interval; subsequent rows give its tail fraction against each baseline. Bold indicates p<0.05 ; <0.001 denotes zero tail counts in 1,000 resamples.
Method
Qwen3-4B
GPT-5.4-mini
MATH-500
BBH
OlympiadBench
BBEH-mini
PE
67.04 ± 1.64
69.83 ± 0.99
53.65 ± 1.92
46.11 ± 1.31
LL
67.67 ± 1.22
71.57 ± 1.59
70.17 ± 0.68
43.34 ± 2.54
SAUCE (Ours)
72.26 ± 2.29
72.58 ± 1.76
82.79 ± 0.87
49.01 ± 1.51
Runs
5
5
6
5
n per run
500
1,250
673–674
460
Appendix
Table 12 : Repeated Debate rollouts (AUROC ↑ , mean ± sample standard deviation, ×100 ). Boldface and underlining denote the best and second-best means.
Switches
n
Accuracy (%)
Mean confidence
0
8,004
76.9
0.915
1
1,595
49.5
0.758
2
486
44.4
0.676
3+
611
38.8
0.622
Appendix
Table 13 : Answer-switch analysis among final-unanimous trajectories on Qwen3-4B / Debate / MMLU-Pro ( n=10,696 ). Switches count changes in individual agents’ answers.
LLM-based multi-agent systems (MAS) solve complex tasks through communication among role-specialized agents. However, inter-agent dependencies introduce reliability risks beyond isolated agent failures. For instance, errors in intermediate messages could be inherited and amplified by downstream agents. Existing uncertainty quantification (UQ) methods mainly target isolated responses or single-agent reasoning, and therefore fail to capture uncertainty propagation in MAS. To this end, we propose PropUQ-MAS, an error propagation-aware UQ framework that represents MAS execution as a communication-structured graph and estimates each step's reliability by combining local uncertainty with uncertainty inherited from upstream messages. Extensive experiments demonstrate that PropUQ-MAS consistently improves UQ in MAS, with average relative gains of +6.10% in AUROC and +47.58% in PRR.
Yaokun Liu, Yifan Liu, Daniel Yue Zhang +3
University of Illinois Urbana-Champaign · Scale AI
This paper investigates how multi-agent systems (MAS)-based on large language models (LLMs) can support actuarial risk modelling, with a particular focus on uncertainty quantification. Actuarial workflows represent a high-stakes decision-support setting where unreliable outputs may lead to incorrect risk assessment, unfair pricing, and regulatory non-compliance. To address uncertainty introduced by the probabilistic nature of LLMs and dependencies between agents, a multi-agent framework is proposed in which specialised agents perform data preparation, modelling, review, and explanation tasks under a central hub. The main contribution is a novel approach to uncertainty propagation using token-level log-probabilities and a Bayesian Network. Importantly, log probabilities are not treated as direct probabilities of correctness or task success. Instead, length-normalised log-probability summaries are transformed into calibrated task-level confidence estimates before incorporation into the Bayesian Network. Results show that the framework reproduces baseline actuarial performance while providing additional insight into workflow stability and runtime uncertainty propagation.
Multi-agent systems (MAS) assume that collaborating inherently improves Large Language Model (LLM) reasoning. We challenge this by demonstrating that simulated social pressure triggers an algorithmic Bystander Effect,'' inducing severe cognitive loafing. By evaluating 22,500 deterministic trajectories across 3 dataset contexts (GAIA, SWE-bench, Multi-Challenge) with 3 state-of-the-art (SOTA) models, we semantically audit internal reasoning traces against external outputs. We formalize the \textit{Interaction Depth Limit} ($D_L$), the exact plurality threshold where an agent's logical sovereignty collapses into social compliance. Crucially, we uncover the \textit{Sovereignty Gap}: models frequently compute the correct derivation internally but suffer Alignment Hallucinations'' -- actively subjugating empirical evidence to sycophantically appease a simulated swarm. We prove that multi-agent social load is strictly non-commutative; the "brand" identity of the ``Lead Anchor'' auditor disproportionately dictates the swarm's integrity. These findings expose architectural vulnerabilities, proving that unstructured multi-agent topologies can degrade independent reasoning.