A multi-LLM \emph{council} lets several large language models (LLMs) deliberate on a question and return an answer together with a confidence estimate. As these systems become increasingly used for reasoning, that confidence should represent a calibrated \emph{probability of being correct}, and the decision should remain robust when some agents are persistently unreliable. Existing \emph{council aggregation} methods fail on both fronts: their confidence estimates measure decisiveness rather than correctness, and they cannot identify or discount persistently unreliable agents. We introduce Bayesian Dialectical Argumentation (BDA), which treats the council's \emph{typed} moves---who proposed, challenged, or conceded which answer---as observations of a classical annotator model with \emph{per-agent} reliabilities. This formulation recasts multi-agent deliberation as a reliability estimation problem, using the deliberation trace to infer agent reliability under persistent adversarial behavior. By weighting evidence according to inferred agent reliability, BDA yields calibrated posterior probabilities over candidate answers while allowing persistently unreliable agents to be inverted rather than merely outvoted. Across binary and multi-class benchmarks, BDA achieves the best calibration among zero-cost council aggregation methods, requiring no additional LLM calls, and improves robustness under persistent adversarial coalitions while remaining competitive in clean settings.
Figures & tables
Figure 1: The paper in one picture. Top—the mechanism. Two distrusted seats (red) and one genuine (blue) deliberate on a task with truth g . BDA reads typed moves as per-seat tallies ci:di , weighs them by learned reliability wi , and assigns negative weight to distrusted seats, so rejecting g becomes evidence for g : a calibrated π(g∣T)=0.89 where plurality is wrong. Bottom—the payoff. Among zero-cost aggregators, the BDA family occupies the calibration frontier, while stacking, gated vote, DS-EM, and paid baselines trail. Robustness under attack is shown in Figure 2 .
Figure 2: The fair, role-matched attack : the same anti-correct adversary mounted inside each method’s own generation pipeline ( K=2,3,4 ; full sets, three seeds; error bars are seed standard deviations). Self-MoA, with no internal check, collapses; ArgLLMs’ clean evaluator is a partial defense that erodes as the answer space grows; the council, facing the adversary as two of its three seats , holds near clean. Protocol in Appendix F.
Binary
PubMedQA
MMLU
Aggregator
k=1
k=2
k=1
k=2
k=1
k=2
Majority
.573
.369
.442
.209
.671
.433
Log. stacking
.814
.865
.722
.753
.813
.743
Gated vote
.751
.661
.596
.570
.767
.683
DS-EM
.664
.178
.564
.282
.774
.001
MoA
.561
.483
.380
.227
.733
.661
Table 1: Accuracy under the live stealth coalition ( k persistent adversarial seats of three; three seeds). Clean references (majority/BDA per-agent): binary 0.734/0.751 ; K>2 in Appendix E. The regime split: the confusion leads binary and PubMedQA, the stacker MMLU (same budget, zero cost). Methods marked † never read the council (Appendix B), so their cell is their clean accuracy (invariance within 1.1 points)—abstention (Figure 2 ).
Aggregator
Acc.
ECEew
ECEem
Brier
$/task
Majority
0.734
0.160
0.154
0.212
0
Logistic stacking
0.751
0.120
0.118
0.194
0
Gated vote
0.726
0.150
0.149
0.210
0
DS-EM (unsup.)
0.659
0.338
0.339
0.338
0
MoA
0.718
—
—
—
.0020
Self-MoA
0.736
—
—
—
.0017
Table 2: Clean binary results ( 4,500 trials, three seeds; dashes: no informative confidence; every free arm shares the 25 -label budget). Rows follow the grouping of Section 5 , baselines first. ArgLLMs leads accuracy and pays per task; of the free arms BDA leads calibration and beats plurality’s decision ( Δ Acc +0.017 ; pooled and question-clustered 95% CIs both [+0.010,+0.024] ).
Figure 3: The label budget is not the bottleneck. Left: held-out accuracy per supervised arm with n=10 labels (open), n=25 —the operating budget—(filled), and all available labels (tick); a filled dot sitting on its tick is the claim. The one arm that needs its 25 is the MMLU stacker; only the confusion’s PubMedQA accuracy still gains beyond the budget ( +0.05 ), its K×K rows being the hungriest fit. Right: what the same 25 labels already buy in calibration: on K>2 the confusion is 4× better-calibrated than the one-coin ( 0.09 vs. 0.36 ECE on PubMedQA). Majority (dotted) uses no labels. Full grid in Appendix E.
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Meaning
Symbol
Meaning
one-coin
uniform-error likelihood (main Eq. 1); as an arm name, BDA per-agent (priors of main Eq. 3).
B(a,b)
Beta function Γ(a)Γ(b)/Γ(a+b) .
confusion
learned full-matrix composite likelihood (main Eq. 4); the arm BDA confusion.
ri,rˉ
fitted reliability of i ; council mean.
ECEew/em
ECE, equal-width / equal-mass binning.
μi,s
prior mean and strength; αi=μis .
θi∈[0,1]
latent reliability of agent i .
Pi(g′∣g)
agent i ’s confusion: prob. a move is consistent with g′ under truth g (main Eq. 4).
ci(g),di(g)
moves of i consistent / inconsistent with g ; ci(g′) generalizes to every candidate (main Eq. 4).
Table 3: Notation used throughout the paper and appendix, in order of first use (left column, then right).
Method
Evidence GenΦ
Rule AggΦ
Reads T ?
ρ corrupts
γΦ
Majority
Propose answers ( T )
plurality count
yes
2 of 3 seats
2/3
Logistic stacking
per-seat tallies ( T )
logistic regression
yes
2 of 3 seats
2/3
Gated vote
Propose answers ( T )
gated plurality
yes
2 of 3 seats
2/3
DS-EM
Propose answers ( T )
EM posterior
yes
2 of 3 seats
2/3
MoA / MoA-Lite
proposer surfaces ( T )
synthesizer LLM
yes
2 of 3 seats
2/3
Self-MoA
6 samples of one model on q
synthesizer LLM
no
sampled model
1
Appendix
Table 4: Every method as AggΦ∘GenΦ (Eq. 7 ). Reading the council is dependence on the trace T ; the fair attack applies the single role ρ (Eq. 8 ) to whichever LLM the method’s evidence comes from, leaving its aggregation rule intact. Here γΦ is the fraction of that evidence corrupted: it is 2/3 for the redundant council readers (one clean seat remains) and all of it for the single-pipeline methods.
Arm
PubMedQA
MMLU
ArgLLMs (canonical, clean)
100%
100%
MoA (clean)
99.4%
84.9%
Self-MoA (clean)
99.9%
99.9%
MoA (attacked, k=1/2 )
100/100%
99.9/100%
MoA-Lite (attacked)
100%
≥99.9 %
Appendix
Table 5: Parse rates: the fraction of answers committing to a candidate (after the LLM-judge pass on the prose-only residual). Unparseable answers score 0 . Read faithfully (multi-option dispatch), ArgLLMs returns a canonical option label by construction, so its parse rate is 100% ; this replaces the artifactual 62% / 92% of a wrapper that returned proposer paragraphs.
Aggregator
Acc.
ECEew
ECEem
Brier
$/task
Majority
0.734±0.005
0.160±0.002
0.154±0.002
0.212±0.002
0
Verbalized conf.
0.732
0.117
0.116
0.202
0
Best member
0.744
0.067
0.075
0.182
0
MoA
0.718±0.010
—
—
—
.0020
Self-MoA
0.736±0.003
—
—
—
.0017
ArgLLMs
0.788±0.008
0.165±0.007
0.167±0.010
0.199±0.004
.0015
Appendix
Table 6: Clean binary pool ( 4,500 trials); cells are mean ± standard deviation across the three seeds. Per-agent versus majority: Δ Acc =+0.0171 , 95% CI [+0.0104,+0.0240] , pwin=1.000 (paired bootstrap). “Verbalized conf.” is a confidence-weighted vote reporting the winner’s mean stated confidence; “best member” is the strongest seat, selected leak-free (Appendix E.3 ).
Figure 4: Reliability diagrams for every confidence-emitting arm, clean binary pool ( 4,500 trials; gray bars are bin populations; the MoA family emits near-constant scores, not applicable). The BDA posteriors (bottom right pair, the confusion model almost exactly) hug the diagonal; a vote fraction, a strength margin, an unsupervised mode, a gate, and a raw stacking softmax do not.
Figure 5: Equal-width calibration error per (dataset, seed) cell, clean binary family. The BDA rows are uniformly light across all nine cells.
Ablation
Setting
Acc.
ECEem
Brier
granularity
per-agent
0.722
0.134
0.212
per-force
0.703
0.266
0.279
per-(agent, force)
0.703
0.258
0.273
estimator
all moves
0.730
0.080
0.195
propose only
0.714
0.117
0.210
round-0 only
0.722
0.150
0.218
Appendix
Table 7: Reliability-key and reliability-estimator ablations (binary pool, three seeds). Granularity uses a fixed prior (b,κ,s)=(0.58,2,16) ; the estimator selects its hyperparameters by inner cross-validation. The per-agent key and the all-moves estimator win.
evidence window
Acc.
ECEem
decisions changed
≤ round 0
0.748
0.044
—
≤ round 1
0.754
0.035
vs. round 0: +
≤ round 2 (full)
0.751
0.044
net +17/4500
Appendix
Table 8: Binary deliberation is roughly decision-neutral; its net value rises with K (Appendix E.4 ).
Dataset
Arm
Acc.
ECEew
ECEem
Brier
PubMedQA
Majority
0.619
0.264
0.263
0.294
BDA shared
0.607
0.311
0.304
0.327
BDA per-agent
0.614
0.304
0.295
0.326
BDA confusion
0.604
0.072
0.085
0.229
MMLU
Majority
0.770
0.093
0.093
0.150
BDA shared
0.759
0.190
0.186
0.193
Appendix
Table 9: Clean K>2 : plurality wins accuracy narrowly; the one-coin BDA is over-confident (the concentration of Proposition 3 without heterogeneity to earn it); the tempered confusion model repairs the calibration at unchanged accuracy.
Figure 6: The fix/break ratio of deliberation rounds—decisions the round- 1 -and-later moves repair versus damage—rises with the number of candidate answers K .
selected
clean ECEew
Cell
λ
τ
one-coin
temp. 1-coin
confusion
Binary
0.3
0.25
0.016
0.012
0.017
PubMedQA
27.0
0.15
0.304
0.051
0.072
MMLU
93.4
0.24
0.182
0.037
0.063
Appendix
Table 10: Left: the confusion model’s CV-selected concentration λ and temperature τ (mean over folds and seeds); λ is large on clean K>2 , so the learned matrix is shrunk toward the one-coin. Right: a CV-tempered one-coin already matches the confusion’s clean K>2 calibration—the fix is the temperature. One-coin denotes the BDA per-agent arm.
one-coin
confusion
log. stacking
gated vote
Dataset
n
acc
ECEew
acc
ECEew
acc
ECEew
acc
ECEew
Binary
25
0.707
0.064
0.741
0.086
0.712
0.080
0.721
0.163
100
0.718
0.077
0.736
0.089
0.711
0.125
0.716
0.163
500
0.722
0.048
0.732
0.038
0.710
0.140
0.716
0.163
PubMedQA
25
0.577
0.355
0.517
0.091
0.580
0.099
0.630
0.256
100
0.573
0.344
0.560
0.116
0.580
0.139
0.630
0.256
Appendix
Table 11: Label efficiency of the whole supervised family on clean pools (disjoint calibrate/holdout, three seeds; full grid in the results file). Accuracy is flat from tiny n for every arm; the confusion’s K>2 calibration advantage is present already at n=25 ; the gated vote’s chance-gate saturates immediately; the stacker’s binary ECEew drifts up with n as the stack overfits its tallies. One-coin denotes the BDA per-agent arm.
pwrong
Arm
0
.2
.4
.5
.6
.8
1.0
persistent, k=1
Majority
.734
.705
.671
.657
.640
.611
.581
Shared
.730
.712
.693
.686
.676
.660
.638
Per-agent
.751
.710
.717
.722
.755
.824
.923
Confusion
.758
.732
.718
.719
.723
.748
.792
Appendix
Table 12: Synthetic count-level adversary, accuracy (binary pool, three seeds). Both per-agent parameterizations are non-monotone and then rising under a persistent adversary. Under rotation the centered one-coin falls monotonically while the confusion model’s absolute rows recover at high dose (council-level identifiability); the unlearnable regime is moderate-dose rotation.
Binary
PubMedQA
MMLU
ECEem
k=1
k=2
k=1
k=2
k=1
k=2
Majority
.237
.493
.338
.558
.167
.332
Log. stacking
.097
.047
.124
.066
.108
.070
BDA shared
.234
.352
.427
.755
.239
.620
BDA per-agent
.143
.152
.349
.389
.234
.246
BDA confusion
.082
.053
.090
.045
.072
.109
Appendix
Table 13: Calibration under the live attack (council-reading arms). The tempered confusion model is genuinely calibrated under fire; MoA-family confidences remain degenerate (not applicable); ArgLLMs is council-invariant and keeps its clean calibration, poor on K>2 .
one-coin
confusion
Dataset
k
r0
full
Δ
r0
full
Δ
Binary
k=1
0.914
0.794
-0.120
0.914
0.812
-0.102
Binary
k=2
0.933
0.736
-0.197
0.936
0.892
-0.045
PubMedQA
k=1
0.670
0.608
-0.062
0.873
0.744
-0.129
PubMedQA
k=2
0.588
0.471
-0.117
0.896
0.840
-0.056
MMLU
k=1
0.796
0.733
-0.063
0.813
0.752
-0.061
Appendix
Table 14: Accuracy under the live coalition on round- 0 -truncated (r0) versus full traces, leak-free refit per basis, three seeds. Round- 0 one-coin is the classical supervised weighted vote (main Proposition 1). Deliberation rounds reduce attacked accuracy in eleven of twelve cells: they are the contagion channel. One-coin denotes the BDA per-agent arm.
Figure 7: The sleeper contamination axis. Top: the monitoring statistic of Appendix F.13 —total-variation drift of the fitted confusion rows against the clean window—exits its noise floor by f≈0.2 : the alarm precedes the recovery below. Bottom: per-arm accuracy under the persistent k=2 coalition versus the fraction f of the calibration window in which the seat has turned ( f=0 is the pure sleeper; all supervised arms start in the below-majority band). The one-coin recovers fastest at low f ; the confusion needs more of the turn in-window but overtakes from f≈0.6 and ends highest—the bias–variance trade-off of Appendix C.5 . Three seeds.
Cell
Arm
Acc.
ECEew raw → iso
Br. iso
clean bin.
Majority
0.734
0.160 → 0.006
0.184
BDA per-agent
0.751
0.011 → 0.005
0.182
BDA confusion
0.758
0.012 → 0.013
0.172
clean pub.
Majority
0.619
0.255 → 0.067
0.228
BDA per-agent
0.614
0.301 → 0.030
0.234
BDA confusion
0.604
0.043 → 0.028
0.228
Appendix
Table 15: Leak-free isotonic recalibration. On clean data it equalizes calibration (BDA’s edge vanishes); under attack it perfects calibration but cannot rescue a flipped vote’s accuracy, so BDA confusion’s accuracy lead survives. Interim, not a headline.
Init
clean
k=1
k=2
majority
0.659
0.668
0.175
spectral
0.659
0.668
0.175
anti (oracle mode)
0.341
0.332
0.825
switch rate
1.00
1.00
1.00
Appendix
Table 16: Unsupervised DS–EM accuracy by initialization (seed 0 ; the collapse is structural across seeds). At k=2 the deployable (majority/spectral) modes are the coalition’s mirror; only the label-requiring “anti” mode is correct.
PubMedQA k=2
MMLU k=2
Arm
orig.
permuted
orig.
permuted
Logistic stacking
0.753
0.750
0.743
0.735
BDA per-agent
0.471
0.478
0.699
0.699
BDA confusion
0.840
0.740
0.721
0.716
Appendix
Table 17: Per-task candidate-label permutation, live k=2 (clean cells are invariant for every arm). The label-blind arms do not move; the confusion model loses exactly its label-identity component on the semantic dataset and nothing on the positional one.
PubMedQA
MMLU
cal → dep
1-coin
conf.
stack
1-coin
conf.
stack
A → A
0.600
0.676
0.682
0.707
0.693
0.713
B → B
0.487
0.842
0.722
0.695
0.712
0.731
A → B (shift)
0.531
0.553
0.602
0.584
0.532
0.632
B → A (shift)
0.607
0.578
0.571
0.708
0.528
0.319
Appendix
Table 18: Wrong-answer policy shift at k=2 (A = synthetic random-wrong, B = live best-defensible; disjoint calibrate/deploy rows, three seeds). Across policies both learned-structure arms (confusion, stacker) lose their edge—the stacker catastrophically on MMLU B → A—while the one-coin (the BDA per-agent arm) is policy-robust in every shifted cell.
genuine flip
effective vs nominal
Cell
k=1
k=2
k=1
k=2
Binary
0.109
0.239
0.48/0.33
0.68/0.67
PubMedQA
0.186
0.353
0.59/0.33
0.72/0.67
MMLU
0.029
0.122
0.44/0.33
0.64/0.67
Appendix
Table 19: Contagion: genuine-seat flip rate to the coalition’s answer, and the resulting effective corrupted fraction against the nominal k/I .
seat
ri
μi
wi
(c,d)
Δℓi(True)
A1 (adv)
0.21
0.37
−0.55
(0,3)
+1.18
A2 (adv)
0.29
0.44
−0.24
(0,1)
+0.24
A3 (gen)
0.60
0.75
+1.12
(2,1)
+0.71
total log-odds
+2.12
Appendix
Table 20: The worked example. The two distrusted seats argue False , but their negative weights turn those moves into evidence for True —more from the more distrusted A1 . Even though the genuine seat flipped once (its (2,1) leaves only a weak +0.71 ), the two inverted adversaries carry +1.42 of the +2.12 total, which gives π(True∣T)=0.89 —a calibrated confidence, not a saturated vote.
Dataset
Arm
Acc.
ECEew
ECEem
Brier
Binary
Majority
0.692
0.151
0.151
0.223
BDA shared
0.684
0.024
0.041
0.202
BDA per-agent
0.701
0.022
0.029
0.197
BDA confusion
0.697
0.031
0.039
0.199
PubMedQA
Majority
0.574
0.285
0.281
0.317
BDA shared
0.558
0.315
0.310
0.341
Appendix
Table 21: Small council, clean. On binary, per-agent wins everything and the calibration-error reduction over voting is about sevenfold (per-agent versus majority Δ Acc +0.009 , pwin=0.96 ). On K>2 , with near-chance, near-homogeneous reliabilities there is no per-seat accuracy signal and plurality leads the decision—the heterogeneity assumption as a real boundary—while the tempered confusion model still repairs the calibration ( ECEew0.033 / 0.050 against plurality’s 0.285 / 0.157 ), the temperature needing no heterogeneity.
School of Computer Science & Technology, Beijing Jiaotong University, Beijing, China · Beijing Key Laboratory of Traffic Data Mining and Embodied Intelligence, Beijing, China