A latent dimension of Condorcet's jury theorem for multiple AI advisers
Organizations: Institute of Science Tokyo, Tokyo, Japan · Carnegie Mellon University, Pittsburgh, PA, USA
Abstract
When the same question is asked of multiple AI advisers, as in self-consistency and LLM-as-a-judge panels, Condorcet's jury theorem predicts that adding independent, competent advisers makes the majority more reliable. The theorem, however, has a latent dimension when viewed from the user's vantage: adding advisers also makes disagreement more visible. A binomial model reveals that this ``visible dissent'' becomes nearly inevitable as the number of advisers grows, and that reliability and disagreement approach certainty at rates that cross at an adviser accuracy of 4/5 (0.8); below it, visible dissent eventually becomes more likely than a correct majority. Even ideal panels of independent and competent advisers can be correct in aggregate but appear divided; such disagreement does not by itself indicate aggregation failure. The way advisers split also provides a common basis for predictive multiplicity, reconciliation load, and reliance miscalibration. These results separate aggregation from disclosure and turn the latter into testable questions about how disagreement should be presented and interpreted.
Figures & tables
| View | Focus | Question | Representative work |
|---|---|---|---|
| Oracle/System | Dependence, diversity and participation | What shapes the reliability of collective judgements? | Dependent software failures [ 30 ] ; algorithmic monoculture [ 31 ] ; ensemble diversity [ 32 , 33 ] ; interpreted versus generated signals [ 34 ] ; pooling medical judgements [ 35 ] ; confidence-based abstention [ 36 ] |
| Oracle/System and Observer/User | Vote distributions and multiplicity | How do verdicts differ across advisers and models? | Jury verdict-split models [ 26 , 27 ] ; ensemble diversity measures [ 37 ] ; predictive multiplicity [ 23 ] ; positive dissensus [ 38 ] |
| Observer/User | Uncertainty and inference | What can advisers’ outputs tell us about uncertainty, competence, or truth? | Deep ensembles [ 39 ] ; disagreement-based UQ [ 40 ] ; abstention on arbitrary predictions [ 41 ] ; abstention based on vote thresholds [ 42 ] ; posterior updating from jury size and margin [ 28 ] ; algebraic evaluation from answer patterns [ 43 ] ; consensus–accuracy laws across models [ 44 ] ; unanimity as evidence of systematic failure [ 45 ] ; correlation neglect [ 46 ] ; decisions with misleading correlated signals [ 47 ] |
| Observer/User | Visible dissent under the classical ideal assumptions ( this work ) | How often will users encounter a split as aggregate reliability improves? | Visible-dissent probability of independent, equal-accuracy, better-than-chance advisers compared with aggregate reliability; rate boundary ; disclosure reference scale |