Organizations: School of Systems Science and Industrial Engineering, Binghamton University, Binghamton, NY, USA · School of Management, Binghamton University, Binghamton, NY, USA · IBM T. J. Watson Research Center, Yorktown Heights, NY, USA
Multi-agent systems of LLMs add discussion to majority voting and are therefore expected to be more capable. However, empirical reports conflict on whether discussion improves accuracy or leads to an incorrect consensus. Here, we introduce a parsimonious model that explains when discussion improves accuracy and when it ends in an incorrect consensus, built from four behaviors repeatedly observed in LLM agents: (1) withholding dissent, (2) internalizing a stated answer, (3) reconsidering after seeing dissent, and (4) correcting toward the correct answer. The model shows that discussion can overturn an incorrect initial majority only when the withholding rate c is below a critical rate c∗=γ/(γ+a), set by the net correction rate γ and the internalization rate a. We estimate these rates from conversation logs with a Bayesian method and place LLM teams relative to c∗. As the model predicts, the gain from discussion shrinks as withholding rises, across LLMs and on a hidden profile benchmark, HiddenBench, and MedEInst. Instructing agents not to withhold dissent increases this gain. Turning reasoning off also increases the gain, because reasoning raises the internalization rate a and keeps agents from reconsidering a minority answer. These findings reconcile the conflicting reports and identify when discussion outperforms majority voting.
Figures & tables
Figure 1: (a) Each agent holds a private belief z (inner circle) and states a public opinion y (outer ring), and it sees the statements of the agents it observes and their majority mi ( q=4 in the schematic). (b) When the visible majority differs from its belief, an agent either withholds (probability c ), states the majority, and adopts it with probability a , or states its belief, reconsiders with probability ρdi , and reaches the right answer with probability r . (c) Share of teams ending on the correct answer against the single-agent accuracy p and the withholding rate c : agent-based simulation (color) and the mean-field boundary (black curve) agree. Below p=1/2 , discussion reaches the correct answer only while c stays below c∗ ( a=0.5 , ρ=1 , r=0.8 , q=3 ).
Figure 2: (a) Hidden profile task: only the pooled private clues reveal the correct answer. (b) Withholding rate c^ and (c) critical value c∗ per LLM under instructions A (honest, blue), B (none, grey), and C (value cohesion, open). (d) Gain from discussion against the margin c^−c∗ ; marker shape is the LLM, striped marks have reasoning switched off, and the curve is a quadratic fit with its 95% bootstrap band. The gain falls as the margin grows. (e–g) Internalization a^ , net correction γ^ , and gain from discussion with reasoning on (filled) and off (striped), under instruction A. Intervals are 95%.
Figure 3: (a–c) Gain from discussion against the margin c^−c∗ on HiddenBench (HB), MuSiQue (MuS), and MedEInst (MedE), with markers as in Fig. 2 d and a linear fit (orange) with its 95% band; settings whose LLM solved no item and the MedEInst run of qwen3.8 are not drawn (Appendix B.1 ). The gain falls as the margin grows, clearly on HB and MedE and weakly on MuS. (d) Effects of the benchmark, the LLM, and the instruction (relative to B) on the logit of the withholding rate, from one fit over the four benchmarks (HP: our hidden profile task). (e) Share of the shared evidence that supports the correct answer (black) or the agent’s own answer (orange) on the hidden profile task. (f) Evidence pooling and utilization rates on the hidden profile task under instruction A. Intervals are 95%.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Evidence pooling rate (left) and evidence utilization rate (right) on the hidden profile task, one row per model in the order of Fig. 3 e, with the three instructions on each row (A blue, B grey, C open). Intervals are 95% Wilson intervals.
Figure 5: Reasoning switched on (filled) and off (open) in the same model on the hidden profile task, one row per model, pooled over A, B and C, unlike the instruction-A panels of Fig. 2 e–g: (a) withholding rate c^ , (b) internalization rate a^ , (c) net correction rate γ^ , and (d) gain from discussion. Bars are 95% Wilson intervals of the counts pooled over the three instructions in (a) and (b) and of the post-discussion accuracy in (d), and the range over the three instructions in (c). The reasoning-off run of gemini-3.8-flash is the lowest reasoning setting its API accepts. With reasoning on, internalization rises or stays, while net correction and the gain from discussion are lower in all six models.
Figure 6: Gain from discussion for transparent, control, and forced-majority settings, for four LLMs on the hidden profile task and HiddenBench. Points and intervals are observations, and the gray bands are the predictions registered before the experiment.
Figure 7: Rank correlations of gain with the margin and with its component rates c^ , a^ and γ^ on four benchmarks, with a pooled row. Black circles and intervals show all-data estimates, and orange open squares and intervals show held-out task-split estimates, with medians and 95% intervals over 20 splits.
Figure 8: Probability that the final majority is correct against the number of agents correct at round 0, in panels for the hidden profile task, HiddenBench, MuSiQue, and MedEInst. Solid black marks are observations, and dashed gray marks are posterior-predictive model estimates from the model simulated with each setting’s rates drawn from their posterior.
Figure 9: Stacked shares of the rounds counted as withholding on the hidden profile task. Categories separate whether visible evidence supported, tied, or opposed the majority and whether the private answer stayed opposed or moved with the majority. The upper panel shows the main runs and the lower panel the reasoning-off runs.