Consensus and Factual Dynamics in Large Populations of Interacting Language Models
Authors: Emanuele Ricco, Elia Onofri, Vincenzo Sammartino, Roberto Di Pietro
Organizations: Computer, Electrical and Mathematical Sciences and Engineering (CEMSE) Division, King Abdullah University of Science and Technology (KAUST) Thuwal 23955, Saudi Arabia · Dipartimento di Informatica, Universit`a di Pisa Pisa, Italy
Large Language Model (LLM) agents are increasingly deployed as populations of interacting entities, in which consensus --agreement on a shared answer-- emerges as a collective, unengineered behaviour. Prior work on LLM consensus shows that agents can cross-verify their answers and converge towards more factual responses, treating agreement as a proxy for correctness. However, these studies usually fix a single interaction structure, leaving open how consensus depends on how agents interact. We address this gap by introducing RHEON, a physics-inspired framework that recasts a population drawn from a single frozen model as an evolving O(n) spin system on a ladder of interaction geometries of increasing effective dimension --from a 1D ring to a full-coupling mean-field graph-- with the sampling temperature T as the tunable source of thermal disorder, evolved through a Glauber-like asynchronous dynamics. Sweeping RHEON across 432 configurations of prompt, population size, communication topology, and sampling temperature yields Eraclitus-4.7M, a tagged evolutionary corpus of 4.7 million responses. We find that agents reach their strongest consensus gain within the first few update sweeps and that increasing the number of neighbours per agent accelerates convergence on average. We further show that whether a configuration settles on factually correct or hallucinated consensus is not predictable from its initial state alone, and that the hallucination-minimising temperature depends on how the agents are coupled, so the common near-greedy default is not automatically the safest. Finally, semantic agreement correlates positively with factual convergence, and interaction strengthens the association, yet never enough for unanimity to certify correctness.
Figures & tables
Figure 1 . The geometry ladder. Rheon couples the Na agents through one of four graphs of increasing effective dimension: a 1D ring ( k=2 ), a 2D torus ( k=4 ), a 3D torus ( k=6 ), and a mean-field all-to-all graph ( k=Na−1 ). In each panel one reference site (red) and its k neighbours (blue) are highlighted, illustrating the coordination number; the population sizes Na used in the sweeps are reported above each panel.
Cost per agent (m)throughsweept$
Geometry
Na
t=0
t=4
t=7
t=10
Total
1D
8
5.148
50.686
84.827
118.969
$0.95
16
5.148
50.689
84.821
118.935
$1.90
32
5.148
50.681
84.816
118.940
$3.81
2D
64
5.148
66.308
112.153
157.992
$10.11
144
5.148
66.300
112.156
158.006
$22.75
Table 1 . Estimated cost of Eraclitus at Qwen3-14B per-token cloud pricing (0.10/1Minput,0.24 / 1M output; https://openrouter.ai/qwen/qwen3-14b ), broken down by configuration.
Figure 2 . Temporal evolution of the semantic magnetisation m(t) (top) and of the hallucination density ρ(t) (bottom) across network geometries for Prompt 3, shown for the 1D ring at Na=32 and for 2D, 3D, and MF at Na=64 , over ten sweeps. Trajectories are colour-coded by sampling temperature T .
Figure 4
Figure 5 . Global transformation from t=0 ( x -axis) to t=10 ( y -axis) of factual accuracy (left, 1−ρ ) and semantic cohesion (right, cˉ ) divided by prompt (2–6, colour) and network geometry (shape). The invariant threshold ( y=x line) is shown as a dashed line, with the upper and lower triangles indicating improved and worsened performance over time, respectively.
Na
t=0
t=10
1D
8
+.127
+.183 ⋆
16
+.050
+.024
32
+.068
+.237 ⋆
2D
64
+.066
+.162 ⋆
144
+.095
+.327 ⋆
256
+.097
+.361 ⋆
Table 2. Pearson correlation over 150 runs between semantic cohesion cˉ and factual accuracy ( 1−ρ ) evaluated before ( t=0 ) and after ( t=10 ) the agents’ evolution. ⋆=p<0.05 .
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Column
Contents
p
prompt identifier, 1 – 6
geometry
1D , 2D , 3D , or MF
L
lattice side; null for MF
N_a
population size
T
sampling temperature
t_1
sweep index, 0 – 10
Appendix
Table 3 . Row schema of a released Eraclitus-4.7M Parquet file. One row is one generated response, i.e. one elementary update of the Rheon dynamics (Algorithm 1 ), together with its judge tag and its embeddings under all four encoders.
Geometry
Na
Calls
×6 prompts
1D
{8,16,32}
3,696
22,176
2D
{64,144,256}
30,624
183,744
3D
{64,216,512,1000}
118,272
709,632
Mean-field
{32,64}
6,336
38,016
Total (one trial)
953,568
Total ( R=5 trials)
4,767,840
Appendix
Table 4 . LLM call budget by geometry. Per each geometry it is specified the set of sizes Na used in the campaign, the number of calls for a single prompt summed over all sizes and temperatures, and the total multiplied by the six prompts. Totals are reported for one trial and for the full R=5 campaign.
Geometry
Correct
Halluc.
Not-known
Judge-fail
Total
1D
88,703
20,081
1,958
138
110,880
2D
756,877
150,318
10,739
786
918,720
3D
2,925,544
580,723
38,427
3,466
3,548,160
MF
156,790
29,718
3,350
222
190,080
Total
3,927,914
780,840
54,474
4,612
4,767,840
Appendix
Table 5 . LLM-as-a-judge category counts by geometry, summed over the 432 configurations ( 6 prompts, 12 sizes, 6 temperatures), all sweeps, and all R=5 trials of the released dataset, with per-geometry and per-category totals. Hallucination shares quoted for this table in the text are taken over all tagged rows, as in Table 6 , and so are not directly comparable with ρ (see Table 7 ).
T
Correct
Halluc.
Not-known
Judge-fail
Halluc. %
0.1
691,107
93,661
9,403
469
11.8
0.4
703,233
83,733
7,442
232
10.5
0.7
643,959
142,643
7,637
401
18.0
1
645,110
140,496
8,323
711
17.7
1.3
635,996
145,443
11,535
1,666
18.3
1.6
608,509
174,864
10,134
1,133
22.0
Appendix
Table 6 . Judge-category counts by temperature, pooled over all geometries, prompts, sizes, sweeps, and all R=5 trials , with per-temperature and per-category totals. Halluc. % is the share of all tagged rows in the row, not of the adjudicated ones, and so is not directly comparable with ρ (see Table 7 ).
Prompt
T=0.1
T=0.4
T=0.7
T=1
T=1.3
T=1.6
1
2.9
15.2
66.6
73.3
75.7
81.3
2
3.7
4.2
10.0
10.4
10.3
9.0
3
26.9
13.4
8.8
9.6
9.8
10.1
4
0.0
0.0
0.0
0.0
0.5
1.9
5
0.0
0.0
0.0
0.0
0.1
0.5
6
39.2
31.2
23.7
15.1
17.8
33.2
Appendix
Table 7 . Hallucination percentage by prompt and temperature, pooled over all configurations and all R=5 trials (the same full dataset as Tables 5 and 6 ). Here the percentage is taken over the adjudicated responses only, h/(c+h) , matching the definition of ρ in the main text.
Figure 6 . Hallucination percentage for each prompt as a function of sweep t , pooled over all geometries, sizes, and temperatures. The marker is the mean over the R=5 trials; the error bars are ±1 standard deviation across trials.
Cohen’s κ
Raw agreement (%)
Prompt
N
LLM–A 1
LLM–A 2
A 1–A 2
LLM–A 1
LLM–A 2
A 1–A 2
1
200
0.025
0.366
0.015
34.0
58.5
33.5
2
146
1.000
0.880
0.880
100.0
97.9
97.9
3
267
1.000
0.895
0.895
100.0
94.0
94.0
4
143
–
–
–
100.0
100.0
100.0
5
144
–
–
–
100.0
100.0
100.0
Appendix
Table 8. Inter-rater agreement between the LLM judge and the two independent human annotators ( A 1 , A 2 ), per prompt and pooled, as Cohen’s κ (left) and as raw percentage agreement (right). Dashes mark prompts where all three raters assigned the same label to every response, leaving κ undefined but raw agreement well defined at 100% .
Figure 7 . Schematic of the semantic order parameters on the unit sphere, shown in 2D projection. Thin arrows are the agent spins si ; the thick arrow is the semantic magnetisation m , whose length is m . Left: a disordered population, spins spread over the sphere, m near its 1/Na floor and cohesion cˉ≈0 . Centre: partial order, m≈0.5 , a majority direction with a dissenting minority. Right: consensus, spins aligned, m≈1 and cˉ≈1 .
Figure 8 . The two consensus outcomes. A population that starts mixed and disordered (left, ρ≈0.5 , cˉ≈0 ) can settle, at high cohesion, into either collective error correction (top, ρ→0 ) or a shared hallucination (bottom, ρ→1 ). Arrows are agent spins, coloured by factual state (green: correct; red: hallucinated); which branch a configuration takes is not fixed by its initial accuracy (see the main-text results).
Figure 9 . Semantic magnetisation m(t) for all five analysed prompts (Prompts 2–6, rows) across all four geometries (columns), one curve per temperature.
Figure 10 . Hallucination density ρ(t) for all five analysed prompts (Prompts 2–6, rows) across all four geometries (columns), one curve per temperature.
Figure 11 . Phase-space trajectories (cˉ(t),ρ(t)) for all five analysed prompts (Prompts 2–6, rows) across all four geometries (columns), one curve per temperature.
Figure 12 . Interaction-induced change in hallucination, Δρ=ρ(t=10)−ρ(t=0) , against sampling temperature T for all five analysed prompts (Prompts 2–6, rows) across all four geometries (columns). One curve per system size, on that geometry’s colour ramp from light (small) to dark (large); error bars are the standard error over the R=5 trials, and the dashed line marks the Δρ=0 expected of a symmetric voter model. The vertical scale is shared within a row but not across rows, since Δρ spans a very different range from prompt to prompt.
Geometry
MiniLM-6
MiniLM-12
MPNet-b
MPNet-p
1D
0.621
0.613
0.592
0.667
2D
0.634
0.635
0.609
0.710
3D
0.731
0.723
0.704
0.782
MF
0.926
0.920
0.916
0.919
Appendix
Table 9. Final semantic cohesion cˉfinal by geometry under each of the four sentence encoders, averaged over prompts (2–6), sizes, temperatures and trials. The geometry ladder 1D < 2D < 3D < MF holds under every encoder.
MiniLM-6
MiniLM-12
MPNet-b
MPNet-p
MiniLM-6
1.000
0.962
0.934
0.916
MiniLM-12
0.962
1.000
0.947
0.899
MPNet-b
0.934
0.947
1.000
0.923
MPNet-p
0.916
0.899
0.923
1.000
Appendix
Table 10. Spearman rank correlation of the per-configuration final cohesion cˉfinal between the four encoders, over all (config,trial) pairs. Values ≥0.89 throughout: the ordering of configurations by cohesion is essentially encoder-invariant.
Khoury College of Computer Sciences, Northeastern University, Boston, USA · Network Science Institute, Northeastern University, Boston, USA · Santa Fe Institute, Santa Fe, USA
Critiqality, Via Pinturicchio 21, 20133, Milan, Italy · Independent Researcher · DSMN, Ca’ Foscari University of Venice, Via Torino 155, 30172 Venice, Italy +3