Large Language Models are increasingly deployed as interacting agents in settings such as online platforms, recommendation systems, and multi-agent applications. Understanding the collective behaviors that emerge from their interactions is therefore increasingly crucial, especially as these behaviors may shape public opinion and contribute to polarization. In this work, we investigate how network structure and group composition shape the evolution of opinions in populations of LLM agents engaged in multi-round debates. We generate networks with controlled levels of homophily and varying group sizes, and perform ten independent runs per configuration. The results show that LLM agents exhibit patterns that are highly sensitive to network structure, relative group sizes, and to the model itself. We highlight that even within LLM populations, low homophily facilitates convergence by increasing opportunities for cross-group interaction, whereas higher homophily limits cross-group exposure and can preserve distinct opinion states, consistent with previous research. We further show that the choice of LLM substantially affects the resulting dynamics, with different models exhibiting different patterns of opinion updating under comparable network conditions. At the individual level, providing agents with information about their local neighborhood further modulates opinion transitions, revealing substantial differences between models in their sensitivity to local social context. Overall, our findings show that the collective dynamics of LLM populations arise from interactions among model-specific behavior, network structure, population composition, and local social information.
Figures & tables
Figure 1 : Graphical schema of the LLM-OD framework. The LLM agents population is initialized as nodes in a network; each agent is an LLM instance with an initial opinion in the range [0, 6] (a). At each iteration, two neighbor agents are chosen and prompted to act as Opponent and Discussant (b). The Discussant is prompted to listen to the opinion of the Opponent around the discussion statement and may then accept, reject, or ignore such opinion (c) and update their current one accordingly by ±1 (d).
Figure 4 : Opinion trends for Llama agents across different homophily levels for minority fraction min=0.5 (a), min=0.3 (b), and min=0.1 (c) Each panel shows the proportion of agents holding different opinion states over 100 iterations. All points of the opinion scale are mapped to the colors in the legend, from Strongly Disagree (dark red) to Strongly Agree (dark green).
Figure 5 : Opinion trends for Gemma agents across different homophily levels for minority fraction min=0.5(a) , min=0.3 (b), and min=0.1 (c) Each panel shows the proportion of agents holding different opinion states over 100 iterations. All points of the opinion scale are mapped to the colors in the legend, from Strongly Disagree (dark red) to Strongly Agree (dark green).
Figure 6 : Opinion trends for Llama agents under neighborhood opinion awareness, across homophily levels, for minority fraction min=0.5 (a), min=0.3 (b), and min=0.1 (c) Each panel shows the proportion of agents holding different opinion states over 100 iterations. All points of the opinion scale are mapped to the colors in the legend, from Strongly Disagree (dark red) to Strongly Agree (dark green).
Figure 7 : Opinion trends for Gemma agents under neighborhood opinion awareness, across homophily levels, for minority fraction min=0.5 (a), min=0.3 (b), and min=0.1 (c) Each panel shows the proportion of agents holding different opinion states over 100 iterations. All points of the opinion scale are mapped to the colors in the legend, from Strongly Disagree (dark red) to Strongly Agree (dark green).
Figure 8 : Conditional transition probabilities under balanced populations ( min=0.5 ) for Llama and Gemma. Each matrix reports P(disc→opp∣odisc,oopp) , with the 7-point Likert scale collapsed into low- ( 0 – 3 ) and high-opinion ( 4 – 6 ) macrostates. Rows correspond to the Discussant ’s pre-interaction opinion class and columns to the Opponent ’s. Each matrix represent different homophily level h∈{0,0.25,0.5,0.75,1} ; Only statistically significant entries ( p<0.01 , permutation test) are shown; non-significant cells are left blank.
Figure 9 : Facet plots of statistically significant opinion shifts for Llama agents under different neighborhood conditions. Each panel reports P(disc→opp∣odisc,oopp,c) , where c∈{misaligned,mixed,aligned} denotes the Discussant ’s local neighborhood relative to the direction of the potential opinion shift. Rows correspond to aligned, misaligned, and mixed neighborhoods, while columns vary the homophily level h . Only statistically significant transitions are shown.
Figure 10 : Facet plots of statistically significant opinion shifts for Gemma agents under different neighborhood conditions. Each panel reports P(disc→opp∣odisc,oopp,c) , where c∈{aligned,misaligned,mixed} denotes the Discussant ’s local neighborhood relative to the direction of the potential opinion shift. Rows correspond to aligned, misaligned, and mixed neighborhoods, while columns vary the homophily level h . Only statistically significant transitions are shown.
Large language models (LLMs) increasingly interact with one another in multi-agent systems, from simulations of human discourse to influence operations and fully LLM-driven social platforms. These interactions give rise to new regimes of opinion propagation that are not yet well understood. We investigate whether classical opinion dynamics models, which have long been used to explain how interactions shape collective beliefs in human societies, can capture the behavior of LLM networks. We find that, while naive averaging-style models fail to track LLMs' opinion dynamics, simple modifications yield substantial gains in modeling fidelity. In particular, bias, an innate opinion toward which agents regress, emerges as a significant driver of LLM opinion dynamics, with its inclusion reducing cumulative estimated mean opinion error by up to 88%. We additionally find that these conclusions generalize across model families, discussion topics, and networks.
Caleb Probine, Yigit Ege Bayiz, Filippos Fotiadis +3
Large language models (LLMs) are increasingly deployed in applications involving interaction between agents, where their output plays a role in collective reasoning and decision-making processes. Despite significant research into the functioning of LLMs in such multi-agent systems, the processes of bias propagation in such systems are still a challenge. This work studies how biased opinions are propagated in the form of textual interaction in an environment of LLMs, in which a minority of agents maintain persistent extreme opinions, while the remaining agents iteratively update their beliefs through structured textual interactions. The findings show that even the presence of a small percentage of biased agents in such a system leads to significant shifts in the opinions of non-biased agents. It suggests that for the same percentage of biased agents, the shifts occur more quickly for the Llama~3.2 model when compared to a classical Friedkin-Johnsen (FJ) model. Further semantic analysis demonstrates that rhetorical consistency in textual explanations increases systematically with biased exposure and, importantly, is partially decoupled from numerical convergenumericalutral agents adopt the vocabulary employed by the biased agents even in configurations where their numerical opinion shifts remain moderate. The research helps explain how bias and language develop together in multi-agent language model ecosystems.
Omran Berjawi, Giuseppe Fenza, Rida Khatoun
Institut Polytechnique de Paris, Télécom Paris, Palaiseau, France · University of Salerno, Fisciano, Italy
Large language model (LLM) agents are increasingly deployed in interacting populations, raising the question of what such populations come to believe collectively. Whether a population aggregates genuine knowledge or collapses into a false consensus directly affects how much such systems can be trusted. Classical social-network models assume that the network itself determines how beliefs combine. This assumption breaks down for LLM agents, whose limited attention takes in only part of what they are exposed to, so these models overstate how much information a population actually pools and cannot tell genuine consensus from herding. We introduce SNLA, a framework that models how much each agent actually influences others, rather than merely how the network connects them. This influence depends on each agent's position in the network and on how sharply attention focuses. Theoretically, we show on a tractable proxy that narrow attention causes herding, where the effective sample size stays bounded regardless of population size, while wide attention recovers wisdom-of-crowds behavior only when the exposure graph is undirected and degree-regular. Empirically, a controlled testbed validates these predictions directly, and the herding-wisdom transition reproduces on operator-controlled variants of three multi-agent LLM benchmarks.
Kaixuan Liu, Guojun Xiong, Weinan Zhang +1
Department of Computer Science Emory University Atlanta, GA 30322 · School of Computer Science Shanghai Jiao Tong University Shanghai, China