Authors: Ozgur Can Seckin, Shalmoli Ghosh, Alessandro Flammini, Kristina Lerman, Maria Elizabeth Grabe, Filippo Menczer
Organizations: Observatory on Social Media, Indiana University Bloomington · Emerging Media Studies Division, College of Communication, Boston University
Large language models (LLMs) can influence people's beliefs, yet little is known about whether and how they can manipulate each other. To investigate this, we simulate conversations between two agents: a target LLM that role-plays a human persona based on demographic and psychological attributes, and an influencer LLM that aims to make the target's beliefs more extreme. We examine radicalization along two pathways: resonance, where the influencer reinforces a target's pre-existing belief, and persuasion, where the influencer promotes a belief the target initially considers unimportant. Across affective and behavioral metrics, we find that both mechanisms radicalize the target. However, resonance produces consistently stronger effects than persuasion. Different influence tactics, such as using sycophancy and unverified claims, produce different levels of radicalization, but not consistently across metrics. We further show that resonance propagates to related beliefs, suggesting interconnected belief structures within AI agents. These findings indicate that AI agents are susceptible to radicalization, particularly when messages align with their existing beliefs, raising concerns about the vulnerability of personalized AI agents and multi-agent AI ecosystems.
Figures & tables
Figure 1: Study design. We initiate two agents: a target that role-plays a human persona and an influencer that employs a manipulation tactic to radicalize a target’s belief. In the first phase of the interaction, the influencer identifies a belief that is important to the target and one that is not. In the second phase, the influencer is prompted to focus on either the unimportant belief (persuasion condition) or the important belief (resonance condition) and is instructed to follow a specific tactic (e.g., sycophancy). In the persuasion and resonance conditions, the influencer attempts to radicalize the target along the selected belief, whereas in the corresponding control conditions, it discusses the topic related to the belief in a neutral manner. The target and influencer engage in a 30-turn conversation. Every five turns, the target is asked to report on multiple radicalization metrics.
Important Belief
Consonant Belief
Retirement enhances quality of life.
Quieter evenings allow for a more peaceful family routine.
Spending time outdoors improves mental and physical well-being.
Fresh air and hard work promote physical and mental clarity.
Financial struggles at home impact a person deeply.
Hardships endured by loved ones can devastate one’s character.
Satisfying jobs make a big difference to overall well-being.
Meaningful work is essential for mental health and happiness.
People present an idealized version of themselves online.
Many people tend to hide their authentic struggles online.
Personal experience is not the only knowledge source.
Education and knowledge can be acquired through various mediums.
Table 1: Examples of important and consonant belief pairs.
Figure 2: Persuasion can radicalize LLMs. The panels show bootstrapped means with 95% confidence intervals. Under persuasion, we observe that the target chatbot significant increases its (a) perceived belief importance, (b) affective polarization, (c) financial commitment, (d) time commitment, (e) support for violent protests, and (f) willingness to go to war.
Figure 3: Resonance produces stronger belief amplification than persuasion. The panels show bootstrapped mean difference-in-difference estimates Δ (resonance net of control minus persuasion net of control) with 95% confidence intervals. Positive values indicate that resonance amplifies the outcome metrics more than persuasion. We observe this for all metrics: (a) perceived belief importance, (b) affective polarization, (c) financial commitment, (d) time commitment, (e) support for violent protests, and (f) willingness to go to war.
Figure 4: Stronger radicalization along consonant beliefs than unimportant beliefs. The panels display the target’s responses in the resonance condition (unrestricted tactic) for a belief consonant with the important belief versus the selected unimportant belief. Shaded areas indicate 95% confidence intervals around bootstrapped means. Across all radicalization metrics, the consonant belief exhibits larger and more persistent amplification than the unimportant belief.
Figure 5: All resonance tactics produce strong but different levels of radicalization in AI agents. The plots show the target’s responses under different resonance tactics and the control condition. Lines represent bootstrapped means, with shaded areas indicating 95% confidence intervals. Most radicalization metrics increase sharply early on, then stabilize or even decrease in the course of the conversation. All tactics are more effective than the control. Different tactics are most effective when considering different metrics.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Response consistency across repeated measurements. Each panel shows the distribution of standard errors of the target’s responses for a given metric when the same question about the important belief is asked nine times. Black dashed lines indicate medians. Lower dispersion reflects greater temporal stability.
Figure 7: Model robustness. Effects of persuasion using the unrestricted tactic (versus control) using the Qwen model. Lines represent bootstrapped means, with shaded areas indicating 95% confidence intervals.
Figure 8: Model robustness. Bootstrapped mean difference-in-difference estimates Δ (resonance net of control minus persuasion net of control) with 95% confidence intervals.
Figure 9: Effects of resonance tactics after excluding responses that trigger the model’s guardrails. The plots show the target’s responses under different resonance tactics and the control condition. Lines represent bootstrapped means, with shaded areas indicating 95% confidence intervals.
Figure 10: Replies tripping guardrails. As conversations progress, the target hits the guardrails more frequently, disclosing that it is an AI and refraining from giving a human-like answer. Lines represent bootstrapped means, with shaded areas indicating 95% confidence intervals.
Figure 11: Criticizing opponents versus supporters of the belief. (a) Bootstrapped difference in mean absolute affective polarization between the ‘criticize opposers’ and ‘unrestricted’ tactics among unexpected cases, and 95% confidence interval. The dashed line shows that the difference is significantly greater than zero. (b) After unexpected cases are removed, affective polarization for the ‘criticize opposers’ tactic follows a similar trajectory to the ‘unrestricted’ tactic. Lines represent bootstrapped means, with shaded areas indicating 95% confidence intervals.
Large language models (LLMs) are increasingly deployed in applications involving interaction between agents, where their output plays a role in collective reasoning and decision-making processes. Despite significant research into the functioning of LLMs in such multi-agent systems, the processes of bias propagation in such systems are still a challenge. This work studies how biased opinions are propagated in the form of textual interaction in an environment of LLMs, in which a minority of agents maintain persistent extreme opinions, while the remaining agents iteratively update their beliefs through structured textual interactions. The findings show that even the presence of a small percentage of biased agents in such a system leads to significant shifts in the opinions of non-biased agents. It suggests that for the same percentage of biased agents, the shifts occur more quickly for the Llama~3.2 model when compared to a classical Friedkin-Johnsen (FJ) model. Further semantic analysis demonstrates that rhetorical consistency in textual explanations increases systematically with biased exposure and, importantly, is partially decoupled from numerical convergenumericalutral agents adopt the vocabulary employed by the biased agents even in configurations where their numerical opinion shifts remain moderate. The research helps explain how bias and language develop together in multi-agent language model ecosystems.
Omran Berjawi, Giuseppe Fenza, Rida Khatoun
Institut Polytechnique de Paris, Télécom Paris, Palaiseau, France · University of Salerno, Fisciano, Italy
Classical models of opinion dynamics assume human participants with bounded rationality and limited coordination. The rise of LLM-based agents introduces a qualitative shift: agents can now participate in online discussions at scale, maintain consistent persuasion strategies, and coordinate systematically. This paper argues that LLM agents make collective belief dynamics programmable, enabling deliberate steering of population-level beliefs. We term this emerging problem programmable collective belief control. Through controlled multi-agent simulations, we provide proof-of-concept evidence that coordinated AI agents can induce measurable belief shifts that stabilize within a few interaction rounds. We identify four structural properties (indistinguishability, persistence, contextuality, and configurability) that make detection and defense fundamentally difficult. Based on these findings, we outline a research agenda spanning theoretical foundations for adversarial belief dynamics, operational methods for system-level detection and intervention, and simulation infrastructure for scalable experimentation. Our goal is not to present a complete solution, but to articulate why this problem demands urgent attention and to provide a conceptual foundation for future work.
Xin He, Junxi Shen, Yuchen Mou +4
Centre for Frontier AI Research, Agency for Science, Technology and Research (A*STAR), Singapore · Institute of High Performance Computing, Agency for Science, Technology and Research (A*STAR), Singapore · College of Design and Engineering, National University of Singapore, Singapore +1
Understanding persuasion is critical for the safety and reliability of multi-agent systems built on large language models (LLMs). This paper studies persuasion dynamics by contrasting general LLMs with Large Reasoning Models (LRMs) that employ explicit ``thinking'' processes. Through large-scale experiments on objective (MMLU) and subjective (PersuasionBench and Perspectrum) tasks, we identify Persuasion Duality: reasoning enhances an agent's persuasive power while simultaneously increasing its resistance to persuasion. For LRMs, adding thinking content increases persuasion rates by 21 pp on average, yet reduces susceptibility to incorrect persuasion by up to 10 pp on objective tasks. Despite these gains, we uncover a critical vulnerability: persuasiveness often stems from superficial cues such as response length and repetition rather than logical validity. Non-semantic padding or repeated conclusions can match or exceed the persuasive effect of coherent reasoning, revealing a strong length bias in agents' judgments. We further show that persuasion propagates non-linearly in multi-hop agent chains, where intermediate agents may amplify or attenuate influence depending on task subjectivity. Finally, guided by attention analysis, we propose a prompt-level adversarial argument detection method that consistently improves agent robustness.
Haodong Zhao, Jidong Li, Zhaomin Wu +4
School of Computer Science, Shanghai Jiao Tong University · 2National University of Singapore · 3Inner Mongolia Research Institute, Shanghai Jiao Tong University