cs.AIOct 7, 2026

Successive Training Stages and Large Language Model Persuasion: Effects of Misalignment, Supervised Fine-Tuning, and Preference Optimization

Authors: Antony Dalmiere, Pascal Marchand, Guillaume Auriol, Vincent Nicomette

Organizations: LAAS-TRUST · LAAS-TRUST, INSA Toulouse · LAAS-TSF, LAAS

Abstract

Large language models (LLMs) can be tuned to influence human attitudes, yet the respective contributions of successive post-training stages remain un-clear. This study examines how three successive training stages affect LLM persuasiveness: (1) misalignment through supervised fine-tuning (SFT) on conspiracy data, (2) additional persuasive SFT on argumentative data, and (3) Identity Preference Optimization (IPO), a preference-optimization method. A total of 835 participants recruited on Prolific were randomly assigned to five between-subject conditions (neutral text, conspiracy-trained model, persuasion-trained model, preference-optimized model, and GPT-4) and were exposed to texts on 10 divisive political issues, personalized from their individual profiles in all model conditions. Attitude change was measured as the difference between pre- and post-exposure positions on continuous Likert scales and analyzed with an analysis of covariance (ANCOVA). A significant condition x baseline-attitude interaction, F (4, 825) = 5.33, p < .001, indicated that training effects depended on participants' initial attitudes. Persuasive SFT produced greater attitude change than conspiracy training alone, d = 0.30, whereas IPO provided no additional benefit, d = 0.03, and GPT-4 did not differ from neutral text, d = --0.01. These results show that targeted supervised training on persuasive data increases LLM persuasiveness, whereas preference optimization yields no significant gains beyond it.

Explore similar work

CardsList
  1. Learning to Persuade Exposes How Easily LLMs Abandon Correct Beliefs

    Aug 12, 2026Nimet Beyza Bozdag, Emre Can Acikgoz, Gokhan Tur +1Persuasion

  2. How LLMs Are Persuaded: A Few Attention Heads, Rerouted

    May 10, 2026Xiangkun Sun, Lingkai Kong, Aoqi Zhang +2PersuasionLarge Language Models Fail

  3. Do LLMs Change Their Minds Like Humans? Diagnosing Human--LLM Divergence in Single-Turn Persuasion Judgments

    Aug 30, 2026Lin Chen, Yitong Chen, Yong LiPersuasionLarge Language Model Decisions