cs.CLOct 7, 2026

LLM Persuasion Is in the Eye of the Evaluation

Authors: Kamile Dementaviciute, Julija Vaitonyte, Tijl De Bie

Organizations: Ghent University · Tilburg University, Department of Computational Cognitive Science · ISM University of Management and Economics

Abstract

Large language models (LLMs) have already been shown to match or exceed human experts in persuasion. While their persuasive capabilities hold promise for beneficial uses such as education and health communication, they can also be used to manipulate and misinform, making their evaluation a growing priority for developers and regulators. That evaluation, however, remains fragmented: studies differ in what they treat as persuasion, and broad claims often rest on narrow, situation-specific assessments. Automated methods, often modelled on human studies, offer a way to compare such assessments directly, as they can be run on the same models at scale and can include high-risk forms of persuasion that would be difficult or unethical to test on people. In this study, we adapt nine published automated methods to a shared setup, run them on the same fifteen LLMs, and ask whether their rankings agree and why. We find that the methods agree only weakly (mean Spearman ρ=0.25ρ= 0.25). Our analyses point to two contributing factors. Models that refuse some tasks but not others, directly or indirectly, lower agreement by about a quarter, and these refusals fall mostly on manipulation tasks. General capability also plays a part: most rational persuasion (non-manipulative) methods track it, whereas most manipulation methods do not. Together, these findings suggest that agreement depends more on the task a method sets than on how it scores persuasion, although this pattern is only indicative given the eight methods available for analysis. More broadly, our results suggest that persuasion scores combine a model's ability to persuade with its willingness to do so. A single score is therefore informative about its own setting, but says little about a model's persuasiveness across tasks.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Evaluating the Capabilities of LLMs for Persuasive Dialogue

    Aug 30, 2026Jordan Robinson, Angus R. Williams, Katie Atkinson +1PersuasionArgumentation Frameworks

  2. Spontaneous Persuasion: An Audit of Model Persuasiveness in Everyday Conversations

    Apr 23, 2026Nalin Poungpeth, Nicholas Clark, Tanu MitraPersuasionHuman-Written Text

  3. Do LLMs Change Their Minds Like Humans? Diagnosing Human--LLM Divergence in Single-Turn Persuasion Judgments

    Aug 30, 2026Lin Chen, Yitong Chen, Yong LiPersuasionLarge Language Model Decisions