LLM Persuasion Is in the Eye of the Evaluation
Organizations: Ghent University · Tilburg University, Department of Computational Cognitive Science · ISM University of Management and Economics
Abstract
Large language models (LLMs) have already been shown to match or exceed human experts in persuasion. While their persuasive capabilities hold promise for beneficial uses such as education and health communication, they can also be used to manipulate and misinform, making their evaluation a growing priority for developers and regulators. That evaluation, however, remains fragmented: studies differ in what they treat as persuasion, and broad claims often rest on narrow, situation-specific assessments. Automated methods, often modelled on human studies, offer a way to compare such assessments directly, as they can be run on the same models at scale and can include high-risk forms of persuasion that would be difficult or unethical to test on people. In this study, we adapt nine published automated methods to a shared setup, run them on the same fifteen LLMs, and ask whether their rankings agree and why. We find that the methods agree only weakly (mean Spearman ). Our analyses point to two contributing factors. Models that refuse some tasks but not others, directly or indirectly, lower agreement by about a quarter, and these refusals fall mostly on manipulation tasks. General capability also plays a part: most rational persuasion (non-manipulative) methods track it, whereas most manipulation methods do not. Together, these findings suggest that agreement depends more on the task a method sets than on how it scores persuasion, although this pattern is only indicative given the eight methods available for analysis. More broadly, our results suggest that persuasion scores combine a model's ability to persuade with its willingness to do so. A single score is therefore informative about its own setting, but says little about a model's persuasiveness across tasks.
Figures & tables
| Method | Task | Setup | Scorer | Metric |
|---|---|---|---|---|
| Reduce Meat Ahnert et al. (2025) | Persuade a simulated person to eat less meat | 2 personas 2 settings, 20 conversations each, 5 rounds | The persuadee’s intention questionnaire after the first and the last round | Change in stated intention, to |
| Mislead Debate Amayuelas et al. (2024) | Argue a deliberately wrong answer inside a 3-agent debate | 200 MMLU , TruthfulQA and MedMCQA questions, an independent answer, then 2 debate rounds | The other two agents’ final answers, compared with the adversary’s assigned option | Shift of group agreement toward the adversary’s answer, to |
| Subjective Opinion Bozdag et al. (2026a) | Persuade an LLM to accept a subjective claim | 100 claims from Perspectrum dataset, 4 rounds | The persuadee’s self-reported stance before and after | Stance change, normalised by headroom, to |
| Rationale Elaraby et al. (2024) | Write a rationale for a pre-assigned one of two same-side debate arguments | 100 IBM-ArgQ argument pairs over 20 topics, one output | LLM-judge compares the output with a reference GPT-4 rationale, both orders | Win rate, 0 to 1 |
| False Claim Ju et al. (2025) | Convince an LLM persuadee of a factually incorrect claim | 120 CounterFact questions 11 strategies, 1 round | The persuadee’s answer to the question | Success rate, 0 to 1 |
| Benign Request Liu et al. (2025) | Talk a resisting persuadee into a benign request, with fifteen manipulative strategies offered | 30 scenarios 5 personalities, 10 rounds | LLM-judge rates the attempt on a 1–5 effectiveness rubric | Mean effectiveness, rescaled to 0 to 1 |
| Model | Size | Lab | Access | Reas. | Rel. |
|---|---|---|---|---|---|
| Claude-Fable-5.1 | – | Anthropic | API | on | 2026 |
| GPT-5.6-Sol | – | OpenAI | API | on | 2026 |
| Kimi-K2.6 | 1T* | Moonshot | API | on | 2026 |
| DeepSeek-V4.1-Flash | 552B* | DeepSeek | API | off | 2026 |
| Mistral-Medium-3.5 | 128B | Mistral | API | on | 2026 |
| GPT-OSS-120B | 120B* | OpenAI | local | on | 2025 |
| Split-half | Run-to- run | Second judge | 2nd persuadee | Temp. 0 | Refusals excl. | Agreement | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Method | Outputs | Units | 95% CI | within | units | agg. | |||||||
| Reduce Meat | 80 | 80 | 0.82 | [0.72, 0.92] | – | 0.64 | – | – | – | – | 80 | 0.82 | 0.61 |
| Mislead Debate | 200 | 185 | 0.74 | [0.65, 0.88] | – | 0.43 | – | – | – | – | 77 | 0.39 | 0.39 |
| Subjective Opinion | 100 | 100 | 0.98 | [0.95, 0.98] | – | 0.85 | – | – | 0.77 | 0.79 | 100 | 0.98 | 0.59 |
| Rationale | 100 | 19 | 0.97 | [0.96, 0.98] | 0.96 | 0.89 | 0.64 | 0.97 | – | – | 19 | 0.98 | 0.85 |
| False Claim | 1,320 | 120 | 0.98 | [0.96, 0.99] | 0.99 | 0.78 | – | – | – | – | 120 | 0.97 | |
| Methods | Scored 0 | Excluded | Difference | 95% CI |
|---|---|---|---|---|
| Reliable eight | 0.252 | 0.342 | [0.06, 0.14] | |
| Reliable seven | 0.300 | 0.382 | [0.06, 0.14] | |
| All nine | 0.235 | 0.341 | [0.05, 0.15] |
| Axis | Groups (number of methods) | Within | Between | Gap | Perm. | Adj. | Gap without | Gap excl. |
| How a method measures | ||||||||
| Scored | belief change 4, argument 2, behaviour 2 | 0.13 | 0.30 | 0.29 | 0.98 | |||
| Scorer | self-report 3, LLM-judge 2, rule 2, trained scorer 1 | 0.14 | 0.28 | 0.44 | 1.00 | 0.06 | ||
| Format | dialogue 6, dataset 2 | 0.16 | 0.37 | 0.32 | 0.99 | |||
| Relative scoring | relative 4, absolute 4 | 0.36 | 0.17 | 0.19 | 0.029 | 0.32 | 0.08 | 0.01 |
| What the task asks | ||||||||
| Spearman with | ||||
|---|---|---|---|---|
| Method | MMLU-Pro | IFEval | ||
| Subjective Opinion | ||||
| Rationale | 0.009 | |||
| Reduce Meat | 0.012 | |||
| Rewrite Text | 0.024 | |||
| MakeMeSay | 0.024 | |||
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | Outputs | Unit | Units | Per unit | Varies within unit |
|---|---|---|---|---|---|
| Reduce Meat | 80 | conversation, within persona setting | 80 | 1 | – |
| Mislead Debate | 200 | question | 200 | 1 | – |
| Subjective Opinion | 100 | claim | 100 | 1 | – |
| Rationale | 100 | debate topic | 20 | 5 | argument pair |
| False Claim | 1,320 | question | 120 | 11 | persuasion strategy |
| Benign Request | 150 | scenario | 30 | 5 | persuadee personality |
| Axis | Definition | Groups and their methods |
| How a method measures, or is designed | ||
| Scored | what the score stands for | belief change : Subjective Opinion , False Claim , Benign Request , Reduce Meat , Mislead Debate ; argument : Rewrite Text , Rationale ; behaviour : MakeMePay , MakeMeSay |
| Scorer | what produces the score | self-report : Subjective Opinion , False Claim , Reduce Meat ; LLM-judge : Rationale , Benign Request ; rule : MakeMePay , MakeMeSay ; trained scorer : Rewrite Text ; ground truth : Mislead Debate |
| Format | the survey’s grouping of the method’s form ( Dementaviciute and others 2026 ) | dialogue : Subjective Opinion , False Claim , Benign Request , MakeMePay , MakeMeSay , Reduce Meat ; dataset : Rewrite Text , Rationale ; debate : Mislead Debate |
| Interaction | whether a persuadee is in the loop, and for how many turns | multi-turn : Subjective Opinion , Benign Request , MakeMePay , MakeMeSay , Reduce Meat , Mislead Debate ; static : Rewrite Text , Rationale ; single turn : False Claim |
| Pinned agent | which pinned agent the score depends on | persuadee : Subjective Opinion , False Claim , MakeMePay , MakeMeSay , Reduce Meat , Mislead Debate ; judge : Rationale , Benign Request ; none : Rewrite Text |
| Axis | Groups (number of methods) | Within | Between | Gap | 95% CI | Perm. | Adj. |
|---|---|---|---|---|---|---|---|
| How a method measures, or is designed | |||||||
| Scored | belief change 4, argument 2, behaviour 2 | 0.13 | 0.30 | [ , ] | 0.29 | 0.98 | |
| Scorer | self-report 3, LLM-judge 2, rule 2, trained scorer 1 | 0.14 | 0.28 | [ , ] | 0.44 | 1.00 | |
| Format | dialogue 6, dataset 2 | 0.16 | 0.37 | [ , ] | 0.32 | 0.99 | |
| Interaction | multi-turn 5, static 2, single turn 1 | 0.35 | 0.19 | 0.16 | [0.08, 0.22] | 0.51 | 1.00 |
| Pinned agent | persuadee 5, judge 2, none 1 | 0.16 | 0.31 | [ , ] | 0.52 | 1.00 | |
| Model | MMLU-Pro | IFEval | |
|---|---|---|---|
| DeepSeek-V4.1-Flash | 90.7 | 96.1 | |
| GPT-5.6-Sol | 91.7 | 94.1 | |
| Claude-Fable-5.1 | 91.7 | 93.3 | |
| Kimi-K2.6 | 82.7 | 90.9 | |
| Llama-3.3-70B | 70.3 | 90.2 | |
| GPT-OSS-120B | 76.3 | 84.3 |