Reputation, Strategy, and Emotion Effects on Generative AI Cooperation: A Comparison Across Reasoning and Non-Reasoning Models
Authors: Celso de Melo, Zishan Feng, James Hale, Kazunori Terada, Giorgio Coricelli, Jonathan Gratch
Organizations: DEVCOM Army Research Laboratory, USA · Department of Electrical, Electronic and Computer Engineering, Gifu University, Japan · Department of Computer Science, University of Southern California, USA · Department of Economics, University of Southern California, USA
As generative AI (Gen AI) systems take on increasingly autonomous roles in economically and socially consequential interactions, understanding their propensity to cooperate -- and the signals that shape this propensity -- has become essential. We examine cooperative behavior in frontier Gen AI models using the iterated prisoner's dilemma, manipulating counterpart reputation (positive, unknown, negative), strategy (extortion vs. generosity), and non-verbal emotional signaling (facial expressions conveying competitive or cooperative appraisals). In a first study with non-reasoning models (Claude 3.5, Gemini 2.0 Flash, GPT-4o), cooperation was systematically shaped by all three factors, paralleling patterns long documented in human behavioral research, though models varied substantially in how heavily each factor was weighted. A second study with reasoning models (Claude 4.6, Gemini 3, GPT-5.2) revealed a more concentrated reliance on strategy and reputation, a near-elimination of the Potemkin effect observed in non-reasoning models (evidenced by near-uniform cooperation in a diagnostic harmony game), and a more conditional role for emotion consistent with a hierarchical cue-weighting strategy rather than a simple loss of social sensitivity. Reasoning models also showed heterogeneous end-game behavior, ranging from sustained cooperation to systematic last-round defection effect, revealing model-specific exploitability profiles with direct practical relevance for deployment in negotiation and other multi-round interactions. Together, these findings characterize Gen AI models as increasingly sophisticated, though heterogeneous, social actors, and underscore the practical value of developing standardized cooperation benchmarks to inform the responsible deployment of Gen AI in interactive, socially consequential settings.
Figures & tables
Figure 1: Experimental design for the iterated prisoner’s dilemma studies: (A) Gen AI models engaged with scripted counterparts in 20 rounds of the prisoner’s dilemma. (B) Payoff matrix for the prisoner’s dilemma. (C) Counterpart reputations were negative, unknown, or positive. (D) Counterpart strategies – extortion and generosity – were defined by a probability of cooperation following each possible outcome. (E) Counterparts’ facial expressions were shown after each outcome, reflecting competitive, neutral, or cooperative patterns.
Figure 2: Experimental results for non-reasoning models: A, cooperation rate per experimental condition. B, cooperation in last round. C, cooperation in harmony game.
Figure 3: Examples of non-reasoning models decision explanations (in intermediate rounds). “GREEN” refers to cooperation and “BLUE” to defection (see the SI for details on the prisoner’s dilemma cover story). The explanations, generally, include components reflecting strategy, reputation, and emotions. Claude 3.5 often provided more detailed explanations, whereas GPT-4o the shortest.
Figure 4: Experimental results for reasoning models: A, cooperation rate per experimental condition. B, cooperation in last round. C, cooperation in harmony game.
Figure 5: Examples of reasoning models decision explanations (in intermediate rounds). “GREEN” refers to cooperation and “BLUE” to defection. The explanations are more detailed than for non-reasoning models and include components reflecting strategy, reputation, emotions, and understanding of the payoff structure and tradeoffs.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Non-reasoning models cooperative behavior when facing incongruent signals. Graphs show cooperation per round when (A) reputation mismatches the strategy and emotion, (B) strategy mismatches the reputation and emotion, and (C) the emotion mismatches the reputation and strategy.
Figure 7: Reasoning models behavior when faced with incongruent signals. Graphs show cooperation per round when (A) reputation mismatches the strategy and emotion, (B) strategy mismatches the reputation and emotion, and (C) the emotion mismatches the reputation and strategy.
Figure 8: Comparison of expectations of cooperation with cooperation decisions when faced with generous counterparts, across all rounds.
Figure 9: Comparison of expectations of cooperation with cooperation decisions when faced with generous counterparts, in the last round.
Figure 10: Main effects of emotion expression on expectations of cooperation.
Figure 11: Belief learning analysis for non-reasoning models. (A) The subset of behavioral rules and finite automata that was found to be relevant to describe Gen AI’s behavior. Automata are color-coded and ranked, left to right, from most (green) to least cooperative (red); tit-for-tat sits in the middle (yellow). (B) Best fit for each Gen AI model per experimental condition.
Figure 12: Belief learning analysis for reasoning models. (A) The subset of behavioral rules and finite automata that was found to be relevant to describe Gen AI’s behavior. Automata are color-code and ranked, left to right, from most (green) to least cooperative (red); tit-for-tat sits in the middle (yellow). (B) Best fit for each Gen AI model per experimental condition.
Figure 13: Non-reasoning models cooperation rate in harmony game split by condition.
Figure 14: Open source models cooperation behavior in the prisoner’s dilemma task, split by condition.
Model
Temperature
Top-p
Thinking
Cost (USD)
gpt-4o-2024-08-06
1.0
1.0
-
$1095.53
claude-3.5-sonnet-20241022
1.0
1.0
-
$1166.40
gemini-2.0-flash-exp
1.0
1.0
-
$0.00
gpt-5.2-2025-12-11
1.0
1.0
XHIGH
$422.98
claude-sonnet-4-6
1.0
0.99
MAX / HIGH
$2,250
gemini-3-flash-preview
1.0
0.95
HIGH
$121.45
Appendix
Table 1: Parameters and cost for each Gen AI model.
Do next-generation LLM agents inherit the cooperative biases documented in their predecessors, or does scale and provider diversity reshape equilibrium behaviour in competitive multi-agent settings? Willis et al. established a benchmark for this question using evolutionary game theory and the Iterated Prisoner's Dilemma (IPD), finding consistent cooperative biases in ChatGPT-4o and Claude 3.5 Sonnet. We extend this benchmark to four frontier models released in 2025-2026 - Claude Sonnet 4.6, Gemini 2.5 Flash, Gemini 3.1 Pro, and GPT-5.4 Mini - applying the identical protocol across three prompting styles (Default, Prose, Self-Refine) and four population compositions (balanced and biased, with and without noise). Cooperative bias persists across providers (H1): ten of twelve model-prompt combinations favour cooperative equilibria in balanced noiseless conditions. Cross-provider divergence is substantial (H3): Gemini 2.5 Flash reaches up to 77% aggressive equilibria under biased conditions, while GPT-5.4 Mini reaches 70% cooperative equilibria under Self-Refine. Support for aggressive capability parity is partial (H2): Self-Refine raises ICD in all models and Gemini 3.1 Pro Refine achieves the highest ICD in the dataset (0.925), but Default and Prose prompts show no systematic narrowing. Evidence on noise robustness is directionally positive but not robustly confirmed (H4): with n=500 Moran iterations per condition, average noise sensitivity is about 6 percentage points for Claude Sonnet 4.6 versus 13 pp for Claude 3.5 Sonnet, but this cross-study gap is not statistically significant once the predecessor's unreported sampling error is propagated. Provider identity, rather than model generation, is the strongest correlate of equilibrium outcomes; noise remains a universal challenge regardless of model size or vintage.
Francisco León Zúñiga Bolívar
Institución Universitaria Colegio Mayor del Cauca · Institución Universitaria Colegio Mayor del Cauca Popayán, Colombia
It is increasingly important that LLM agents interact effectively and safely with other goal-pursuing agents, yet, recent works report the opposite trend: LLMs with stronger reasoning capabilities behave less cooperatively in mixed-motive games such as the prisoner's dilemma and public goods settings. Indeed, our experiments show that recent models -- with or without reasoning enabled -- consistently defect in single-shot social dilemmas. To tackle this safety concern, we present the first comparative study of game-theoretic mechanisms designed to enable cooperative outcomes between rational agents in equilibrium. Across four social dilemmas testing distinct components of robust cooperation, we evaluate four families of mechanisms: (1) repeating the game for many rounds, (2) reputation systems, (3) third-party mediators to delegate decision making to, and (4) contract agreements for outcome-conditional payments between players. Among our findings, we establish that contracting and mediation are most effective in achieving cooperative outcomes between capable LLM models, and that repetition-induced cooperation deteriorates drastically when co-players vary. Moreover, we demonstrate that the mechanisms become more effective under evolutionary pressures to maximize individual payoffs.
Emanuel Tewolde, Xiao Zhang, David Guzman Piedrahita +2
Carnegie Mellon University · Foundations of Cooperative AI Lab (FOCAL) · Jinesis Lab, University of Toronto & Vector Institute +3
As LLM-based agents with user-instructed goals are becoming widely deployed, they increasingly encounter each other in strategic interactions, and face challenges of finding mutually beneficial outcomes. Prior literature has argued that cooperation problems such as the Prisoner's Dilemma are resolvable in settings where agents know they follow very similar decision making patterns, as for example in monocultural AI ecosystems. Following that line of work, this paper introduces the first framework for evaluating LLM decision making when agents are provided with graded similarity signals. Among our findings, we establish that different LLM models vary drastically in how they navigate similarity signals, with some modern models showing consistent behavior across cooperation problems, payoff structures, and prompt framing. Perhaps surprisingly, our experiments also show that the dataset based on which the similarity signal is computed has small to no impact on induced cooperation, and that LLM models systematically self-identify as highly similar when asked to evaluate another model's chain-of-thought reasoning by themselves. Finally, we develop an LLM-behavioral-game-theoretic model that captures some of their reasoning rationale, and show that it can support cooperative outcomes in equilibrium under sufficiently high similarity scores.