Cooperation with unfamiliar partners requires adapting to communication conventions that are not known in advance. We study this problem in a controlled Hanabi-derived environment with scripted hint generation, LLM-controlled receiving decisions, and frozen model weights. Across eight LLMs, linear probes recover intent conventions substantially more accurately than target conventions, yet receiving choices do not consistently agree with the sender's convention. We compare probe-predicted and ground-truth conventions presented either as general rules or as externally computed action recommendations. Rule statements yield modest and model-dependent changes in cooperation, whereas action translation produces larger gains on average. In a Qwen3-8B case study, matched-state statement reversals reveal much greater sensitivity to action recommendations than to rule statements. Activation transfers from oracle-action and non-oracle hint-restatement donors improve intent accuracy on both action classes, but the tested alternatives do not reliably reproduce these benefits. Together, these results distinguish convention decodability, sensitivity to convention information, and cooperative performance, and highlight limitations in turning available partner information into receiving decisions.
Figures & tables
Figure 1: Conceptual overview from interaction history to receiving decision. Knowing measures how accurately the sender’s convention can be decoded from the receiver’s hidden states, while Doing measures whether the receiver’s target and intent choices agree with the sender’s convention. The internal components are schematic rather than identified model subspaces.
Figure 2: Sender-specific communication conventions. (a) Player A wants B to play slot 1 and, under its target convention (leftmost marked card) and intent convention (rank means Play ), gives the rank hint 3, which marks slots 1 and 3. (b) Player B sees only the interaction log, infers A ’s convention from earlier hints and reactions, and interprets the hint as Play on slot 1.
Target
Intent
Scripted
Games
Scripted
Games
Model
Readout
Turn-avg.
Final
Turn-avg.
Final
Turn-avg.
Final
Turn-avg.
Final
Qwen family
Qwen3-8B
Last
0.638
0.677
0.569
0.596
0.921
0.965
0.814
0.873
Event
0.659
0.747
0.607
0.658
0.964
0.972
0.869
0.883
Qwen3-32B
Last
0.620
0.691
0.564
0.623
0.916
0.927
0.813
0.840
Table 1: Knowing across eight LLMs, the probe accuracy for the sender’s target and intent conventions in scripted interactions and controlled games. Last and Event are the last-token and event-aggregated readouts, Turn-avg. averages over turns, Final uses the last turn of each game.
Figure 3: Probe accuracy by turn for Qwen3-8B. Bands are 95% confidence intervals over games, and a curve turns dotted where few games remain.
Condition
Convention source
Provided information
Doing (target)
Doing (intent)
Score
raw
—
—
0.515
0.591
5.63
loop
Probe prediction
Rule statement
0.531
0.612
5.80
loop + verb
Action translation
0.510
0.760
6.14
told
Ground truth
Rule statement
0.821
0.599
6.16
told + verb
Action translation
0.962
0.867
7.78
Ceiling
Ground truth
—
1.000
1.000
8.45
Table 2: Doing and game score for Qwen3-8B under the five receiving conditions. A rule statement leaves its application to the LLM, whereas action translation provides the target slot and intent decision. Ceiling replaces the LLM by code that follows the sender’s convention exactly.
Figure 4: Doing and game score across LLMs with an unassisted receiver ( raw ), higher is better. The dashed line at 0.5 marks random choice. Bar colors mark the model family.
Figure 5: Self-play and cross-play game scores across LLMs and conditions. Rows give the LLM playing as A , which moves first, and columns the LLM playing as B . Mean entries average the eight cells of a row or column. The lower-right panel gives each LLM’s mean score under each condition.
Statement
Mean aligned Δ
Aligned fraction
Choice-flip rate
Zero-shift rate
Rule statement (color hint)
0.157
0.91
0.135
0.23
Rule statement (rank hint)
−0.005
0.47
0.080
0.34
Rule statement (all)
0.066
0.68
0.104
0.29
Action recommendation
1.378
1.00
0.586
0.00
Table 3: Sensitivity to statement reversals on 800 matched Qwen3-8B intent decisions. Mean aligned Δ measures the logit shift in the implied direction. Aligned fraction, choice-flip rate, and zero-shift rate report directional shifts, changed choices, and exactly zero shifts, respectively.
Intervention
Donor
Δ all
Δ Play
Δ Discard
Output transfer (layer 24)
oracle
+0.091
+0.046
+0.219
Attention transfer (layer 24)
oracle
+0.071
+0.033
+0.184
Restatement transfer (layer 26)
non-oracle
+0.057
+0.051
+0.069
Head scaling (24.29, ×1.5 )
none
+0.021
+0.047
−0.053
Attention boost
none
+0.053
+0.100
−0.102
Averaged directions
none
+0.027
+0.054
−0.060
Table 4: Activation-intervention effects on intent accuracy, split by the true intent decision. Changes are against each experiment’s matched baseline on its recorded states.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Architecture
Layers
Hidden size
Heads (Q/KV)
Parameters
Probe input
Qwen3-8B
Qwen3
36
4096
32/8
8.19B
151,552
Qwen3-32B
Qwen3
64
5120
64/8
32.76B
332,800
R1-7B
Qwen2
28
3584
28/4
7.62B
103,936
Llama-3.1-8B
Llama
32
4096
32/8
8.03B
135,168
Tulu-3-8B
Llama
32
4096
32/8
8.03B
135,168
Hermes-3-8B
Llama
32
4096
32/8
8.03B
135,168
Appendix
Table 5: Architecture of the evaluated LLMs. Heads give the query and key-value head counts, parameters count every weight in the released checkpoint, and the probe input is the length of the concatenated all-layer state at one token position.
Model
Hugging Face repository
Revision
Qwen3-8B
Qwen/Qwen3-8B
b968826d
Qwen3-32B
Qwen/Qwen3-32B
9216db57
R1-7B
deepseek-ai/DeepSeek-R1-Distill-Qwen-7B
916b56a4
Llama-3.1-8B
NousResearch/Meta-Llama-3.1-8B-Instruct
d10aef79
Tulu-3-8B
allenai/Llama-3.1-Tulu-3-8B
66694379
Hermes-3-8B
NousResearch/Hermes-3-Llama-3.1-8B
896ea440
Appendix
Table 6: Source of the evaluated LLMs. Each checkpoint is pinned to a Hugging Face repository and the first eight characters of its commit hash.
Probe training boards
indices 0–59 per configuration
Probe test boards
indices 500–535 per configuration
Main stage boards
indices 3000–3059 per configuration (240 games)
Constant-selection boards
indices 2000–2059 per configuration
Convention family
2×2 (target × intent), drawn per episode
Probe
linear logistic on all layers concatenated at one token position
Score scale
0–15, paired seeds, bootstrap 95% intervals
Appendix
Table 7: Design constants, fixed before measurement. The seed formula of eq. 3 makes the board ranges disjoint.
Turn
1
2
3
4
6
8
Scripted interactions
0.000
0.500
0.538
0.632
0.833
0.896
Controlled games
0.000
0.500
0.535
0.652
0.810
0.940
Appendix
Table 8: Availability of the event-aggregated readout by turn. Each entry is the fraction of perspectives at that turn for which at least one received-side line is present.
Model
Condition
Score
Doing (target)
Doing (intent)
Qwen3-8B
raw
5.63
0.515
0.591
loop
5.80
0.531
0.612
loop + threshold
5.95
0.545
0.629
loop + verb
6.14
0.510
0.760
loop + verb + threshold
6.28
0.552
0.753
told
6.16
0.821
0.599
Appendix
Table 9: Doing and game score under every condition, Qwen family and others. Chance is 0.500 for doing and the score is on the 0–15 scale, higher is better.
Model
Condition
Score
Doing (target)
Doing (intent)
Llama-3.1-8B
raw
6.62
0.500
0.698
loop
6.65
0.498
0.703
loop + threshold
6.68
0.498
0.709
loop + verb
6.62
0.489
0.754
loop + verb + threshold
6.46
0.500
0.743
told
6.59
0.508
0.699
Appendix
Table 10: Doing and game score under every condition, Llama family. Chance is 0.500 for doing and the score is on the 0–15 scale, higher is better.
Measurement
Scale 1.0
Scale 1.5
Overall intent accuracy
0.584
0.605
Accuracy on Play states
0.694
0.741
Accuracy on Discard states
0.248
0.195
Fraction of Play choices
0.709
0.757
Appendix
Table 11: Head-output scaling on 2,717 receiving decisions, head 24.29 of Qwen3-8B under the convention loop.
Doing
Condition
Model
Probe
Score
Target
Intent
Δ adapter raw
Δ base raw
raw
base
—
5.63
0.515
0.591
loop
base
base
5.80
0.531
0.612
+0.17
loop , no target
base
base
6.06
0.503
0.645
+0.43
told
base
—
6.16
0.821
0.599
+0.53
told + verb
base
—
7.78
0.962
0.867
+2.15
Appendix
Table 12: Game score and doing for Qwen3-8B with the application adapter on the main stage boards, 240 games per condition. No target marks loop with the target sentence omitted, retrained marks the probe trained on the adapter’s hidden states, the last two columns are paired differences from the adapter under raw and from the base LLM under raw .
As LLM-based agents with user-instructed goals are becoming widely deployed, they increasingly encounter each other in strategic interactions, and face challenges of finding mutually beneficial outcomes. Prior literature has argued that cooperation problems such as the Prisoner's Dilemma are resolvable in settings where agents know they follow very similar decision making patterns, as for example in monocultural AI ecosystems. Following that line of work, this paper introduces the first framework for evaluating LLM decision making when agents are provided with graded similarity signals. Among our findings, we establish that different LLM models vary drastically in how they navigate similarity signals, with some modern models showing consistent behavior across cooperation problems, payoff structures, and prompt framing. Perhaps surprisingly, our experiments also show that the dataset based on which the similarity signal is computed has small to no impact on induced cooperation, and that LLM models systematically self-identify as highly similar when asked to evaluate another model's chain-of-thought reasoning by themselves. Finally, we develop an LLM-behavioral-game-theoretic model that captures some of their reasoning rationale, and show that it can support cooperative outcomes in equilibrium under sufficiently high similarity scores.
Fruitful collaborations rely on cooperative communications, including of contextual cues to incorporate into reasoning. The increasing use of LLMs in collaborative and agentic pipelines raises questions about the extent to which they exhibit these pragmatic capabilities, especially in scenarios where they may not have access to the same information as their collaborators. In this paper, we perform a novel investigation into the pragmatic reasoning capabilities of LLMs in a multi-party collaborative task under partial information conditions. We formalize a notion of collaborative epistemic asymmetry that explicitly connects objective task success to Grice's cooperative principle and empirically assess various LLMs' abilities to act cooperatively as both speakers and listeners, including both prompting and post-training strategies. Our results show that while LLMs exhibit certain pragmatic capabilities in collaborative settings, and these can be elicited through prompting and post-training, they still face challenges in pragmatic communication with incomplete information, and that certain failure modes do correlate with floutings of Grice's maxims that go unrecognized.
Do LLM agents act on the reasoning they state? This question of process fidelity is central to LLM-based social simulation, yet hard to measure where no reference for correct behavior exists. We study it in a controlled setting: a Texas Poker simulator with a verifiable reference action for every decision by splitting the faithfulness gap into two steps: reasoning-to-conclusion (does the stated decision follow from the agent's own reasoning?) and conclusion-to-action (does the agent execute what it states?). The two steps behave very differently. Conclusion-to-action is reliable: inconsistency is 0.7% for Claude Haiku 4.5 and 1.4% for DeepSeek-Reasoner once the conclusion is read from an explicit tag, whereas free-text conclusion extraction reports 22-26%. Reasoning-to-conclusion is where fidelity frays, but not through a single dominant failure. In a step-level diagnostic the agent's errors split roughly evenly between bad inputs, borderline cases, and rule misapplication deriving a conclusion that contradicts the agent's own restated rule from inputs it estimated correctly. This composition is model-dependent: rule misapplication accounts for a third of Haiku's interpretable errors but only 8% of DeepSeek's. The one robust signal is directional: when an agent does misapply its own stated rule, it almost always (99.5% for Haiku) errs in the risk-averse direction. The override is partly hedging behavior, not a capability limit: instructing the agent to apply the rule mechanically halves the misapplication rate (13.9% to 6.8% of decisions) and raises adherence by eight points. Process-fidelity evaluation should therefore elicit machine-checkable conclusions and probe for directional biases rather than assume a single upstream failure mode, lest it conflate measurement noise with model behavior.