Large foundation models have accelerated progress toward general-purpose agents that interact with humans and other agents through language and multimodal signals. However, robust multi-agent decision-making requires reasoning about what other agents know, intend, and are likely to do under partial observability. Current agentic systems often operate through prompt design, memory, or end-to-end behavioral shaping, but typically do not learn an explicit partner-state representation that can be reused as a decision variable across tasks. We introduce \emph{mental-model-enabled agents}, a framework that equips an agent with a latent mental model of its counterpart, allowing it to infer hidden beliefs, intentions, and likely reactions from the observed history and use these inferences to guide action selection. Our method learns an amortized recursive Theory-of-Mind representation, with first- and second-order mental-state structure, jointly with a belief-conditioned reward model that evaluates candidate actions relative to the inferred partner state. A policy is then learned under this belief-aware signal, yielding an agent that can act independently at inference time while retaining the benefits of explicit partner modeling. We evaluate the same framework on both language-only and multimodal benchmarks. Across these settings, explicit mental-state modeling consistently improves interaction quality and Theory-of-Mind performance over base agentic systems, showing that structured partner modeling is a useful inductive bias for general multi-agent systems. Our code is publicly available at https://github.com/hananshafi/Mental-Models
Figures & tables
Figure 1: Mental-model-enabled decision making under partial observability. Observed behavior only partially reveals another agent’s hidden goals, beliefs, constraints, and likely strategy. By inferring these latent mental states, a mental-model-enabled agent can choose more adaptive actions, improving outcomes while preserving cooperation across multi-agent settings.
Figure 2: Method overview. Phase 1 jointly learns an amortized recursive mental-model and a mental-conditioned reward model from multi-agent context. The latent contains first- and second-order belief, intent, and thought subspaces, which condition reward prediction. Phase 2 freezes the mental-state encoder and reward model, scores sampled policy actions with the learned mental reward, and optimizes the policy with mental-reward guidance. Phase 3 deploys only the trained policy in the multi-agent setting.
Table 3Table 4Table 5
Model
Average
Joint
w/o FB 1st
w/o FB 2nd
w/o FB Reality
w/o FB Memory
FB 1st
FB 2nd
Classic TOMi baselines
MemNN
77.20
44.30
85.45
82.67
93.39
98.90
12.62
17.27
RelNet
86.00
57.40
96.42
95.37
100.0
99.90
10.40
17.81
EntNet
90.60
66.80
94.29
85.08
100.0
100.0
54.95
36.55
Closed source LLMs
GPT-3.5-Instruct
79.78
36.94
94.42
84.05
97.30
94.79
34.26
7.82
Table 7: ToMi Generalization from the BigToM -trained policy. Ours improves the aggregate Average and Joint scores over the base model and strengthens non-false-belief first-/second-order belief tracking, highlighting OOD generalization of our mental-model trained policy.
Table 7
Figure 4: a) z -scrambling barely affects the compression VAE but sharply increases BIT reward-vector MSE, showing stronger reward-causal structure. b) Pairwise CKA shows complementary belief, intent, and thought slices. c) Frozen-latent probes recover ToM-relevant MMRole factors more separably from BIT than from a flat latent. d) Shuffling z2 weakens second-order probes, supporting the recursive z1→z2 structure.
Figure 5: Latent traversal analyses. a) Belief-axis traversal shifts both the branch classifier and frozen reward head beyond random controls. b) Offer/proposal traversal increases strategy signal and local label enrichment, with stronger second-order contribution.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Multi- modal
Episode Reward
Per-turn Reward
Intent Annot.
1st-order Belief
2nd-order Belief
Rationale / CoT
Hard Negatives
BigToM (original)
✗
✗
✗
✗
✓
✗
✗
✗
BigToM (ours)
✗
✗
✗
✓
✓
✓
✓
✓
Sotopia - π (original)
✗
✓
✗
✗
✗
✗
✗
✗
Sotopia - π (ours)
✗
✓
✓
✓
✓
✓
✓
✓
MMRole (original)
✓
✓
✗
✗
✗
✗
✗
✗
MMRole (ours)
✓
✓
✓
✓
✓
✓
✗
✓
Appendix
Table 10: Comparison of original and augmented datasets across annotation dimensions.
Subset
n
Mean ( 1 – 5 )
Faithfulness (%)
MMRole
50
4.30
86.0
Sotopia
50
4.60
92.0
Overall
100
4.45
89.0
Appendix
Table 11: Blinded human audit of generated mental-state annotations. Faithfulness linearly rescales the mean rating to 0 – 100% .
Model
Belief choice
Belief distance
First order
Second order
Base
25.4
35.2
26.8
49.7
SFT
26.5
33.8
29.4
41.4
Ours
27.9
38.0
29.7
52.2
Appendix
Table 12: Zero-shot BigToM → FANToM transfer on 5,000 probes using the official scorer. Scores are percentages; higher is better.
Figure 6: Mental-latent geometry from an unsupervised UMAP of latent z . Left: coloring by task shows task/scenario macro-clusters. Right: coloring by branch shows aware and not-aware samples remain near paired scenarios but are displaced within task regions, consistent with branch F1 =0.989 .
Without Initial Belief
With Initial Belief
Task
Model
TB
FB
TB ∧ FB
TB
FB
TB ∧ FB
Forward Belief
GPT-4 (0-shot)
99.0
98.0
97.0
91.0
99.0
90.0
GPT-4 (0-shot-CoT)
99.0
99.0
98.0
99.0
99.0
98.0
GPT-4 (1-shot)
99.0
99.0
97.0
97.0
99.0
96.0
GPT-4 (1-shot-CoT)
100.0
99.0
99.0
97.0
99.0
96.0
Base
83.0
94.0
77.0
49.5
98.5
48.0
Appendix
Table 13: BigToM official-protocol results. TB denotes true-belief accuracy, FB denotes false-belief accuracy, and TB ∧ FB denotes paired accuracy, where both the true-belief and false-belief versions of the same scenario must be answered correctly.
Figure 7: Range-normalized lift from Qwen2.5-7B (ours w/ 1st order) to Qwen2.5-7B (ours w/ 2nd order) on Sotopia . Petal lengths encode percentage-of-range changes for each official dimension rather than raw scores. The center reports the relative lift in the average score, highlighting that second-order ToM improves the aggregate profile while keeping the comparison scale-free.
Figure 8: Preference-margin distribution. BIT decomposition produces sharper preference margins than flat or shuffled mental supervision. Kernel density of the per-sample margin r+−r− on 1,508 Sotopia validation turns. Dotted vertical lines mark medians; the dashed line marks zero. BIT ranks positives over negatives with roughly 2× higher confidence (median +25.06 vs. flat +12.76 ), and every BIT sample is correctly ranked with margin ≥+1.40 . Flat overlaps the shuffled-label control, indicating that mental supervision improves calibration when belief, intent, and thought are routed to dedicated latent sub-blocks.
Figure 9: Cross-agent transfer of the mental-reward model on Sotopia . Bars compare the average scores of LLaMA2-7B (ours) with its same-architecture mental-reward model and LLaMA2-7B (ours w/ r-Qwen) with the Qwen mental-reward model. The transferred Qwen mental-reward model preserves overall performance, supporting its generalization across agents.
Figure 10: Role-swap latent intervention. Swapping a BIT role-specific block moves the corresponding role prediction more than unrelated roles. Flat and Shuffled variants do not show this separation, yielding more diffuse effects.
Figure 11: t-SNE visualization The concatenated latent representation [z1; z2] reaches 0.662 5-NN accuracy (silhouette +0.046) on the 5-way emotion task compared to 0.529 for Qwen without the adapter, and a chance baseline of 0.200. The +0.13 gap over the no-adapter baseline indicates that the learned mental latents encode emotion-relevant structure that is not present in the base Qwen representation. The corresponding t-SNE projection provides a qualitative visualization: classes appear more separated in the z1 and [z1; z2] projections, while the no-adapter Qwen projection shows more inter-class overlap.
Figure 12: Visual Grounding. Visual grounding rate per response for Base vs. Ours on 200 held-out MMRole examples. Each point is one example’s share of generated sentences that mention a gold scene object; black markers show means. The metric is computed using a token-overlap classifier between generated sentence tokens and the gold scene_objects token bag, making it a proxy for grounding rather than grounding itself. Ours increases the rate by Δ=+0.103 (paired Wilcoxon p=1.3×10−4 , ***).
Figure 13: Probe signal accessibility. We compare the recursive mental-model latent against a compression-only VAE trained on the same Sotopia - π contexts and responses with matched latent dimensionality. The VAE is optimized only for reconstruction, while the mental-model encoder is trained with ToM supervision and reward coupling. After freezing both encoders, we fit simple probes from z and evaluate held-out recovery of social-label structure and reward-vector targets. Higher macro-F1/ R2 for the mental latent shows that downstream social and reward signals are more accessible than in a generic compression code.