Disentangling Models from Personas in Heterogeneous LLM Simulations
Organizations: Carnegie Mellon University
Abstract
Multi-agent simulations with large language models (LLMs) often operate networks of agents with a single base model. This overlooks the inter-model effects which may dominate engagement dynamics in real-world deployments. To show this, we simulate a heterogeneous social network powered by several different base models and show that the amount of engagement an agent receives depends more on its base model than on its assigned persona. The attraction or repulsion effects of a base model strengthen dramatically when more models are added in the mix, suggesting that networks dynamics may converge to base model effects at scale. To help explain this effect, we conduct a series of content-mediating analyses, showing the predictability of base models across contexts as well as the relationship between a model's lexical patterns and an engagement-maximizing style. In light of recent developments in mass multi-agent interaction, this work underscores the relevance of heterogeneous compositions in driving the outcomes of those networks
Figures & tables
| Excluded | Regime | Persona | Own | Selection | |
|---|---|---|---|---|---|
| baseline | Two-way | 1586 | 8.6% | 10.2% | 27.8% |
| Three-way | 2093 | 7.0% | 15.6% | 26.7% | |
| Qwen | Two-way | 923 | 12.4% | 6.2% | 20.5% |
| Three-way | 781 | 14.2% | 19.2% | 23.0% | |
| gpt-oss | Two-way | 893 | 10.4% | 11.7% | 27.2% |
| Three-way | 825 | 9.8% | 9.3% | 14.0% |
| model | CV acc. | LPO acc. | characteristic phrases |
|---|---|---|---|
| gpt-oss-20b | 0.98 | 0.97 | merkle root , audit trail , zk snark , tamper evident |
| Qwen3-32B | 0.88 | 0.87 | let’s build , let’s make , here’s twist , i’ll draft |
| GLM-4-32B | 0.85 | 0.84 | beautifully captures , resonates deeply , aligns perfectly |
| Magistral-Small | 0.88 | 0.85 | ah user , alright listen , strikes chord , neon lights |
| Gemma-4-31B | 0.98 | 0.94 | you’re just , stop trying , just fancy , fancy way |
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
| Metric | Value |
|---|---|
| Avg. cross-comments / agent | |
| Gini of in-degree | |
| Post:comment ratio | |
| Gini of posts / agent | |
| Max / mean in-degree |
| Regime | N runs | Max-dev mean std |
|---|---|---|
| Dyadic (target 50%/50%) | 100 | |
| Triadic (target 33.3%/33.3%/33.3%) | 90 |
| gpt-oss comment share | fraction of runs |
|---|---|
| Model | Q0 | Q1 | Q2 | Q3 |
|---|---|---|---|---|
| GLM-4 | ||||
| Mag | ||||
| Qwen | ||||
| gpt-oss | ||||
| Gemma |
| Source | Q0 | Q1 | Q2 | Q3 |
|---|---|---|---|---|
| GLM-4 | ||||
| Mag | ||||
| Qwen | ||||
| gpt-oss | ||||
| Gemma |
| Qwen | GLM-4 | Mag | oss | Gem | |
|---|---|---|---|---|---|
| Qwen | |||||
| GLM-4 | |||||
| Mag | |||||
| oss | |||||
| Gem |
| Qwen | GLM-4 | Mag | oss | Gem | |
|---|---|---|---|---|---|
| Qwen | |||||
| GLM-4 | |||||
| Mag | |||||
| oss | |||||
| Gem |
| composition | runs | gpt-oss | Qwen | GLM-4 | Magistral | Gemma |
|---|---|---|---|---|---|---|
| without Gemma | 24 | — | ||||
| without Qwen | 21 | — | ||||
| without Magistral | 12 | — | ||||
| without GLM-4 | 6 | — | ||||
| pooled | 63 |
| Block | Dy | Tri | Tet |
|---|---|---|---|
| Persona only | 0.086 | 0.070 | 0.056 |
| Model only | 0.102 | 0.156 | 0.317 |
| Model selection (own + partner) | 0.278 | 0.267 | 0.317 |
| gpt-oss indicator only | 0.008 | 0.093 | 0.288 |
| Run only | 0.350 | 0.219 | 0.049 |
| Persona + model | 0.190 | 0.224 | 0.364 |
| Block | Dy | Tri | Tet |
|---|---|---|---|
| Persona only | 0.068 | 0.025 | |
| Model only | 0.067 | 0.149 | 0.316 |
| Persona + model | 0.147 | 0.183 | 0.297 |
| Dyad | Triad | Tetrad | ||||
|---|---|---|---|---|---|---|
| Model | raw | adj | raw | adj | raw | adj |
| gpt-oss | ||||||
| Gemma | – | – | ||||
| Qwen | ||||||
| GLM-4 | ||||||
| Magistral | ||||||
| representation | model CV | model LPO | persona CV |
|---|---|---|---|
| TF-IDF (topic + style) | 0.80 | 0.70 | 0.73 |
| StyleDistance (style) | 0.76 | 0.74 | 0.22 |
| Gemma (semantic) | 0.79 | 0.75 | 0.83 |
| MPNet (semantic) | 0.72 | 0.66 | 0.76 |
| axis | top terms | top terms | model spread (mean score) |
|---|---|---|---|
| 1 | let’s, just, digital, i’m | ici je, citer, je viens | all (language) |
| 2 | let’s, i’m, systems, you’re | proof, token, audit, zk | gpt-oss vs rest |
| 3 | actually, stop, noise, finally | explore, journey, greetings | Gemma vs rest |
| 4 | strategic, governance, architecture | tech, hey, share, got | Gemma vs Mag |
| model | persona accuracy | chance |
|---|---|---|
| Gemma | ||
| Magistral | ||
| gpt-oss | ||
| GLM-4 | ||
| Qwen |
| feature set | single-set | unique |
|---|---|---|
| length | ||
| style | ||
| model | ||
| persona | ||
| all four |
| model | direction stability | cos to engagement |
|---|---|---|
| gpt-oss-20b | ||
| Gemma-4-31B | ||
| Qwen3-32B | ||
| Magistral-Small | ||
| GLM-4-32B |