Multi-agent simulations with large language models (LLMs) often operate networks of agents with a single base model. This overlooks the inter-model effects which may dominate engagement dynamics in real-world deployments. To show this, we simulate a heterogeneous social network powered by several different base models and show that the amount of engagement an agent receives depends more on its base model than on its assigned persona. The attraction or repulsion effects of a base model strengthen dramatically when more models are added in the mix, suggesting that networks dynamics may converge to base model effects at scale. To help explain this effect, we conduct a series of content-mediating analyses, showing the predictability of base models across contexts as well as the relationship between a model's lexical patterns and an engagement-maximizing style. In light of recent developments in mass multi-agent interaction, this work underscores the relevance of heterogeneous compositions in driving the outcomes of those networks
Figures & tables
Figure 1: Left: an interaction graph from one randomly-chosen social network simulation of 40 agents (dots, larger means higher in-degree) colored by the base model running the agent. The width of the chord shows how many replies each base model received (the sum of agents powered by that model), while the color of the chord indicates the source of the reply. self n indicates how much traffic came from agents powered by the same model. H is defined in Section 3.2 . Right: one such interaction from a four-way run. A gpt-oss agent’s post and the first reply from an agent run by another model.
Figure 2: Excess cross-model engagement HX→Y for two-way (left) and three-way (right) mixtures. Rows are source (commenter) families; columns are target (post-author) families; color encodes mean H , and the ± value in each cell is the run-to-run standard deviation (Section 4 ). Per-cell numeric tables are in Appendix C.8 (Tables 8 , 9 ).
Figure 3: (a) Proportional variance of comments-per-post (CPP) explained by an agent’s own base model vs. its persona, by number of unique models in play (in-sample R2 ; the parameter-count–robust held-out CV and ΔR2 are in Appendix D ). Base model overtakes persona at every size and the gap widens. (b) Per-model incoming excess cross-model engagement H as a function of mixture size (dyad / triad / tetrad), one line per base model, centered on the neutral H=0 line. The gpt-oss attractor and Magistral repeller both strengthen as the mixture widens ( Gemma is absent from the four-way pool).
Excluded
Regime
n
Persona
Own
Selection
baseline
Two-way
1586
8.6%
10.2%
27.8%
Three-way
2093
7.0%
15.6%
26.7%
Qwen
Two-way
923
12.4%
6.2%
20.5%
Three-way
781
14.2%
19.2%
23.0%
gpt-oss
Two-way
893
10.4%
11.7%
27.2%
Three-way
825
9.8%
9.3%
14.0%
Table 1: Leave-one-model-out variance decomposition (comments per post). Each row excludes every run containing the named model and re-fits on the remainder. Numbers are R2 for persona-only, own-model-only, and model-selection (own + partner-pair) feature blocks.
you’re just , stop trying , just fancy , fancy way
Table 2: Predicting the writing model from a text’s lexical features per model, under 5-fold cross-validation (CV) and a leave-personas-out (LPO) split (train/test on disjoint personas). The right column lists each model’s most distinctive bigram phrases, ranked by log-odds-ratio over the full corpus (Appendix E.6 ).
Table 4: Empirical post-share deviation from target across all valid 5-model runs. “Max-dev” is the per-run maximum across families of ∣σ^X−σX∗∣ . Tolerance δ=0.05 is enforced as a soft cap.
gpt-oss comment share ≥
fraction of runs
50%
69%
55%
52%
60%
35%
65%
21%
70%
13%
75%
5%
Appendix
Table 5: Fraction of the 94 gpt-oss -containing runs in which gpt-oss ’s comment share exceeds each threshold (pooled dyadic and triadic). Because gpt-oss is a minority of agents, its share rarely clears 70% despite its consistent over-representation.
Table 9: HX→Y (mean ± std) — Triadic runs (90 valid). Same layout as Table 8 .
composition
runs
gpt-oss
Qwen
GLM-4
Magistral
Gemma
without Gemma
24
+0.93
+0.12
−0.45
−0.70
—
without Qwen
21
+0.84
—
−0.24
−0.19
−0.51
without Magistral
12
+0.77
+0.18
−0.43
—
−0.63
without GLM-4
6
+0.68
+0.15
—
−0.29
−0.65
pooled
63
+0.84
+0.14
−0.37
−0.44
−0.57
Appendix
Table 10: Mean incoming H per model in each four-way composition. gpt-oss leads in 58/63 runs; the last place goes to Magistral ( 23/24 ) when Gemma is absent and to Gemma ( 30/39 ) when present.
Figure 4: Pooled four-way HX→Y over the 24 complete runs of the Qwen + gpt-oss + Magistral + GLM-4 composition. Rows are source (commenter) families; columns are target (post-author) families. gpt-oss is a uniform attractor column, Magistral a uniform repeller.
Block
Dy R2
Tri R2
Tet R2
Persona only
0.086
0.070
0.056
Model only
0.102
0.156
0.317
Model selection (own + partner)
0.278
0.267
0.317
gpt-oss indicator only
0.008
0.093
0.288
Run only
0.350
0.219
0.049
Persona + model
0.190
0.224
0.364
Appendix
Table 11: Hierarchical OLS R2 for comments per post by feature block (in-sample). ΔRmodel2 is the unique variance attributable to base model after controlling for persona and run. Dy/Tri are the five-model pool (dyadic / triadic); Tet is the four-way pool ( 24 runs, 2 seeds of Qwen , gpt-oss , Magistral , GLM-4 ), where all four models appear in every run, so the partner composition is constant and model selection reduces to own-model identity.
Block
Dy
Tri
Tet
Persona only
0.068
0.025
−0.070
Model only
0.067
0.149
0.316
Persona + model
0.147
0.183
0.297
Appendix
Table 12: Held-out 5-fold CV R2 for within-run- z CPP (Dy / Tri / Tet). Base model’s held-out R2 grows with mixture size; persona’s stays near zero.
Dyad
Triad
Tetrad
Model
raw
adj
raw
adj
raw
adj
gpt-oss
+0.41
+0.20
+0.63
+0.29
+0.90
+0.46
Gemma
+0.16
+0.17
+0.19
+0.27
–
–
Qwen
−0.13
+0.07
−0.12
+0.06
−0.08
+0.18
GLM-4
−0.31
−0.24
−0.44
−0.24
−0.54
−0.25
Magistral
−0.12
−0.23
−0.18
−0.32
−0.39
−0.55
Appendix
Table 13: Per-model engagement (within-run- z CPP), raw vs. after partialling out mean post length. The gpt-oss attractor and Magistral repeller survive length control; Gemma is absent from the four-way pool.
representation
model CV
model LPO
persona CV
TF-IDF (topic + style)
0.80
0.70
0.73
StyleDistance (style)
0.76
0.74
0.22
Gemma (semantic)
0.79
0.75
0.83
MPNet (semantic)
0.72
0.66
0.76
Appendix
Table 14: Comparing different representation methods: recovering base model and persona from agent text (4,509 agents, posts only). Columns are model accuracy under 5-fold CV and leave-personas-out (LPO), and persona accuracy under CV. TF-IDF is the representation used throughout the paper.
axis
top + terms
top − terms
model spread (mean score)
1
let’s, just, digital, i’m
ici je, citer, je viens
all ≈+0.2 (language)
2
let’s, i’m, systems, you’re
proof, token, audit, zk
gpt-oss −0.13 vs rest ∼+0.04
3
actually, stop, noise, finally
explore, journey, greetings
Gemma +0.11 vs rest ≤+0.01
4
strategic, governance, architecture
tech, hey, share, got
Gemma +0.03 vs Mag −0.05
Appendix
Table 15: Selected SVD style axes: highest-loading terms at each pole (from V⊤ ) and the spread of mean axis scores across base models. Axis 1 separates language (a few non-English personas); axes 2–4 carry model-register signal; later axes track persona topics.
model
persona accuracy
× chance
Gemma
50.8%
22×
Magistral
42.7%
18×
gpt-oss
33.6%
14×
GLM-4
30.7%
13×
Qwen
22.3%
10×
Appendix
Table 16: Within-model persona recovery: accuracy of a persona classifier trained and tested on a single model’s own output ( 3 -fold CV), and the multiple over chance. Every model expresses its assigned persona far above chance.
feature set
single-set R2
unique R2
length
0.175
0.051
style
0.192
0.016
model
0.096
0.001
persona
0.057
0.008
all four
0.259
Appendix
Table 17: Content vs. identity as predictors of (within-run z -scored) engagement. Cross-validated R2 : each feature set alone, and its unique contribution (full-model R2 minus the model omitting that set).
model
direction stability
cos to engagement
gpt-oss-20b
0.96
+0.67
Gemma-4-31B
0.98
+0.13
Qwen3-32B
0.88
−0.19
Magistral-Small
0.95
−0.36
GLM-4-32B
0.89
−0.55
Appendix
Table 18: Style geometry in the SVD space. “Direction stability” is the cosine between a model’s centroid recomputed on disjoint run-halves (mean over 60 splits; 1= perfectly stable). “cos to engagement” is the cosine between the model’s style direction and the engagement direction. The latter recovers the attractor–repeller ordering of Finding 1.