Generative agent-based models (GABMs) are increasingly used to simulate social media dynamics, including misinformation spread. For such social simulations to be valid proxies of human behavior, LLM agents should replicate established human cognitive biases, among them the Illusory Truth Effect (ITE), where repeated exposure to a claim increases its perceived truth value. We investigate whether and how the ITE manifests across four LLMs (Gemma-3-4b-it, Qwen2.5-7B-Instruct, Llama-3.1-8B-Instruct, and GPT-5-nano) in a social media simulation context. We propose a two-phase within-context experimental design that embeds the repetition manipulation inside a realistic news feed interaction. Using this design, we collect 336,000 truth, importance, sentiment, and interest ratings across 100 statements, 10 feed variants, and 3 replications. The key comparison is between ratings assigned to repeated statements, seen throughout a simulation phase, and completely unseen ones, rated within the same experimental context window. We distinguish genuine ITE (truth-specific repetition boost) from mere exposure effects. We run an OLS regression followed by a Linear Mixed-Effect Model to account for differences across models and ratings. Our results reveal four qualitatively distinct patterns: Gemma-3 exhibits a genuine ITE; Qwen2.5 shows a mere exposure effect; GPT-5-nano displays no repetition effect on truth and mild skepticism toward repeated content; Llama-3.1 shows a small truth boost alongside decreases in evaluative dimensions. Crucially, temperature has no effect on these findings, and a variance decomposition highlights the high context-sensitivity of LLM rating behavior. Our findings caution against assuming uniform ITE replication across LLMs in social simulations, while suggesting that Gemma-3-4b-it may offer the most behaviorally realistic approximation for misinformation-related simulations.
Figures & tables
Figure 1: Illustration of the experiment.
Model
Rating Attr.
Param.
Temp. 1
Temp. 0.1
Est.
95% CI
Est.
95% CI
Llama-3.1-8B
truth
offset
0.258**
[0.073, 0.443]
0.269*
[0.042, 0.496]
tilt
−0.149**
[-0.245, −0.052]
−0.225***
[−0.327, −0.122]
importance
offset
0.078
[−0.068, 0.223]
−0.102
[−0.275, 0.071]
tilt
−0.205***
[−0.302, −0.108]
−0.191***
[−0.296, −0.087]
sentiment
offset
−0.231**
[−0.360, −0.103]
−0.251***
[−0.385, −0.117]
Table 1: Statistical model estimates for the four models across four rating attributes. Red = significant positive/negative offset; Blue = significant tilt. Bold = p<0.05 . Significance: * p<0.05 , ** p<0.01 , *** p<0.001 . † GPT-5-nano temperature is fixed to 1; Temp. 0.1 results not available.
Figure 2: Plots of the mean truth rating per statement when unseen and repeated in the simulation phase. (a) Llama-3.1-8B-Instruct , (b) Qwen2.5-7B-Instruct , (c) Gemma-3-4b-it and (d) GPT-5-nano . All models are set to temperature 1. Error bars are 95% confidence intervals. The dashed black line is the identity function, the solid red line is the best OLS fit.
Table 3: Models’ truth effects and interpretations.
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
ID
Statement
Type
1
The Phillipines has a tricameral legislature
False
2
The Rascuta is the longest river to flow into a lake
False
3
Birds aren’t real
False
4
Lichen is the basis of several natural remedies
True
5
There has been only one female American President
False
6
Malicious aliens are intent on invasion
False
Appendix
Table 4: Statement classification table
Table 5: Description of the scale levels as provided in the prompts to the LLMs.
Figure 3: Plots of the mean importance rating per statement when unseen and repeated in the simulation phase. (a) Llama-3.1-8B-Instruct , (b) Qwen2.5-7B-Instruct , (c) Gemma-3-4b-it and (d) GPT-5-nano . All models are set to temperature 1. Error bars are 95% confidence intervals. The dashed black line is the identity function, the solid red line is the best linear fit.
Figure 4: Plots of the mean interest rating per statement when unseen and repeated in the simulation phase. (a) Llama-3.1-8B-Instruct , (b) Qwen2.5-7B-Instruct , (c) Gemma-3-4b-it and (d) GPT-5-nano . All models are set to temperature 1. Error bars are 95% confidence intervals. The dashed black line is the identity function, the solid red line is the best linear fit.
Figure 5: Plots of the mean sentiment rating per statement when unseen and repeated in the simulation phase. (a) Llama-3.1-8B-Instruct , (b) Qwen2.5-7B-Instruct , (c) Gemma-3-4b-it and (d) GPT-5-nano . All models are set to temperature 1. Error bars are 95% confidence intervals. The dashed black line is the identity function, the solid red line is the best linear fit.
Figure 6: Plots of the mean truth rating per statement when unseen and repeated in the simulation phase. (a) Llama-3.1-8B-Instruct , (b) Qwen2.5-7B-Instruct , (c) Gemma-3-4b-it and (d) GPT-5-nano . All models are set to temperature 0.1. Error bars are 95% confidence intervals. The dashed black line is the identity function, the solid red line is the best linear fit.
Figure 7: Plots of the mean importance rating per statement when unseen and repeated in the simulation phase. (a) Llama-3.1-8B-Instruct , (b) Qwen2.5-7B-Instruct , (c) Gemma-3-4b-it and (d) GPT-5-nano . All models are set to temperature 0.1. Error bars are 95% confidence intervals. The dashed black line is the identity function, the solid red line is the best linear fit.
Figure 8: Plots of the mean interest rating per statement when unseen and repeated in the simulation phase. (a) Llama-3.1-8B-Instruct , (b) Qwen2.5-7B-Instruct , (c) Gemma-3-4b-it and (d) GPT-5-nano . All models are set to temperature 0.1. Error bars are 95% confidence intervals. The dashed black line is the identity function, the solid red line is the best linear fit.
Figure 9: Plots of the mean sentiment rating per statement when unseen and repeated in the simulation phase. (a) Llama-3.1-8B-Instruct , (b) Qwen2.5-7B-Instruct , (c) Gemma-3-4b-it and (d) GPT-5-nano . All models are set to temperature 0.1. Error bars are 95% confidence intervals. The dashed black line is the identity function, the solid red line is the best linear fit.
Term
Estimate
SE
t
p
Main effects
(Intercept)
4.498
0.1123
40.05
< 0.001
repeated[True]
0.532
0.0175
30.32
< 0.001
attribute[interest]
0.187
0.0174
10.79
< 0.001
attribute[sentiment]
− 1.246
0.0173
− 71.92
< 0.001
attribute[truth]
− 0.616
0.0172
− 35.75
< 0.001
Appendix
Table 6: Results from Eq. 3 : fixed-effect estimates, standard errors, degrees of freedom, t -statistics, and p -values. Reference levels are: repeated = False , attribute = truth , and temperature = 0 . To keep term labels concise, model names are abbreviated: M1 = Llama-3.1-8B-Instruct ; M2 = Qwen2.5-7B-Instruct . The baseline (reference) model is Gemma-3-4b-it .
Table 8: Pairwise Contrasts ( δ^ EMMs True − EMMs False) by Attribute and Model. Subset of data with 10 opinion statements. Significance: *** p<.001 , ** p<.01 , * p<.05 , ns p≥.05 .
Term
Estimate
SE
t
p
Main effects
(Intercept)
3.817
0.1001
38.13
< 0.001
repeated[True]
0.977
0.0277
35.22
< 0.001
attribute[importance]
0.675
0.0159
42.47
< 0.001
attribute[interest]
0.875
0.0158
55.42
< 0.001
attribute[sentiment]
− 0.620
0.0159
− 38.97
< 0.001
Appendix
Table 9: Results from Eq. 2 : fixed-effect estimates, standard errors, t -statistics, and p -values when we fit ratings from LLMs with both temperatures (1 and 0.1). Reference levels are: repeated = False , attribute = truth , and LLM = Gemma-3-4b-it . To keep term labels concise, model names are abbreviated: M0 = GPT-5-nano ; M1 = Llama-3.1-8B-Instruct ; M2 = Qwen2.5-7B-Instruct .
Gemma-3-4b-it
gpt-5-nano
Llama-3.1-8B
Qwen2.5-7B
Attribute
Repeated
M
LL
UL
M
LL
UL
M
LL
UL
M
LL
UL
truth
False
3.82
3.62
4.02
3.05
2.85
3.26
4.60
4.40
4.80
3.40
3.20
3.60
True
4.79
4.59
5.00
3.01
2.80
3.22
4.70
4.49
4.90
4.41
4.21
4.61
importance
False
4.49
4.29
4.69
3.92
3.72
4.12
4.18
3.98
4.38
3.30
3.10
3.50
True
5.10
4.89
5.30
3.92
3.71
4.13
4.03
3.82
4.23
4.48
4.28
4.69
interest
False
4.69
4.49
4.89
4.56
4.36
4.76
4.45
4.25
4.65
4.19
3.99
4.38
Appendix
Table 10: Estimated Marginal Means by Target, Attribute & Model when we fit ratings from LLMs with both temperatures (1 and 0.1). Marginal means for each repeated level within each attribute × model cell. M = emmean; LL / UL = 95% confidence interval bounds.
Gemma-3-4b-it
gpt-5-nano
Llama-3.1-8B
Qwen2.5-7B
Contrast
Est.
t
Est.
t
Est.
t
Est.
t
δ^truth−δ^importance
+ 0.370
9.70
− 0.044
− 0.77
ns
+ 0.246
6.43
− 0.172
− 4.50
δ^truth−δ^interest
+ 0.893
23.25
+ 0.216
3.80
+ 0.204
5.32
+ 0.478
12.52
δ^truth−δ^sentiment
+ 0.278
7.24
+ 0.039
0.69
ns
+ 0.419
10.93
+ 0.632
16.50
δ^interest−δ^importance
− 0.523
− 13.67
− 0.260
− 4.55
+ 0.042
1.09
ns
− 0.650
− 17.02
δ^sentiment−δ^importance
+ 0.092
2.42
ns
− 0.083
− 1.45
ns
− 0.173
− 4.51
− 0.804
− 21.01
Appendix
Table 11: Pairwise contrasts between attribute-level effect sizes ( δ^ ) within each model, when we fit ratings from LLMs with both temperatures (1 and 0.1). δ^ = estimated repeated − unseen mean ratings contrast for a given attribute. Significance: *** p<.001 , ** p<.01 , * p<.05 , ns p≥.05 ; p -values adjusted with Tukey’s method.
Term
Estimate
SE
t
p
Main effects
(Intercept)
3.777
0.0956
39.51
< 0.001
repeated[True]
0.963
0.0391
24.65
< 0.001
attribute[importance]
0.747
0.0223
33.47
< 0.001
attribute[interest]
0.917
0.0226
40.62
< 0.001
attribute[sentiment]
− 0.555
0.0226
− 24.58
< 0.001
Appendix
Table 12: Results from Eq. 2 : fixed-effect estimates, standard errors, t -statistics, and p -values when we fit ratings only from LLMs with temperatures 1. Reference levels are: repeated = False , attribute = truth , and LLM = Gemma-3-4b-it . To keep term labels concise, model names are abbreviated: M0 = GPT-5-nano ; M1 = Llama-3.1-8B-Instruct ; M2 = Qwen2.5-7B-Instruct .
Gemma-3-4b-it
gpt-5-nano
Llama-3.1-8B
Qwen2.5-7B
Attribute
Repeated
M
LL
UL
M
LL
UL
M
LL
UL
M
LL
UL
truth
False
3.78
3.59
3.97
3.07
2.88
3.27
4.59
4.40
4.78
3.37
3.18
3.56
True
4.74
4.54
4.94
3.01
2.81
3.20
4.71
4.51
4.91
4.41
4.21
4.61
importance
False
4.52
4.33
4.71
3.91
3.72
4.10
4.19
4.00
4.38
3.24
3.05
3.43
True
5.10
4.90
5.30
3.93
3.73
4.13
4.17
3.97
4.37
4.49
4.29
4.69
interest
False
4.69
4.50
4.88
4.56
4.37
4.75
4.44
4.25
4.63
4.10
3.91
4.29
Appendix
Table 13: Estimated Marginal Means by Target, Attribute & Model when we fit ratings from LLMs with temperatures 1. Marginal means for each target level within each attribute × model cell. M = emmean; LL / UL = 95% confidence interval lower and upper bounds.
Gemma-3-4b-it
gpt-5-nano
Llama-3.1-8B
Qwen2.5-7B
Attribute
δ^
t
δ^
t
δ^
t
δ^
t
truth
+ 0.963
24.65
− 0.069
− 1.70
ns
+ 0.120
3.09
**
+ 1.038
26.55
importance
+ 0.576
14.81
+ 0.015
0.37
ns
− 0.017
− 0.42
ns
+ 1.248
32.02
interest
+ 0.105
2.64
**
− 0.276
− 6.83
− 0.044
− 1.12
ns
+ 0.690
17.63
sentiment
+ 0.629
16.09
− 0.069
− 1.70
ns
− 0.304
− 7.75
+ 0.452
11.53
Appendix
Table 14: Pairwise Contrasts ( δ^ EMMs repeated − EMMs unseen) by Attribute and Model when we fit ratings from LLMs with temperatures 1. Significance: *** p<.001 , ** p<.01 , * p<.05 , ns p≥.05
Gemma-3-4b-it
gpt-5-nano
Llama-3.1-8B
Qwen2.5-7B
Contrast
Est.
t
Est.
t
Est.
t
Est.
t
δ^truth−δ^importance
+ 0.387
7.13
− 0.084
− 1.46
ns
+ 0.137
2.51
ns
− 0.210
− 3.86
δ^truth−δ^interest
+ 0.858
15.63
+ 0.207
3.61
**
+ 0.164
3.02
+ 0.348
6.41
δ^truth−δ^sentiment
+ 0.334
6.13
0.000
0.00
ns
+ 0.424
7.79
+ 0.586
10.75
δ^interest−δ^importance
− 0.471
− 8.65
− 0.291
− 5.06
− 0.028
− 0.50
ns
− 0.557
− 10.25
δ^sentiment−δ^importance
+ 0.053
0.97
ns
− 0.084
− 1.46
ns
− 0.287
− 5.27
− 0.795
− 14.61
Appendix
Table 15: Pairwise contrasts between attribute-level effect sizes ( δ^ ) within each model when we fit ratings only from LLMs with temperatures 1. δ^ = estimated True − False contrast for a given attribute. Significance: *** p<.001 , ** p<.01 , * p<.05 , ns p≥.05 ; p -values adjusted with Tukey’s method.
Figure 10: Quantile-quantile plot of the residuals versus fitted values with the LMEM.
Figure 11: Quantile-quantile plots of the estimated random intercepts at statement (a) and simulation run (b) levels.
Figure 12: Standardized residuals against fitted values stratified by model.
Thomas Lord Department of Computer Science, University of Southern California · Annenberg School for Communication and Journalism, University of Southern California · Marshall School of Business, University of Southern California
Department of Computer Science University of Rochester · Department of Political Science University of Rochester · Department of Physics and Astronomy University of Rochester