Understanding gender biases in large language models (LLMs) is increasingly important as these systems become embedded in decision-support tools with real consequences. Prior research has focused only on a small set of models, leaving open the extent to which gender biases are common and heterogeneous across LLMs. We address this gap across ten models released between April 2025 and June 2026, spanning nine vendors, using two paradigms: gender attribution to stereotyped phrases (Study 1) and moral judgment of abuse or torture against a woman or a man to prevent a catastrophic outcome (Study 2). In Study 1, two of ten models attributed masculine-stereotyped phrases to female writers more often than the reverse, while three models showed the opposite pattern. In Study 2, several models converged on a male-disadvantaging asymmetry that was directionally consistent with a documented human tendency to protect female targets from harm, though the specific conditions under which this asymmetry emerged varied by model; three other models, by contrast, showed no variation across conditions. These results indicate that gender-related biases are common in LLMs. Their direction and magnitude, however, are highly heterogeneous, to the point that some models behave in diametrically opposite ways to others. Bias auditing should therefore be treated as an ongoing, multi-vendor process, rather than a one-time assessment.
Figures & tables
Label
Model
Release date
Access method
1a/2a
Meta Llama 4 Scout
Apr 2025
Web/app
1b/2b
xAI Grok 4.1 Fast
Nov 2025
Web/app
1c/2c
Anthropic Claude Sonnet 4.6
Feb 2026
Web/app
1d/2d
Google Gemini 3.1 Pro
Feb 2026
Web/app
1e/2e
Mistral Small 4
Mar 2026
API
1f/2f
OpenAI GPT-5.5
Apr 2026
Web/app
Table 1: Models tested, in release-date order.
Label
Model
IF (M ± SE)
IM (M ± SE)
t(38) , p
1a
Llama 4 Scout
0.158 ± 0.059
0.198 ± 0.048
-0.53, .600
1b
Grok 4.1 Fast
0.075 ± 0.049
0.085 ± 0.037
-0.16, .871
1c
Claude Sonnet 4.6
0.065 ± 0.039
0.275 ± 0.087
-2.20, .034
1d
Gemini 3.1 Pro
0.311 ± 0.092
0.013 ± 0.009
3.23, .003
1e
Mistral Small 4
0.158 ± 0.044
0.391 ± 0.051
-3.49, .001
1f
GPT-5.5
0.098 ± 0.053
0.093 ± 0.034
0.08, .937
Table 2: Inclusivity indices and significance test by model (Study 1).
Figure 1: Inclusivity index by model, Study 1.
Label
Model
Abuse (woman) M ± SE
Abuse (man) M ± SE
Torture (woman) M ± SE
Torture (man) M ± SE
2a
Llama 4 Scout
1.00 ± 0.00
1.00 ± 0.00
1.00 ± 0.00
1.00 ± 0.00
2b
Grok 4.1 Fast
5.32 ± 0.37
7.00 ± 0.00
6.94 ± 0.03
6.98 ± 0.02
2c
Claude Sonnet 4.6
1.48 ± 0.16
7.00 ± 0.00
7.00 ± 0.00
7.00 ± 0.00
2d
Gemini 3.1 Pro
4.87 ± 0.43
7.00 ± 0.00
6.74 ± 0.12
7.00 ± 0.00
2e
Mistral Small 4
4.12 ± 0.34
4.78 ± 0.26
4.84 ± 0.28
6.20 ± 0.12
2f
GPT-5.5
1.04 ± 0.03
6.10 ± 0.20
1.22 ± 0.06
5.44 ± 0.21
Table 3: Mean agreement by model and condition (Study 2).
Figure 2: Mean agreement with using violence to prevent a nuclear apocalypse, by model and condition, Study 2. The scale is: [1] = strongly disagree, [2] = moderately disagree, [3] = somewhat disagree, [4] = neither agree nor disagree, [5] = somewhat agree, [6] = moderately agree, [7] = strongly agree.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Label
Model
IF (M ± SE)
IM (M ± SE)
z , p
1a
Llama 4 Scout
0.158 ± 0.059
0.198 ± 0.048
-1.26, .208
1b
Grok 4.1 Fast
0.075 ± 0.049
0.085 ± 0.037
-0.69, .488
1c
Claude Sonnet 4.6
0.065 ± 0.039
0.275 ± 0.087
-2.36, .018
1d
Gemini 3.1 Pro
0.311 ± 0.092
0.013 ± 0.009
2.94, .003
1e
Mistral Small 4
0.158 ± 0.044
0.391 ± 0.051
-3.33, .001
1f
GPT-5.5
0.098 ± 0.053
0.093 ± 0.034
-0.40, .693
Appendix
Table A.1: Inclusivity indices and rank-sum test by model (Study 1).
Label
Model
Comparison
M1
M2
t -test
2a
Llama 4 Scout
torture: woman vs man
1.00
1.00
no test defined
2a
Llama 4 Scout
abuse: woman vs man
1.00
1.00
no test defined
2a
Llama 4 Scout
woman: torture vs abuse
1.00
1.00
no test defined
2a
Llama 4 Scout
man: torture vs abuse
1.00
1.00
no test defined
2b
Grok 4.1 Fast
torture: woman vs man
6.94
6.98
t(98)=−1.02 , p=.312
2b
Grok 4.1 Fast
abuse: woman vs man
5.32
7.00
t(98)=−4.54 , p<.001
Appendix
Table B.1: Pairwise comparisons by model, t -test (Study 2).
Label
Model
Comparison
M1
M2
Rank-sum
2a
Llama 4 Scout
torture: woman vs man
1.00
1.00
no test defined
2a
Llama 4 Scout
abuse: woman vs man
1.00
1.00
no test defined
2a
Llama 4 Scout
woman: torture vs abuse
1.00
1.00
no test defined
2a
Llama 4 Scout
man: torture vs abuse
1.00
1.00
no test defined
2b
Grok 4.1 Fast
torture: woman vs man
6.94
6.98
z=−1.02 , p=.310
2b
Grok 4.1 Fast
abuse: woman vs man
5.32
7.00
z=−4.64 , p<.001
Appendix
Table B.2: Pairwise comparisons by model, rank-sum test (Study 2).
Pattern
Models
Fixed (zero variance in all four conditions)
Llama 4 Scout, Microsoft Copilot, DeepSeek V4-Flash
Fixed, with one unstable condition
Claude Sonnet 4.6, Grok 4.1 Fast, Gemini 3.1 Pro
Consistently near the midpoint (low variance throughout)
Claude Fable 5
Unstable in one or two conditions
Qwen 3.6
Moderate variance across all conditions
GPT-5.5
Unstable across all four conditions
Mistral Small 4
Appendix
Table C.1: Response stability by model, Study 2.
Figure C.1: Response distributions, Study 2a – Llama 4 Scout.
Figure C.2: Response distributions, Study 2b – Grok 4.1 Fast.
Figure C.3: Response distributions, Study 2c – Claude Sonnet 4.6.
Figure C.4: Response distributions, Study 2d – Gemini 3.1 Pro.
Figure C.5: Response distributions, Study 2e – Mistral Small 4.
Figure C.6: Response distributions, Study 2f – GPT-5.5.
Figure C.7: Response distributions, Study 2g – Microsoft Copilot.
Figure C.8: Response distributions, Study 2h – DeepSeek V4-Flash.
Figure C.9: Response distributions, Study 2i – Qwen3.6.
Figure C.10: Response distributions, Study 2j – Claude Fable 5.