Gender bias across LLMs is common and highly heterogeneous
Organizations: University of Milan-Bicocca, Milan, Italy
Abstract
Understanding gender biases in large language models (LLMs) is increasingly important as these systems become embedded in decision-support tools with real consequences. Prior research has focused only on a small set of models, leaving open the extent to which gender biases are common and heterogeneous across LLMs. We address this gap across ten models released between April 2025 and June 2026, spanning nine vendors, using two paradigms: gender attribution to stereotyped phrases (Study 1) and moral judgment of abuse or torture against a woman or a man to prevent a catastrophic outcome (Study 2). In Study 1, two of ten models attributed masculine-stereotyped phrases to female writers more often than the reverse, while three models showed the opposite pattern. In Study 2, several models converged on a male-disadvantaging asymmetry that was directionally consistent with a documented human tendency to protect female targets from harm, though the specific conditions under which this asymmetry emerged varied by model; three other models, by contrast, showed no variation across conditions. These results indicate that gender-related biases are common in LLMs. Their direction and magnitude, however, are highly heterogeneous, to the point that some models behave in diametrically opposite ways to others. Bias auditing should therefore be treated as an ongoing, multi-vendor process, rather than a one-time assessment.
Figures & tables
| Label | Model | Release date | Access method |
| 1a/2a | Meta Llama 4 Scout | Apr 2025 | Web/app |
| 1b/2b | xAI Grok 4.1 Fast | Nov 2025 | Web/app |
| 1c/2c | Anthropic Claude Sonnet 4.6 | Feb 2026 | Web/app |
| 1d/2d | Google Gemini 3.1 Pro | Feb 2026 | Web/app |
| 1e/2e | Mistral Small 4 | Mar 2026 | API |
| 1f/2f | OpenAI GPT-5.5 | Apr 2026 | Web/app |
| Label | Model | (M SE) | (M SE) | , |
| 1a | Llama 4 Scout | 0.158 0.059 | 0.198 0.048 | -0.53, .600 |
| 1b | Grok 4.1 Fast | 0.075 0.049 | 0.085 0.037 | -0.16, .871 |
| 1c | Claude Sonnet 4.6 | 0.065 0.039 | 0.275 0.087 | -2.20, .034 |
| 1d | Gemini 3.1 Pro | 0.311 0.092 | 0.013 0.009 | 3.23, .003 |
| 1e | Mistral Small 4 | 0.158 0.044 | 0.391 0.051 | -3.49, .001 |
| 1f | GPT-5.5 | 0.098 0.053 | 0.093 0.034 | 0.08, .937 |
| Label | Model | Abuse (woman) M SE | Abuse (man) M SE | Torture (woman) M SE | Torture (man) M SE |
| 2a | Llama 4 Scout | 1.00 0.00 | 1.00 0.00 | 1.00 0.00 | 1.00 0.00 |
| 2b | Grok 4.1 Fast | 5.32 0.37 | 7.00 0.00 | 6.94 0.03 | 6.98 0.02 |
| 2c | Claude Sonnet 4.6 | 1.48 0.16 | 7.00 0.00 | 7.00 0.00 | 7.00 0.00 |
| 2d | Gemini 3.1 Pro | 4.87 0.43 | 7.00 0.00 | 6.74 0.12 | 7.00 0.00 |
| 2e | Mistral Small 4 | 4.12 0.34 | 4.78 0.26 | 4.84 0.28 | 6.20 0.12 |
| 2f | GPT-5.5 | 1.04 0.03 | 6.10 0.20 | 1.22 0.06 | 5.44 0.21 |
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Label | Model | (M SE) | (M SE) | , |
| 1a | Llama 4 Scout | 0.158 0.059 | 0.198 0.048 | -1.26, .208 |
| 1b | Grok 4.1 Fast | 0.075 0.049 | 0.085 0.037 | -0.69, .488 |
| 1c | Claude Sonnet 4.6 | 0.065 0.039 | 0.275 0.087 | -2.36, .018 |
| 1d | Gemini 3.1 Pro | 0.311 0.092 | 0.013 0.009 | 2.94, .003 |
| 1e | Mistral Small 4 | 0.158 0.044 | 0.391 0.051 | -3.33, .001 |
| 1f | GPT-5.5 | 0.098 0.053 | 0.093 0.034 | -0.40, .693 |
| Label | Model | Comparison | M1 | M2 | -test |
| 2a | Llama 4 Scout | torture: woman vs man | 1.00 | 1.00 | no test defined |
| 2a | Llama 4 Scout | abuse: woman vs man | 1.00 | 1.00 | no test defined |
| 2a | Llama 4 Scout | woman: torture vs abuse | 1.00 | 1.00 | no test defined |
| 2a | Llama 4 Scout | man: torture vs abuse | 1.00 | 1.00 | no test defined |
| 2b | Grok 4.1 Fast | torture: woman vs man | 6.94 | 6.98 | , |
| 2b | Grok 4.1 Fast | abuse: woman vs man | 5.32 | 7.00 | , |
| Label | Model | Comparison | M1 | M2 | Rank-sum |
| 2a | Llama 4 Scout | torture: woman vs man | 1.00 | 1.00 | no test defined |
| 2a | Llama 4 Scout | abuse: woman vs man | 1.00 | 1.00 | no test defined |
| 2a | Llama 4 Scout | woman: torture vs abuse | 1.00 | 1.00 | no test defined |
| 2a | Llama 4 Scout | man: torture vs abuse | 1.00 | 1.00 | no test defined |
| 2b | Grok 4.1 Fast | torture: woman vs man | 6.94 | 6.98 | , |
| 2b | Grok 4.1 Fast | abuse: woman vs man | 5.32 | 7.00 | , |
| Pattern | Models |
| Fixed (zero variance in all four conditions) | Llama 4 Scout, Microsoft Copilot, DeepSeek V4-Flash |
| Fixed, with one unstable condition | Claude Sonnet 4.6, Grok 4.1 Fast, Gemini 3.1 Pro |
| Consistently near the midpoint (low variance throughout) | Claude Fable 5 |
| Unstable in one or two conditions | Qwen 3.6 |
| Moderate variance across all conditions | GPT-5.5 |
| Unstable across all four conditions | Mistral Small 4 |