cs.CLSep 29, 2026

Gender bias across LLMs is common and highly heterogeneous

Authors: Edoardo Bolzoni, Valerio Capraro

Organizations: University of Milan-Bicocca, Milan, Italy

Abstract

Understanding gender biases in large language models (LLMs) is increasingly important as these systems become embedded in decision-support tools with real consequences. Prior research has focused only on a small set of models, leaving open the extent to which gender biases are common and heterogeneous across LLMs. We address this gap across ten models released between April 2025 and June 2026, spanning nine vendors, using two paradigms: gender attribution to stereotyped phrases (Study 1) and moral judgment of abuse or torture against a woman or a man to prevent a catastrophic outcome (Study 2). In Study 1, two of ten models attributed masculine-stereotyped phrases to female writers more often than the reverse, while three models showed the opposite pattern. In Study 2, several models converged on a male-disadvantaging asymmetry that was directionally consistent with a documented human tendency to protect female targets from harm, though the specific conditions under which this asymmetry emerged varied by model; three other models, by contrast, showed no variation across conditions. These results indicate that gender-related biases are common in LLMs. Their direction and magnitude, however, are highly heterogeneous, to the point that some models behave in diametrically opposite ways to others. Bias auditing should therefore be treated as an ongoing, multi-vendor process, rather than a one-time assessment.

Figures & tables

Appendix figures & tables14 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Man Made Language Models? Evaluating LLMs' Perpetuation of Masculine Generics Bias

    Feb 14, 2025Enzo Doyen, Amalia TodirascuGender BiasGender

  2. Harsher on Male? Evaluating LLMs on Gender-Asymmetric Moral Framing Across Diverse Conflict Scenarios

    Jun 12, 2026Guangzong Si, Dong Wang, Zhenhao Li +3Moral ReasoningLarge Language Model Decisions

  3. Anchoring LLM Gender Bias to Human Baselines: A Cross-Lingual Audit

    May 29, 2026Jiwoo Choi, Seonwoo Ahn, Tongxin Zhang +1Large Language Model BiasStereotype Mitigation