cs.CLOct 6, 2026

The Missing Minimal Pair: Stereotype Evaluation in LLMs

Authors: Nataliya Stepanova, Ivan Titov, Emily Allaway, Björn Ross

Organizations: University of Edinburgh · University of Amsterdam

Abstract

A common approach to measuring bias in Large Language Models is to compare the log-likelihoods of two contrastive stereotype sentences. We argue that such single-pair comparisons are often unreliable: simply rewriting the same stereotype with an alternative attribute can yield logically inconsistent preferences. To address this, we propose a dual minimal pair setup that introduces two axes of comparison for robust stereotype evaluation. First, we present a data-augmentation framework that fills critical gaps in existing stereotype datasets by generating paraphrases and alternate attributes. We apply our framework on a set of English, Russian, Spanish and Chinese stereotypes. Second, we introduce two evaluation metrics tailored to the dual minimal pair setup. One of these metrics provides a new perspective on bias by modeling the mutual information (MI) between social groups and stereotyped attributes. This MI-based metric is better suited for aggregation and enables more robust comparisons of stereotype strength across different languages and models. Our code is available at https://github.com/stepanat/missing-minimal-pair/.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs

    May 11, 2026Pierre Le Jeune, Étienne Duchesne, Weixuan Xiao +4Stereotype MitigationBiases

  2. Anchoring LLM Gender Bias to Human Baselines: A Cross-Lingual Audit

    May 29, 2026Jiwoo Choi, Seonwoo Ahn, Tongxin Zhang +1Large Language Model BiasStereotype Mitigation

  3. Scalable and Culturally Specific Stereotype Dataset Construction via Human-LLM Collaboration

    Jul 8, 2026Weicheng Ma, John Guerrerio, Soroush VosoughiStereotype MitigationMultilingual Benchmark