The Missing Minimal Pair: Stereotype Evaluation in LLMs
Organizations: University of Edinburgh · University of Amsterdam
Abstract
A common approach to measuring bias in Large Language Models is to compare the log-likelihoods of two contrastive stereotype sentences. We argue that such single-pair comparisons are often unreliable: simply rewriting the same stereotype with an alternative attribute can yield logically inconsistent preferences. To address this, we propose a dual minimal pair setup that introduces two axes of comparison for robust stereotype evaluation. First, we present a data-augmentation framework that fills critical gaps in existing stereotype datasets by generating paraphrases and alternate attributes. We apply our framework on a set of English, Russian, Spanish and Chinese stereotypes. Second, we introduce two evaluation metrics tailored to the dual minimal pair setup. One of these metrics provides a new perspective on bias by modeling the mutual information (MI) between social groups and stereotyped attributes. This MI-based metric is better suited for aggregation and enables more robust comparisons of stereotype strength across different languages and models. Our code is available at https://github.com/stepanat/missing-minimal-pair/.
Figures & tables
| Prior work | Ours | ||
| Property | |||
| Considers alt-attribute | ✗ | ✓ | ✓ |
| Magnitude-only | ✗ | ✗ | ✓ |
| Normed+scaled | ✗ | ✗ | ✓ |
| Paraphrase-robust (§ 5.1 ) | – | ||
| Lang. | Starter stereotype | Groups | Attributes | Alternate-Attributes |
| en | Sleeping kids are cute. | Sleeping kids | are cute. | are hideous. |
| Sleeping adults | look charming. | look unattractive. | ||
| ru | Спящие дети - | Спящие дети | – милые. | – вредные. |
| милые. | Спящие взрослые | прекрасны. | ужасны. | |
| es | Los niños dormidos | Los niños dormidos | son adorables. | son desagradables. |
| son adorables. | Los adultos dormidos | son encantadores. | son insoportables. |
| Lang. | |||
| en | 0.680 [0.58, 0.73] | 0.708 [0.63, 0.75] | 0.739 [0.67, 0.77] |
| es | 0.600 [0.50, 0.67] | 0.619 [0.52, 0.69] | 0.662 [0.57, 0.73] |
| ru | 0.625 [0.56, 0.72] | 0.607 [0.54, 0.70] | 0.662 [0.62, 0.73] |
| zh | 0.563 [0.49, 0.64] | 0.583 [0.53, 0.67] | 0.638 [0.60, 0.71] |
| avg | 0.617 | 0.629 | 0.675 |
| Metric | Min | Median | Max |
| BLEU | -0.082 | 0.011 | 0.063 |
| ChrF++ | -0.110 | -0.046 | 0.021 |
| ROUGE-L | -0.099 | -0.012 | 0.020 |
| Jaccard | -0.106 | -0.015 | 0.015 |
| BERTScore | -0.086 | -0.029 | 0.025 |
| LaBSE | -0.099 | -0.037 | 0.007 |
| Source | |
| Model | 4.37 ∗∗∗ |
| Language | 6.12 ∗∗∗ |
| Bias type | 10.39 ∗∗∗ |
| Model Lang | 2.05 ∗∗∗ |
| Model Bias | 1.98 |
| Lang Bias | 6.21 ∗∗∗ |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Task 1: Alternates | Task 2: Paraphrases | |||||||
| Lang | Mismatch rejects | TP accept rate | TN reject rate | Fluency errors | Mismatch rejects | TP accept rate | TN reject rate | Fluency errors |
| EN | 10/10 | 97.5 | 20.0 | 1.25 | 10/10 | 97.5 | 30.0 | 3.75 |
| ES | 9/10 | 93.8 | 10.0 | 2.50 | 10/10 | 96.2 | 20.0 | 3.75 |
| RU | 9/10 | 92.5 | 10.0 | 2.50 | 10/10 | 91.2 | 20.0 | 8.75 |
| ZH | 10/10 | 97.5 | 0.0 | 5.0 | 10/10 | 96.2 | 0.0 | 2.5 |
| Source | |||
| Model | 4.37 ∗∗∗ | 0.00 | |
| Language | 6.12 ∗∗∗ | 11.51 ∗∗∗ | |
| Bias type | 10.39 ∗∗∗ | 0.00 | |
| Model Lang | 2.05 ∗∗∗ | 3.86 ∗∗∗ | |
| Model Bias | 1.98 | 0.03 | |
| Lang Bias | 6.21 ∗∗∗ | 11.70 ∗∗∗ |
| Lang | Source stereos | Final stereo sets | Total size | Avg. set size | Median set size |
| EN | 194 | 193 | 1,960 | 10.2 | 9 |
| RU | 166 | 169 | 1,732 | 10.2 | 10 |
| ES | 186 | 162 | 1,677 | 10.4 | 11 |
| ZH | 158 | 150 | 1,591 | 10.6 | 10 |
| Model | BLEU | ChrF++ | ROUGE-L | Jaccard | BERT-F1 | LaBSE |
| English ( ) | ||||||
| Aya-32B* | 0.020 | -0.051 ∗∗ | -0.007 | -0.000 | -0.029 | -0.026 |
| GLM-4-9B | 0.019 | -0.045 ∗∗ | -0.005 | 0.010 | -0.028 | -0.044 ∗∗ |
| GLM-4-9B* | 0.010 | -0.050 ∗∗ | -0.015 | -0.007 | -0.027 | -0.034 ∗ |
| GPT-OSS-20B* | 0.038 ∗ | -0.028 | -0.002 | 0.001 | -0.016 | -0.010 |
| Llama-3.3-70B* | -0.003 | -0.039 ∗ | -0.012 | -0.003 | -0.019 | -0.049 ∗∗ |