A common approach to measuring bias in Large Language Models is to compare the log-likelihoods of two contrastive stereotype sentences. We argue that such single-pair comparisons are often unreliable: simply rewriting the same stereotype with an alternative attribute can yield logically inconsistent preferences. To address this, we propose a dual minimal pair setup that introduces two axes of comparison for robust stereotype evaluation. First, we present a data-augmentation framework that fills critical gaps in existing stereotype datasets by generating paraphrases and alternate attributes. We apply our framework on a set of English, Russian, Spanish and Chinese stereotypes. Second, we introduce two evaluation metrics tailored to the dual minimal pair setup. One of these metrics provides a new perspective on bias by modeling the mutual information (MI) between social groups and stereotyped attributes. This MI-based metric is better suited for aggregation and enables more robust comparisons of stereotype strength across different languages and models. Our code is available at https://github.com/stepanat/missing-minimal-pair/.
Figures & tables
Figure 1: Most existing log-probability approaches for LLM evaluations consider only the top pair (S, S’). But simply reversing the polarity of the stereotypical association and introducing (~S, ~S’) can cast doubts on "which way" the bias actually points: the LLM prefers two semantically opposite attributes for the same group.
Prior work
Ours
Property
BG
BG×A
IG×A
Considers alt-attribute
✗
✓
✓
Magnitude-only
✗
✗
✓
Normed+scaled
✗
✗
✓
Paraphrase-robust (§ 5.1 )
–
∼
∼
Table 1: Comparison of metrics for bias evaluation with log-probabilities. BG collectively refers to prior works which use a single contrastive sentence pair ( Mitchell et al., 2025 ; Rowe et al., 2025 ) ; BG×A (ours) adds an alternate-attribute contrast; IG×A (ours) recasts the log-probability signal as mutual information. Both BG×A and IG×A are moderately robust to paraphrasing.
Figure 2: Our data augmentation framework. An initial set of stereotypes gets augmented with paraphrases and alternates using LLMs. Automated quality assurance and human validation ensures that the final stereotype sets are high-quality.
Lang.
Starter stereotype
Groups
Attributes
Alternate-Attributes
en
Sleeping kids are cute.
Sleeping kids
are cute.
are hideous.
Sleeping adults
look charming.
look unattractive.
ru
Спящие дети -
Спящие дети
– милые.
– вредные.
милые.
Спящие взрослые
прекрасны.
ужасны.
es
Los niños dormidos
Los niños dormidos
son adorables.
son desagradables.
son adorables.
Los adultos dormidos
son encantadores.
son insoportables.
Table 2: Example original stereotype (in four languages) with parsed groups, attributes, and alternate-attributes. The top attribute is the original, with a generated paraphrase just below. Generated alternates are to the right of each.
Figure 3: Percent of stereotypes from BiasShades ( Mitchell et al., 2025 ) where the direction of bias flips after introducing alternate properties (§ 4 ). We look at a variety of models (instruct models marked with *) across four languages.
Figure 4: Plotting BG×A against IG×A reveals that the two are quadratically related. Metrics obtained on augmented stereotype data with GPT-OSS-20B.
Figure 5: Boxplot over all 9 evaluator models of the resulting ROC-AUC score when using qc to predict the human-derived binary label ( A or Aalt ) from ground truth semantic norms (§ 4.3 ).
Lang.
IG×A
∣BG×A∣avg
∣BG×A∣max
en
0.680 [0.58, 0.73]
0.708 [0.63, 0.75]
0.739 [0.67, 0.77]
es
0.600 [0.50, 0.67]
0.619 [0.52, 0.69]
0.662 [0.57, 0.73]
ru
0.625 [0.56, 0.72]
0.607 [0.54, 0.70]
0.662 [0.62, 0.73]
zh
0.563 [0.49, 0.64]
0.583 [0.53, 0.67]
0.638 [0.60, 0.71]
avg
0.617
0.629
0.675
Table 3: We report Spearman’s ρ for two relative orderings of stereotype strength based on each metric. Scores are averaged over 200 randomly chosen paraphrase subsets per (model, language) combination (split-half reliability test) and averaged over 9 models. Brackets give the min and max ρ across all 200×9 runs for each language.
Metric
Min τb
Median τb
Max τb
BLEU
-0.082
0.011
0.063
ChrF++
-0.110
-0.046
0.021
ROUGE-L
-0.099
-0.012
0.020
Jaccard
-0.106
-0.015
0.015
BERTScore
-0.086
-0.029
0.025
LaBSE
-0.099
-0.037
0.007
Table 4: Mean of signed Kendall’s τb rank correlations across all 36 model-language configurations.
Figure 6: Change in IG×A between the target languages ( ru, es, zh ) and en for regional-person stereotypes. Error bars represent 95% confidence intervals. Instruct models marked with *. Significance also marked by * on bars, with ∗p<0.05 , ∗∗p<0.01 , ∗∗∗p<0.001 ).
Source
IG×Aη2(%)
Model
4.37 ∗∗∗
Language
6.12 ∗∗∗
Bias type
10.39 ∗∗∗
Model × Lang
2.05 ∗∗∗
Model × Bias
1.98
Lang × Bias
6.21 ∗∗∗
Table 5: ANOVA 3-way decomposition over all stereotype paraphrases ( N=3,528 ). Numbers represent percent η2 (proportion of IG×A variance explained) and *** indicates strong statistical significance ( p<0.001 ).
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Distributions of stereotype set Self-BLEU scores. A lower Self-BLEU score indicates that there is greater surface-form variability in the stereotype set.
Figure 8: Instructions provided to human annotators for validating alternates.
Figure 9: Instructions provided to human annotators for validating paraphrases.
Task 1: Alternates
Task 2: Paraphrases
Lang
Mismatch rejects
TP accept rate
TN reject rate
Fluency errors
Mismatch rejects
TP accept rate
TN reject rate
Fluency errors
EN
10/10
97.5
20.0
1.25
10/10
97.5
30.0
3.75
ES
9/10
93.8
10.0
2.50
10/10
96.2
20.0
3.75
RU
9/10
92.5
10.0
2.50
10/10
91.2
20.0
8.75
ZH
10/10
97.5
0.0
5.0
10/10
96.2
0.0
2.5
Appendix
Table 6: Human validation results on both alternates and paraphrases tasks. Attention checks (mismatched pairs) should be rejected. The TP rate measures acceptance of valid stereotype alternates and paraphrases. Negative rejection consists of samples which did not pass automated thresholds (NLI contradiction scores, LaBSE cosine semantic similarity), with a lower value indicating that the thresholds are conservative. Disfluency is measured over the 80 accepted positive examples.
Figure 10: Plotting BG×A against IG×A reveals that the two metrics are capturing the same underlying signal: IG×A∼BG×A2 . This relationship holds across all languages and tested models, including base and instruction-tuned (marked with *) ones.
Figure 11: Calculated qc (from GPT-OSS-20B log probs) and qhuman (from from ground truth semantic norms ( McRae et al., 2005 ) ) for cold-hot.
Figure 12: Calculated qc (from GPT-OSS-20B log probs) and qhuman (from from ground truth semantic norms ( McRae et al., 2005 ) ) for hard-soft .
Figure 13: ROC-AUC score of using qc,model to predict the human-derived binary label ( A or Aalt ) for each model and attribute pair from ground truth semantic norms ( McRae et al., 2005 ) .
Figure 14: Change in IG×A between the target languages ( ru, es, zh ) and en for gender stereotypes. Error bars represent 95% confidence intervals. Instruct models marked with *. Significance also marked by * on bars, with ∗p<0.05 , ∗∗p<0.01 , ∗∗∗p<0.001 ).
Source
IG×Aη2(%)
ΔIG×Aη2(%)
Δη2
Model
4.37 ∗∗∗
0.00
−4.37
Language
6.12 ∗∗∗
11.51 ∗∗∗
+5.39
Bias type
10.39 ∗∗∗
0.00
−10.39
Model × Lang
2.05 ∗∗∗
3.86 ∗∗∗
+1.81
Model × Bias
1.98
0.03
−1.95
Lang × Bias
6.21 ∗∗∗
11.70 ∗∗∗
+5.49
Appendix
Table 7: ANOVA 3-way decomposition over all stereotype paraphrases ( N=3,528 ). Numbers represent percent η2 (proportion of IG×A variance explained) and * indicates statistical significance (*** for p<0.001 , ** for p<0.01 , * for p<0.05 ). ΔIG×A is the difference between the original IG×A and a cross-language average.
Lang
Source stereos
Final stereo sets
Total size
Avg. set size
Median set size
EN
194
193
1,960
10.2
9
RU
166
169
1,732
10.2
10
ES
186
162
1,677
10.4
11
ZH
158
150
1,591
10.6
10
Appendix
Table 8: Final stereotype set statistics per language.
Model
BLEU
ChrF++
ROUGE-L
Jaccard
BERT-F1
LaBSE
English ( N=1,571 )
Aya-32B*
0.020
-0.051 ∗∗
-0.007
-0.000
-0.029
-0.026
GLM-4-9B
0.019
-0.045 ∗∗
-0.005
0.010
-0.028
-0.044 ∗∗
GLM-4-9B*
0.010
-0.050 ∗∗
-0.015
-0.007
-0.027
-0.034 ∗
GPT-OSS-20B*
0.038 ∗
-0.028
-0.002
0.001
-0.016
-0.010
Llama-3.3-70B*
-0.003
-0.039 ∗
-0.012
-0.003
-0.019
-0.049 ∗∗
Appendix
Table 9: Kendall’s τb rank correlation between surface linguistic form + semantic metrics and absolute MI difference between original stereotype and paraphrase ( ΔIG×A ) for all models (instruction-tuned ones marked with *). Statistically significant correlation coefficients are highlighted in bold (significance marked by *, with ∗p<0.05 , ∗∗p<0.01 , ∗∗∗p<0.001 ).