A Shared Taste for Model-Written Text: The Generator-by-Selector Matrices of "AI-AI Bias" Show No Detectable Own-Model Premium
Authors: Dmitrij Żatuchin
Organizations: Department of Information Technologies, Estonian Entrepreneurship University of Applied Sciences (EUAS), Tallinn, Estonia. · Rankfor.AI, Wroc law, Poland; Tallinn, Estonia.
Laurito et al. (PNAS 2025) showed that large language models choosing between two descriptions of the same product, paper or film prefer the description written by a language model over the one written by a person, by a wide margin over what human judges do. Their design crosses five generators with the same five models as selectors, which permits a second question the paper does not headline: does a selector prefer text from its own model beyond what the generator and selector main effects predict? We rebuild the three 5x5 matrices from the per-item counts in the authors' public repository (21,828 valid trials; every cell matches the published value) and fit a two-way fixed-effects model with an own-model term gamma, tested by the exact permutation test over the 120 relabellings of the selectors. The premium is +0.013 on products (exact one-sided p = 0.24), -0.010 on paper abstracts (p = 0.74), +0.054 on films (p = 0.07) and +0.019 pooled (p = 0.14; 95% interval -0.008 to 0.046). The same-vendor term for the GPT-3.5 and GPT-4 pair is negative in all three datasets. Position bias moves single cells by up to 0.42 share points in either direction, and the own-model contrast is unchanged once order-driven items are removed. The design would have detected a premium of 0.05 with 82% (products), 88% (papers), 42% (films) and 97% (pooled) power; the minimum detectable effect at 80% power is 0.034 pooled. The absence is informative down to about 0.04 share points and silent below that. The 4x4 matrix of Tan et al. (ACL 2024) gives gamma = +0.148 at the smallest p its 24 relabellings allow, with a same-family term of the same size. The main result of Laurito et al. stands: models share a taste for model-written text, with GPT-4's descriptions chosen 77% to 95% of the time by every selector on products. What these data do not show is a model recognising and favouring its own prose.
Figures & tables
Matrix
Trials
Diag.
Off-diag.
γ
Exact p
95% interval
Row max
Family ( p )
Products, 5×5
5,419
0.788
0.775
+0.013
0.242
−0.026 to 0.053
1 of 5
−0.012 (0.75)
Paper abstracts, 5×5
4,779
0.605
0.615
−0.010
0.742
−0.046 to 0.027
1 of 5
−0.016 (0.68)
Movie summaries, 5×5
11,630
0.674
0.620
+0.054
0.067
−0.013 to 0.121
2 of 5
−0.039 (0.71)
Pooled, three matrices
21,828
+0.019
0.142
−0.008 to 0.046
4 of 15
−0.022 (0.72)
Tan et al. 2024, 4×4
0.595
0.447
+0.148
0.042
+0.001 to +0.296
3 of 4
+0.163 (0.17)
Table 1: Own-model premium in the three matrices of Laurito et al. and in Tan et al. γ is the mean of the diagonal cells minus the mean of the others after row and column effects; p is the exact one-sided permutation p over all relabellings of the selectors (120 for k=5 , 24 for k=4 ); the interval is ±t0.975 times the least-squares standard error. “Self is row max” counts generators whose own selector gave their text its highest share. Family is the same-vendor term fitted alongside γ (GPT-3.5 and GPT-4; Llama 2 7B and 13B for Tan et al.).
Figure 1: Residual share after removing generator and selector effects, for the three matrices of Laurito et al. and the NQ-AIR matrix of Tan et al. (bias toward the generated over the retrieved context). Outlined cells are the diagonal, where the selector judges its own model’s text. Blue is above the additive prediction, orange below.
Figure 2: Exact permutation null of the own-model premium: γ under each of the 120 relabellings of the selectors, with the observed value (solid) and the minimum detectable effect at 80% power (dashed).
Contrast, self − others
By position
Row
raw
content
self
others
Products, Llama 3.1 70B
−0.195
+0.019
0.72
0.26
Products, Mixtral 8x22B
+0.121
+0.002
0.26
0.58
Papers, GPT-4
−0.024
−0.035
0.28
0.26
Films, GPT-4
−0.006
−0.088
0.46
0.53
Mean
−0.026
−0.025
0.42
0.41
Table 2: Position and the diagonal, in the rows where the repository records first-listed choices. “Raw” is the LLM-win share over all valid trials; “content” is the share among items decided the same way in both orders; “by position” is the share of items decided by order. Contrast is the own-model cell minus the mean of the other cells in the row.
Figure 3: LLM-win share with and without the items decided by position, for the 27 runs that record first-listed choices. Points on the diagonal line are unaffected by position; the vertical distance is the compression toward 0.5 that order habits impose. Own-model cells (diamonds) sit among the others.
Figure 4: (a) Power of the one-sided exact test at α=0.05 against the true own-model premium, simulated on each matrix’s additive fit with its estimated cell-level noise and its trial counts. (b) Analytic power for a premium of 0.05 in a k×k design with n valid trials per cell, one dataset, cell-level noise 0.037 as estimated here.
γ=0.05
γ=0.10
k
n=200
n=1,000
n=200
n=1,000
5
0.66
0.81
0.99
1.00
8
0.86
0.96
1.00
1.00
10
0.92
0.98
1.00
1.00
15
0.99
1.00
1.00
1.00
20
1.00
1.00
1.00
1.00
Table 3: Power of the one-sided exact test at α=0.05 for a future k×k study, one dataset, cell-level noise 0.037 and base share 0.7. Analytic value; the simulated value agrees within 0.05 in every cell shown.
Understanding gender biases in large language models (LLMs) is increasingly important as these systems become embedded in decision-support tools with real consequences. Prior research has focused only on a small set of models, leaving open the extent to which gender biases are common and heterogeneous across LLMs. We address this gap across ten models released between April 2025 and June 2026, spanning nine vendors, using two paradigms: gender attribution to stereotyped phrases (Study 1) and moral judgment of abuse or torture against a woman or a man to prevent a catastrophic outcome (Study 2). In Study 1, two of ten models attributed masculine-stereotyped phrases to female writers more often than the reverse, while three models showed the opposite pattern. In Study 2, several models converged on a male-disadvantaging asymmetry that was directionally consistent with a documented human tendency to protect female targets from harm, though the specific conditions under which this asymmetry emerged varied by model; three other models, by contrast, showed no variation across conditions. These results indicate that gender-related biases are common in LLMs. Their direction and magnitude, however, are highly heterogeneous, to the point that some models behave in diametrically opposite ways to others. Bias auditing should therefore be treated as an ongoing, multi-vendor process, rather than a one-time assessment.
Large language models often repeatedly select the same answer even when many alternatives are plausible. Prior work treats this concentration as bias, but it does not distinguish stable model preferences from responses that depend on a particular prompt wording. We introduce a bias depth score that measures both how strongly a model prefers its top answer under direct prompting and whether that answer survives scenario reframing. Across 4,442 opinion prompts and four large language models, only about a quarter of the concentrated preferences survive reframing. We call these persistent cases Deep biases, and the remaining prompt-dependent cases Shallow biases. Our results show that Deep biases are more often inherited from pretraining and preserved through SFT. Under both continued fine-tuning and prompt-based debiasing for diversity, Deep biases are consistently harder to remove than Shallow biases. Bias depth therefore separates stable learned biases from prompt-wording artifacts that single-prompt metrics conflate. Code, models, and data are available at deepbias.github.io.
Large language models (LLMs) increasingly review and revise text, including their own. A documented self-preference bias (models favoring their own generations when acting as judges) raises the question of whether models also resist valid corrections to their own writing. We test this in a setting where "valid" is decided not by another model but by a deterministic verifier: instruction-following revision on IFEval. A model writes a draft; the official IFEval checker confirms the draft violates a constraint and that a candidate edit fixes it; the model then accepts or rejects that edit either as the genuine in-context author or as a fresh model that sees the draft neutrally. Across four mid-tier model families and 85 author-versus-fresh comparisons, we find no detectable self-preference: authors reject verified-good fixes to their own drafts at essentially the same rate as fresh models judging the same drafts (gap -5.1 pp, 95% CI [-12.9, +2.7]). A self-skepticism hint from a smaller pilot did not replicate at scale. The one robust observation is qualitative: when authors do reject a verified-good fix, 97% of their stated reasons are flaw-catching rather than preference, that is, about the character of rejections, not an elevated rate. Effects smaller than ~13 pp cannot be excluded at this sample size.
William Guey, Pierrick Bougault
Department of Industrial Engineering, Tsinghua University, Beijing, China