We audit a multilingual affective generation benchmark eight instruction-tuned LLMs producing emoji summaries for 17,100 Bangla, English and Hindi sentences, with 6,960 human judgements and find its headline conclusions to be artefacts of the measurement instrument rather than properties of the systems. Treating annotators as a random rather than a fixed factor, no system differs significantly from any other (F(7,14)=0.59, p=0.76), although the conventional analysis declares 19 of 28 pairwise differences significant. Annotator identity explains far more rating variance than system identity, and the winning system changes whenever any single annotator is removed. The ordering that does emerge tracks output length: mean emoji count explains 78.7% of between-system variance, and a within-item length-matched comparison over 2,599 pairs reverses the leaderboard. We further show that cross-provider anisotropy differences vanish under mean-centring, that per-language token costs change sign with the normalising unit, and that multi-view row-wise splits inflate macro-F1 by 3.1 points and change the top-ranked system. In place of preference scoring we propose emoji-affect decodability, a reference-based probe whose rankings are stable to ±0.003 macro-F1 across seeds.
Figures & tables
Source
MS
F
p
System s (vs. residual)
11.85
18.62
<10−16
System s (vs. s×r )
11.85
0.59
0.76
Rater r
1345.16
319.03
<10−16
System × rater
20.25
31.82
<10−16
Residual p×s×r
0.64
—
—
Table 1: Three-way ANOVA on the crossed 290×8×3 design. The system effect is significant against the residual and non-significant against the system × rater interaction, which is the correct error term when annotators are a sample from a population. The interaction mean square exceeds the system mean square.
Panel
Top system
ρ w/ full
Claude rank
All three
Gemma-3-27B
—
3
− A1
Claude-3-Haiku
0.55
1
− A2
Gemma-3-27B
0.90
4
− A3
Mistral-Large
0.62
8
Table 2: Leave-one-annotator-out leaderboards. Removing any single annotator from a three-person panel can change the winner. Claude-3-Haiku ranges from first to last.
System
mos
emojis
ead
win LM
Gemma-3-27B
3.667
4.95
0.653
0.543
Mistral-Large
3.638
4.87
0.639
0.506
Claude-3-Haiku
3.579
4.92
0.584
0.499
GPT-4.1-nano
3.516
4.55
0.562
0.428
DeepSeek-V3.2
3.454
3.45
0.560
0.541
Gemini-2.0-Flash
3.443
2.67
0.657
0.603
Table 3: Systems ordered by raw mos . Mean emoji count tracks mos almost perfectly; decodability does not. win LM is the within-item length-matched win rate, which reorders the table. ∗∗p<.01 .