LLM arenas turn pairwise human preferences into model rankings. Those preferences may reflect how an answer is presented as well as what it says. We take a stylometric approach to 137,293 decisive French-language votes from the July 2026 Compar:IA release; the primary formatting analysis includes 137,113 battles across 116 models, and the joint estimates use the 127,092 battles with all required measurements. For each battle, we reconstruct the response visible when the user voted. We then compare the raw ranking with rankings adjusted for formatting, length, readability, vocabulary variety, and sentence structure. Presentation is associated with winning, but length, bold text, and lists tend to occur together, making their individual contributions hard to separate. Across the measured features, two associations change least across specifications: bold usage (+11.0% win odds per standard deviation in the joint model) and moving-average type-token ratio (MATTR), a measure of vocabulary variety that is less sensitive to answer length (+16.8%). The bold association is substantially smaller in observed multi-turn conversations, whereas the MATTR association changes little; because users choose whether to continue, this difference is descriptive rather than causal. The full adjustment moves 36 of 116 models by at least ten ranks. Yet comparisons with external benchmarks do not show that adjusted rankings better measure capability. We therefore recommend publishing raw and adjusted rankings side by side as a transparent sensitivity analysis.
Figures & tables
Raw turns
Decisive French battles
Models ( ≥ 100 battles)
comparia-fr-arena
641,277
137,293
116
Table 1
Feature
Description
Regex Pattern
Headers
Markdown headers (# through ######)
ˆ #{1,6}\s (multiline)
Lists
Ordered and unordered list items
ˆ \s*[-+]\s and ˆ \s\d+.\s
Bold
Bold-formatted text
**[ ˆ *]+**
Code blocks
Fenced code blocks
Emoji
Emoji characters
Unicode emoji ranges
Table 2
Feature
Alone
Together
95% CI
p (BH)
Significant?
Bold
+27.0%
+19.0%
[+16.0%, +22.4%]
0.002
Yes
Headers
+22.3%
+9.1%
[+3.0%, +12.6%]
0.002
Yes
Lists
+16.8%
+6.6%
[+4.5%, +9.0%]
0.002
Yes
Emoji
+8.8%
+3.5%
[+1.5%, +5.4%]
0.002
Yes
Code blocks
+4.8%
+0.8%
[ − 0.8%, +4.5%]
0.580
No
Table 3
Model
Std → Ctrl
Δ Rank
Sig?
mistral-small-2603
9 → 33
− 24
Yes
qwen-3-8b
80 → 102
− 22
Yes
mistral-large-2512
3 → 24
− 21
Yes
gpt-oss-120b
37 → 56
− 19
Yes
gpt-5.3
45 → 27
+18
Yes
qwen3-30b-a3b
78 → 95
− 17
Yes
Table 4
Figure 1: Left: joint style coefficients with 95% bootstrap CIs, coloured by feature family (grey = not significant after BH). Bold and MATTR remain clearly positive; raw TTR is strongly negative because it falls mechanically with length. Right: formatting coefficients shrink as length and then the linguistic features are added. Length has the largest coefficient when added to formatting and shrinks substantially once the linguistic features, which overlap with it, are included.
Feature
Single-turn
Multi-turn
Interaction δ
Sig?
Bold
+30.1%
+7.4%
− 0.191
Yes
Headers
+11.1%
+6.8%
− 0.040
Yes
Lists
+9.0%
+7.8%
− 0.012
No
Code blocks
+5.0%
+0.8%
− 0.041
Yes
Emoji
+2.4%
+6.0%
+0.035
Yes
Table 6
Figure 2: Depth visible at the retained vote. Left: win-odds change per SD for each feature in single-turn (circle) vs multi-turn (diamond) battles. Right: the interaction with 95% bootstrap CIs; grey is not significant after BH.
Topic
Bold
Politics & Government
+63.7%
Law & Justice
+43.9%
Health & Wellness & Medicine
+36.0%
Personal Development & Career
+35.3%
Arts
+34.8%
Food & Drink & Cooking
+24.1%
Table 8
Figure 3: Topic controls. Left: the bold association (win-odds change per SD) estimated within each topic, with unadjusted 95% bootstrap CIs. The point estimate is positive in every subject; the intervals for three smaller subjects include zero. The dashed line is the all-topic estimate under the same capped-contrast specification. Right: the vote-time multi-turn interactions of §4.3 (circle) versus the same model with topic × formatting interactions added (diamond); the two nearly coincide.
Task
Battles
Bold
Headers
Lists
Code blocks
Emoji
explanation
44,870
+26.7%*
+7.3%*
+14.4%*
− 1.0%
+0.9%
writing
12,587
+15.1%*
+9.6%*
+3.8%
− 6.1%*
+11.5%*
code
9,749
+28.1%*
+6.3%
+2.5%
+13.1%*
+0.6%
ideas
4,408
+19.1%*
+13.3%*
+3.2%
+6.4%
− 2.1%
list/table
2,956
+7.8%
+16.9%*
− 3.3%
+16.8%*
+13.1%*
summarisation
2,943
+34.7%*
+10.2%
+2.0%
− 4.5%
+5.7%
Table 10
Specification
Battles
Odds change
MATTR
127,092
+16.8%
MTLD
127,092
+12.7%
MATTR without French function words
119,503
+11.7%
MATTR excluding capitalised tokens
125,364
+13.3%
Table 11
Model
Raw
Style-controlled
Epoch Capabilities
live rank
live rank
Index coverage
GPT-5.3
47
1
Not matched
Mistral Medium 2508
2
28
Not matched
Gemini 3.1 Flash Lite
4
4
Not matched
Gemini 2.5 Flash
5
12
140.33
Gemini 3.1 Pro
15
27
154.90
Table 12
Capability benchmark
Matches
Raw
Formatting-
Full joint-
Formatting-controlled
controlled
controlled
minus raw (95% CI)
Epoch Capabilities Index
38
0.717
0.703
0.635
− 0.014 [ − 0.079, +0.054]
GPQA Diamond
32
0.753
0.741
0.664
− 0.011 [ − 0.063, +0.035]
FrontierMath
13
0.699
0.655
0.534
− 0.044 [ − 0.223, 0.000]
LiveBench
17
0.419
0.277
0.358
− 0.142 [ − 0.485, +0.097]
ARC-AGI-2
10
0.537
0.488
0.303
− 0.049 [ − 0.367, +0.209]
Table 13
Figure 4: Change in Spearman correlation relative to raw Compar:IA for non-arena capability benchmarks. Points show formatting-controlled and full joint-controlled rankings; lines are paired 95% bootstrap intervals.
LMArena preference ranking
Matches
Raw
Formatting-controlled
Full joint-controlled
Raw, overall
49
0.792
0.800
0.710
Style-controlled, overall
49
0.768
0.808
0.735
Raw, French
40
0.779
0.796
0.698
Style-controlled, French
40
0.701
0.773
0.693
Table 15
Feature
Bottom
Middle
Top
Bold
+20.3%
+12.2%
+13.5%
Lists
+13.8%
+8.8%
− 5.4%
Headers
+10.5%
+1.0%
+13.7%
N battles
32,439
13,691
15,928
Table 16
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Script
Section
Output
analyze_core.py
§4.1
formatting Bradley-Terry model, rank changes, position bias
formatting_interactions.py
§4.1
pairwise formatting interactions and likelihood-ratio test