Language models learn content and style jointly, making stylistic variation in their outputs difficult to identify and control. We study whether recurring styles in model responses can be discovered without supervision and explicitly controlled. We design an algorithm that learns to separate representations of content and style from language models' outputs and validate its effectiveness on math questions in a controlled setting. By applying this method to over 100K verified traces from nine distinct teacher models, we discover six recurring yet imbalanced styles. We then fine-tune smaller student models to follow these styles when explicitly conditioned on them, using importance weighting to balance the contribution of the styles represented in the corpus. This approach improves Pass@k over standard fine-tuning on the same data across six math reasoning benchmarks, demonstrating that we can diversify the style of answers effectively. We confirm that this also results in strong correspondence between requested and realized styles. We find that style affects correctness: the probability of solving a problem depends on the style we condition on, and different problems benefit from different styles. In summary, our results show that stylistic variation in model-generated data can be discovered in an unsupervised way, and made explicit, providing a source of both control and improved reasoning performance.
Figures & tables
Figure 1: Pass@ k for the Qwen3 0.6B student. Style SFT (Importance-Weighted) is our full method, combining discovered style conditioning with importance-weighted SFT. Style SFT (Empirical) conditions on the discovered style labels while retaining the empirical training distribution. Random-Style SFT uses randomly assigned style labels as a control. Vanilla SFT uses the same training traces without style information. Our full method substantially improves Pass@ k , while vanilla SFT often provides little benefit or degrades performance.
Style
gemma
gpt-5
gpt-oss
llama
olmo
phi-4
qwen2.5
qwen3
qwen3.5
s1
30.2
0.3
0.1
34.5
4.0
0.0
28.6
0.4
1.9
s2
16.4
10.8
5.5
16.0
12.4
8.1
13.2
6.5
11.0
s3
0.8
29.6
34.7
0.0
18.8
9.3
2.6
0.3
3.9
s4
2.4
24.0
23.3
0.1
19.2
18.7
5.2
2.6
4.5
s5
10.0
10.0
11.9
8.7
12.0
7.9
10.9
16.6
12.1
s6
1.9
4.9
0.7
3.5
4.0
31.8
0.1
24.2
28.8
Table 1: Teacher composition of each discovered style (%; rows sum to 100).
Figure 2: Correctness by style on MATH-500. Changing the style condition produces differences in accuracy at all three model scales.
Figure 3: The most useful style depends on the problem. For each of 400 random splits, 128 of the 256 generations per question and style are used for style selection and the remaining 128 for held-out evaluation. Left: fraction of questions for which each style is selected as best; several different styles are preferred across problems. Right: held-out accuracy under a uniform mixture of styles, the globally best fixed style, and a style selected separately for each question. Question-specific selection consistently outperforms the globally best fixed style, showing that style usefulness cannot be explained by a single global ranking.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
ID
Name
Prompted surface form
0
Academic
Formal paragraphs; passive voice; frequent use of connective phrases such as “therefore” and “it follows that”.
1
ELI5 / tutor
Short sentences; explanatory analogies; encouraging, tutorial-like tone.
2
Algorithmic list
Step-by-step organization, with each sentence explicitly formatted as a numbered or ordered step.
3
Pure equation
Minimal natural language; equation-dominated derivations using symbols such as ⇒ and ∴ .
Appendix
Table 2: Four prompted Gemini writing styles used in the controlled experiment.
Metric
Frozen GTE h
Recon-only
Full objective
c question gap
—
0.36
0.93
c style gap
—
0.09
≈0
z style gap
—
0.10
1.18
z question gap
—
0.38
−0.30
z style ARI / NMI
0.29 / 0.36
0.85 / 0.82
0.77 / 0.78
c style ARI / NMI
—
0.91 / 0.87
≈0 / 0.00
Appendix
Table 3: Full disentanglement results on the controlled Gemini setting. Gap is the mean within-group cosine similarity minus the mean between-group similarity. ARI/NMI are from k -means with k=4 against the ground-truth style_id . The style probe is a 4-way linear classifier with chance accuracy 25% .
Teacher
s1
s2
s3
s4
s5
s6
gemma-4-31b-it
41.6
16.1
0.7
2.3
37.0
2.3
gpt-5-chat
0.4
10.6
23.9
22.1
37.1
5.9
gpt-oss-120b
0.1
5.4
28.1
21.4
44.3
0.8
llama-3.3-70b-instruct
47.6
15.7
0.0
0.1
32.4
4.2
olmo-3.1-32b-instruct
5.6
12.1
15.2
17.7
44.7
4.7
phi-4-reasoning-plus
0.0
8.0
7.5
17.2
29.4
37.8
Appendix
Table 4: Share of each teacher’s traces assigned to each K=6 style (%; rows sum to 100). The final row gives the overall style distribution across all 111,834 traces.
# distinct styles
1
2
3
4
5
6
% of questions
0.8
15.8
36.6
38.3
8.2
0.3
# questions
105
1,963
4,545
4,754
1,021
38
Appendix
Table 5: Number of distinct styles among the nine teacher solutions to the same question. There are 12,426 aligned questions with nine traces each.
Feature
s1
s2
s3
s4
s5
s6
η2
n words
311
330
417
483
486
723
0.09
dens. backtrack
0.01
0.10
0.12
0.20
0.17
0.34
0.07
dens. verification
0.02
0.08
0.04
0.07
0.11
0.20
0.09
dens. equals
6.2
4.7
8.4
9.3
7.0
5.9
0.06
lines / 100w
13.0
17.5
27.9
26.5
19.9
13.8
0.12
Appendix
Table 6: Handcrafted properties of the six discovered styles. η2 gives the fraction of variance in each feature explained by style membership.
LLM style
n
Teacher mix (of 20 per teacher)
Exploratory first-person monologue
20
qwen3 20/20
Telegraphic “We are asked” think-block
20
phi-4 20/20
Structured request-analysis outline
18
qwen3.5 18/20
H2-numbered recipe with stock closer
20
llama 20/20
Restating markdown tutor
34
gpt-5 18/20 , olmo 16/20
Compact contest exposition
25
gpt-oss 20/20 ; also gpt-5 2, qwen3.5 2, qwen2.5 1
Appendix
Table 7: Blind LLM styles on the 180 -trace sample ( 20 writeups per teacher) vs. the oracle teacher that produced each writeup. Most styles recover a single teacher at 18 – 20/20 ; the markdown-tutor bin merges gpt-5 and olmo; the instructional walkthrough merges gemma and qwen2.5.
Teacher
mono.
think
outline
H2
md tutor
contest
walk.
gemma-4-31b-it
0
0
0
0
0
0
20
gpt-5-chat
0
0
0
0
18
2
0
gpt-oss-120b
0
0
0
0
0
20
0
llama-3.3-70b-instruct
0
0
0
20
0
0
0
olmo-3.1-32b-instruct
0
0
0
0
16
0
4
phi-4-reasoning-plus
0
20
0
0
0
0
0
Appendix
Table 8: Same 180 -trace sample, transposed: for each teacher, how many of its 20 writeups the LLM assigned to each invented style. Pure diagonals (e.g. llama → H2 recipe, gpt-oss → contest) are the recoverable fingerprints; olmo and gpt-5 share the markdown-tutor bin.
Figure 4: Qwen3 0.6B importance sampling variants
Figure 5: Qwen3 1.7B importance sampling variants
Figure 6: Qwen3 4B importance sampling variants
Figure 7: Pass@ k Qwen3 1.7B
Figure 8: Pass@ k Qwen3 4B
Figure 9: Requested versus realized style on MATH-500. Rows denote the forced prefix [style_i] and columns denote the realized style s^ assigned by the nearest teacher-style centroid in the handcrafted style feature space. Values report P(s^=j∣[style_i]) . Diagonal mass measures adherence to the requested teacher-style mode, while off-diagonal mass reveals which styles are conflated. (a) Importance-weighted Style SFT yields substantially greater style fidelity than (b) empirical Style SFT, although adherence remains uneven across modes.