Reason in Style: Discovering and Controlling Style in Language Models
Organizations: NYU · NYU & NYU Langone Health
Abstract
Language models learn content and style jointly, making stylistic variation in their outputs difficult to identify and control. We study whether recurring styles in model responses can be discovered without supervision and explicitly controlled. We design an algorithm that learns to separate representations of content and style from language models' outputs and validate its effectiveness on math questions in a controlled setting. By applying this method to over 100K verified traces from nine distinct teacher models, we discover six recurring yet imbalanced styles. We then fine-tune smaller student models to follow these styles when explicitly conditioned on them, using importance weighting to balance the contribution of the styles represented in the corpus. This approach improves Pass@ over standard fine-tuning on the same data across six math reasoning benchmarks, demonstrating that we can diversify the style of answers effectively. We confirm that this also results in strong correspondence between requested and realized styles. We find that style affects correctness: the probability of solving a problem depends on the style we condition on, and different problems benefit from different styles. In summary, our results show that stylistic variation in model-generated data can be discovered in an unsupervised way, and made explicit, providing a source of both control and improved reasoning performance.
Figures & tables
| Style | gemma | gpt-5 | gpt-oss | llama | olmo | phi-4 | qwen2.5 | qwen3 | qwen3.5 |
|---|---|---|---|---|---|---|---|---|---|
| 30.2 | 0.3 | 0.1 | 34.5 | 4.0 | 0.0 | 28.6 | 0.4 | 1.9 | |
| 16.4 | 10.8 | 5.5 | 16.0 | 12.4 | 8.1 | 13.2 | 6.5 | 11.0 | |
| 0.8 | 29.6 | 34.7 | 0.0 | 18.8 | 9.3 | 2.6 | 0.3 | 3.9 | |
| 2.4 | 24.0 | 23.3 | 0.1 | 19.2 | 18.7 | 5.2 | 2.6 | 4.5 | |
| 10.0 | 10.0 | 11.9 | 8.7 | 12.0 | 7.9 | 10.9 | 16.6 | 12.1 | |
| 1.9 | 4.9 | 0.7 | 3.5 | 4.0 | 31.8 | 0.1 | 24.2 | 28.8 |
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
| ID | Name | Prompted surface form |
|---|---|---|
| 0 | Academic | Formal paragraphs; passive voice; frequent use of connective phrases such as “therefore” and “it follows that”. |
| 1 | ELI5 / tutor | Short sentences; explanatory analogies; encouraging, tutorial-like tone. |
| 2 | Algorithmic list | Step-by-step organization, with each sentence explicitly formatted as a numbered or ordered step. |
| 3 | Pure equation | Minimal natural language; equation-dominated derivations using symbols such as and . |
| Metric | Frozen GTE | Recon-only | Full objective |
|---|---|---|---|
| question gap | — | 0.36 | 0.93 |
| style gap | — | 0.09 | |
| style gap | — | 0.10 | 1.18 |
| question gap | — | 0.38 | |
| style ARI / NMI | 0.29 / 0.36 | 0.85 / 0.82 | 0.77 / 0.78 |
| style ARI / NMI | — | 0.91 / 0.87 | / |
| Teacher | ||||||
|---|---|---|---|---|---|---|
| gemma-4-31b-it | 41.6 | 16.1 | 0.7 | 2.3 | 37.0 | 2.3 |
| gpt-5-chat | 0.4 | 10.6 | 23.9 | 22.1 | 37.1 | 5.9 |
| gpt-oss-120b | 0.1 | 5.4 | 28.1 | 21.4 | 44.3 | 0.8 |
| llama-3.3-70b-instruct | 47.6 | 15.7 | 0.0 | 0.1 | 32.4 | 4.2 |
| olmo-3.1-32b-instruct | 5.6 | 12.1 | 15.2 | 17.7 | 44.7 | 4.7 |
| phi-4-reasoning-plus | 0.0 | 8.0 | 7.5 | 17.2 | 29.4 | 37.8 |
| # distinct styles | 1 | 2 | 3 | 4 | 5 | 6 |
|---|---|---|---|---|---|---|
| % of questions | 0.8 | 15.8 | 36.6 | 38.3 | 8.2 | 0.3 |
| # questions | 105 | 1,963 | 4,545 | 4,754 | 1,021 | 38 |
| Feature | |||||||
|---|---|---|---|---|---|---|---|
| words | 311 | 330 | 417 | 483 | 486 | 723 | 0.09 |
| dens. backtrack | 0.01 | 0.10 | 0.12 | 0.20 | 0.17 | 0.34 | 0.07 |
| dens. verification | 0.02 | 0.08 | 0.04 | 0.07 | 0.11 | 0.20 | 0.09 |
| dens. equals | 6.2 | 4.7 | 8.4 | 9.3 | 7.0 | 5.9 | 0.06 |
| lines / 100w | 13.0 | 17.5 | 27.9 | 26.5 | 19.9 | 13.8 | 0.12 |
| LLM style | Teacher mix (of 20 per teacher) | |
|---|---|---|
| Exploratory first-person monologue | 20 | qwen3 |
| Telegraphic “We are asked” think-block | 20 | phi-4 |
| Structured request-analysis outline | 18 | qwen3.5 |
| H2-numbered recipe with stock closer | 20 | llama |
| Restating markdown tutor | 34 | gpt-5 , olmo |
| Compact contest exposition | 25 | gpt-oss ; also gpt-5 2, qwen3.5 2, qwen2.5 1 |
| Teacher | mono. | think | outline | H2 | md tutor | contest | walk. |
|---|---|---|---|---|---|---|---|
| gemma-4-31b-it | 0 | 0 | 0 | 0 | 0 | 0 | 20 |
| gpt-5-chat | 0 | 0 | 0 | 0 | 18 | 2 | 0 |
| gpt-oss-120b | 0 | 0 | 0 | 0 | 0 | 20 | 0 |
| llama-3.3-70b-instruct | 0 | 0 | 0 | 20 | 0 | 0 | 0 |
| olmo-3.1-32b-instruct | 0 | 0 | 0 | 0 | 16 | 0 | 4 |
| phi-4-reasoning-plus | 0 | 20 | 0 | 0 | 0 | 0 | 0 |