Visible Reasoning Is Not a Universal Optimizer: Persona- and Thinking-Dependent Effects in Analytics Code Generation
Organizations: Optum AI
Abstract
Visible Chain-of-Thought (CoT) is often treated as a broadly useful reasoning instruction, yet analytics code generation combines natural-language ambiguity, schema grounding, target-language constraints, and model-specific inference behavior. Because the same analytics request can be expressed in two distinct target languages-SQL and Python (pandas)-this setting provides a natural test of a common but under-examined assumption: that visible reasoning is more effective when its representation matches the requested target, as in "think in SQL" or "think in Python." Together with generic instructions such as "think step-by-step," such recommendations remain insufficiently evaluated under controlled, execution-based comparisons. We study a query matched SQL-pandas benchmark that crosses persona phrasing, target language, visible-CoT format, control prefixes, direct generation, and internal-reasoning configurations. The results do not support either a universal accuracy advantage from visible CoT or a consistent benefit from matching the reasoning representation to the target language. Instead, the effects depend on the persona, target, model configuration, and internal-reasoning setting. The control ablations further distinguish effects of reasoning content from those of prompt format. These findings indicate that reasoning strategies should be selected jointly for the model, persona, target, and internal-reasoning configuration rather than adopted as universal defaults. More broadly, the study provides a controlled framework for identifying when visible reasoning improves executable generation, when it primarily perturbs model behavior, and when the internal-reasoning configuration is the more consequential factor.
Figures & tables
| Model | Full evaluated model and configuration |
|---|---|
| GLM4 | GLM-4-32B-0414; non-reasoning |
| Qwen3 | Qwen3-30B-A3B-Instruct-2507; non-reasoning |
| Gemma4-off | Gemma 4 26B-A4B-IT; internal reasoning disabled |
| Gemma4-on | Gemma 4 26B-A4B-IT; internal reasoning enabled |
| Gemini Flash-0 | Gemini 2.5 Flash; internal-reasoning budget 0 |
| Gemini Flash-24k | Gemini 2.5 Flash; internal-reasoning budget 24,576 |
| Inference family | Tests | BH discoveries |
|---|---|---|
| Mode-direct | 252 | 0 |
| Target matching | 36 | 0 |
| Internal-reasoning toggle | 96 | 16 |
| Persona interactions | 756 | 1 |
| Model | Target | CoT | Control | |||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Direct | NL | SQL | Python | Pseudocode | mean | Filler | Scrambled | Wrong | ||
| GLM4 | SQL | 23.68 | 28.29 | 28.29 | 25.99 | 29.28 | 27.96 | 24.01 | 25.66 | 27.30 |
| GLM4 | Python | 15.13 | 19.74 | 18.42 | 17.11 | 17.76 | 18.26 | 16.45 | 15.13 | 15.79 |
| Qwen3 | SQL | 29.93 | 33.55 | 30.26 | 33.55 | 30.59 | 31.99 | 31.58 | 34.54 | 30.26 |
| Qwen3 | Python | 21.71 | 22.04 | 21.38 | 20.39 | 20.72 | 21.13 | 23.03 | 25.33 | 24.01 |
| Gemma4-off | SQL | 36.18 | 37.17 | 37.83 | 37.50 | 36.51 | 37.25 | 37.17 | 36.84 | 36.51 |
| Ablation | NL | direct | BH outcome (NL; direct) |
|---|---|---|---|
| Length filler | -0.36 | 0.90 | not detected; not detected |
| Scrambled | 0.16 | 1.43 | not detected; higher |
| Wrong content | -0.49 | 0.77 | not detected; not detected |
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
| Evaluation path | Runtime configuration |
|---|---|
| GLM4; Qwen3 | vLLM on AWS g6.12xlarge (4 L4 GPUs, 48 vCPUs, 192 GiB); BF16; 32,768-token context; no internal-reasoning setting. |
| Gemma4-off/on | Same AWS/vLLM host and checkpoint; BF16; 32,768-token context; request-level internal reasoning disabled/enabled. |
| Gemini Flash-0/24k | Provider API, issued from Mac M4 Max with 64 GB unified memory; provider-side serving hardware not observed. |
| All configurations | Temperature 0; fixed experiment seed; bounded repair and execution; deterministic reference normalization. |
| Model | Rows | Queries | Correct (%) |
|---|---|---|---|
| GLM4 | 14,592 | 304 | 37.38 |
| Qwen3 | 14,592 | 304 | 41.18 |
| Gemma4-off | 14,592 | 304 | 43.54 |
| Gemma4-on | 14,592 | 304 | 46.57 |
| Gemini Flash-0 | 14,592 | 304 | 46.15 |
| Gemini Flash-24k | 14,592 | 304 | 45.96 |
| Model | Persona | Target | Direct | NL CoT | SQL CoT | Python CoT | Pseudo code CoT | Length filler | Scrambled | Wrong content |
|---|---|---|---|---|---|---|---|---|---|---|
| GLM4 | expert | SQL | 88.16 | 87.83 | 86.18 | 87.83 | 88.82 | 90.79 | 86.84 | 87.83 |
| GLM4 | expert | PYTHON | 60.20 | 54.61 | 53.95 | 55.59 | 57.89 | 60.20 | 60.86 | 55.92 |
| GLM4 | novice | SQL | 19.08 | 21.38 | 18.42 | 21.05 | 18.42 | 20.72 | 26.32 | 20.72 |
| GLM4 | novice | PYTHON | 14.14 | 14.80 | 13.49 | 15.13 | 15.13 | 14.14 | 16.45 | 13.16 |
| GLM4 | original | SQL | 23.68 | 28.29 | 28.29 | 25.99 | 29.28 | 24.01 | 25.66 | 27.30 |
| GLM4 | original | PYTHON | 15.13 | 19.74 | 18.42 | 17.11 | 17.76 | 16.45 | 15.13 | 15.79 |
| Term family | |
|---|---|
| Persona | |
| Target | |
| Model configuration | |
| Mode | 0.620 |
| Mode persona | 0.00033 |
| Mode target | 0.626 |
| Phrasing | Target | Condition | Direct | Condition | pp | |
|---|---|---|---|---|---|---|
| Original | SQL | Pseudocode CoT | 33.39 | 35.31 | 1.92 | 0.168 |
| Original | SQL | NL CoT | 33.39 | 35.14 | 1.75 | 0.168 |
| Original | SQL | Scrambled | 33.39 | 35.36 | 1.97 | 0.196 |
| Novice | SQL | Scrambled | 27.47 | 32.62 | 5.15 | 0.00052 |
| Novice | Python | Scrambled | 18.86 | 21.11 | 2.25 | 0.111 |
| Novice | SQL | NL CoT | 27.47 | 28.84 | 1.37 | 0.439 |
| Phrasing | Target | Contrast | pp | 95% CI | Raw | |
|---|---|---|---|---|---|---|
| Novice | SQL | Scrambled direct | 5.15 | 0.00045 | ||
| Novice | SQL | Scrambled length filler | 5.04 | 0.000011 | ||
| Novice | SQL | Scrambled wrong content | 3.45 | 0.00063 | 0.0038 | |
| Novice | SQL | Scrambled explicit-CoT mean | 4.39 | 0.00019 | ||
| Novice | Python | Scrambled length filler | 2.47 | 0.0046 | 0.0184 | |
| Novice | Python | Scrambled wrong content | 2.58 | 0.0083 | 0.0250 |
| Persona | Difficulty | vs NL CoT | vs direct | ||
|---|---|---|---|---|---|
| Original | Simple | 0.00 | 1.0000 | 1.85 | 0.1646 |
| Original | Moderate | -0.62 | 0.8076 | 0.23 | 0.9469 |
| Original | Challenging | 1.31 | 0.4539 | 2.34 | 0.1646 |
| Novice | Simple | 2.85 | 0.0183 | 4.01 | 0.0026 |
| Novice | Moderate | 1.17 | 0.4539 | 3.04 | 0.0183 |
| Novice | Challenging | 3.84 | 0.0077 | 4.12 | 0.0034 |
| Reasoning condition | Accuracy (%) | Mean tokens | Mean latency (s) | pp | Extra tokens | Extra latency (s) |
|---|---|---|---|---|---|---|
| Direct | 42.94 | 2,153 | 37.1 | 0.00 | 0 | 0.0 |
| NL CoT | 43.41 | 3,095 | 69.3 | 0.48 | 942 | 32.2 |
| SQL CoT | 43.52 | 3,077 | 75.4 | 0.58 | 924 | 38.4 |
| Python CoT | 42.82 | 3,069 | 76.0 | -0.12 | 916 | 38.9 |
| Pseudocode CoT | 43.51 | 3,194 | 81.3 | 0.58 | 1,041 | 44.2 |
| Condition family | Accuracy (%) | Tokens | Latency (s) | pp | Extra tokens | Extra latency (s) | Dominated (T/L, %) |
|---|---|---|---|---|---|---|---|
| Direct | 42.94 | 2,153 | 37.1 | 0.00 | 0 | 0.0 | 0.0 / 0.0 |
| Explicit CoT mean | 43.32 | 3,109 | 75.5 | 0.38 | 956 | 38.4 | 41.7 / 38.9 |
| Difficulty | Condition family | Accuracy (%) | Mean tokens | Mean latency (s) | pp vs direct |
|---|---|---|---|---|---|
| Simple | Direct | 47.25 | 1,943 | 31.9 | 0.00 |
| Simple | Explicit CoT mean | 47.69 | 2,834 | 61.8 | 0.44 |
| Moderate | Direct | 42.52 | 2,128 | 34.2 | 0.00 |
| Moderate | Explicit CoT mean | 43.12 | 3,079 | 73.1 | 0.60 |
| Challenging | Direct | 38.20 | 2,439 | 46.9 | 0.00 |
| Challenging | Explicit CoT mean | 38.25 | 3,478 | 95.0 | 0.05 |
| Model / CoT | Accuracy | Tokens | Latency (s) |
|---|---|---|---|
| GLM4 direct | 36.73 | 1,537 | 20.5 |
| GLM4 NL CoT | 37.77 | 2,222 | 63.8 |
| Qwen3 direct | 40.13 | 1,549 | 13.6 |
| Gemma4-off direct | 43.70 | 1,873 | 28.8 |
| Gemma4-on direct | 46.05 | 3,743 | 144.7 |
| Gemma4-on SQL CoT | 47.15 | 6,056 | 318.3 |
| Control | Correctness OR | |
|---|---|---|
| Length filler | 1.325 | 0.00048 |
| Scrambled | 0.995 | 0.94542 |
| Wrong content | 1.263 | 0.00264 |