When Reasoning Goes Astray: Attention Dynamics of Uncontrolled Reasoning
Organizations: Beijing University of Posts and Telecommunications · Wuhan University · Nanyang Technological University · Chongqing University of Posts and Telecommunications
Abstract
Large reasoning models (LRMs) improve performance on complex tasks through extended reasoning, yet the same process can degenerate into redundant verification and persistent generation loops. Such uncontrolled reasoning increases inference cost and creates risks of resource exhaustion and service degradation. However, existing mitigations largely truncate long outputs or react to surface repetition, and thus fail to distinguish normal thinking from uncontrolled reasoning or explain how benign reasoning degenerates into harmful behavior. In this paper, we operationalize LRM generation as four states and further introduce Reasoning-state Analysis via Dynamic Attention Responses (RADAR), which identifies the current reasoning state in real time and characterizes how effective reflection can develop into uncontrolled generation. Guided by RADAR's analysis, we further realign abnormal attention distributions toward patterns observed in normal requests and examine how this correction affects excessive reflection and persistent looping. Temporal analyses show that uncontrolled reasoning is characterized by attention distributions that deviate from normal generation, with abnormal trends becoming detectable before repetition begins. Correcting these deviations through Attention Realignment consistently reduces looping while largely preserving benign performance. Together, RADAR provide a mechanistic account of how reasoning becomes uncontrolled, offering actionable guidance for identifying critical failure stages and designing targeted runtime interventions.
Figures & tables
| Attack | |||||||
|---|---|---|---|---|---|---|---|
| Recur | 994.3 | 62.3 | 75.1 | 108.0 | 266.9 | 1052.0 | 1911.3 |
| LoopLLM | 2141.0 | 801.0 | 997.0 | 1265.0 | 1268.7 | 1676.7 | 2461.3 |
| Joint | 0 | 52.2 | 60.6 | 69.2 | 88.1 | 132.8 | 204.7 |
| MiP | 4447.3 | 367.8 | 450.4 | 578.4 | 734.8 | 1004.0 | 1598.8 |
| Final loop rate | Benign accuracy | |||||
|---|---|---|---|---|---|---|
| Model | Original | RADAR | (pp) | Original | RADAR | (pp) |
| DeepSeek-R1-Distill-Llama-8B | 56.0% | 40.0% | 66.4% | 65.6% | ||
| DeepSeek-R1-Distill-Qwen-14B | 40.8% | 24.8% | 72.8% | 74.4% | ||
| QwQ-32B | 22.4% | 18.4% | 71.2% | 76.0% | ||
| GLM-4.7-Flash | 36.8% | 26.4% | 65.6% | 69.6% | ||
| Qwen3.6-27B | 28.8% | 26.4% | 61.6% | 72.0% | ||
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
| State | Reflection behavior | Output behavior | Quantitative cue |
|---|---|---|---|
| No explicit reflection or backtracking. | Completes a single forward solution path, regardless of correctness. | . | |
| Bounded verification, correction, or reconsideration that makes substantive progress. | Reaches a definite answer after useful reflection. | . | |
| Reconsideration continues after a sufficient solution prefix, when useful progress has saturated. | Remains semantically variable but adds little new information. | , without a persistent lexical loop. | |
| Identical or near-identical lexical content recurs without substantive progress. | Fails to terminate naturally before reaching the generation cap. | Generation cap reached and . |
| State | Source record | Prompt and trajectory excerpt | Label evidence |
|---|---|---|---|
| GSM8K | Prompt: Janet’s ducks lay 16 eggs per day; after eating three and using four for muffins, how much does she earn by selling the rest for =3+4=7=16-7=9=9\times$2=$1818$ .” | One forward calculation reaches a definite answer and terminates normally after 488 tokens, without explicit verification or backtracking. | |
| GPQA | Prompt: In a universe with isolated magnetic poles, which Maxwell equation changes? Excerpt: “The equation is , which … states that there are no isolated magnetic poles. … Wait, but let me double-check. … The main change would be in the divergence of … So the correct answer should be option C.” | A bounded double-check examines a possible alternative, confirms the initial physical argument, and ends with a definite answer after 722 tokens. | |
| Recur | Prompt: Infer a nurses’ station’s starting stock from three days of bandage use and orders, while explaining why 19 is incorrect. Excerpt: “At the end, . But starting with 19, using 38 is impossible. … Therefore, the problem is flawed. … However, … perhaps the correct answer is different. Maybe I made a mistake in the equations. Let me try to set up the equations again, carefully.” | The trajectory repeatedly revisits delivery timing and the same equations after useful progress has saturated. It remains semantically variable, produces no final answer, and reaches the 16,384-token cap without collapsing to one fixed lexical span. | |
| LoopLLM | Prompt: Remove all vowels from a sentence containing an adversarial suffix. Excerpt: “Wait, but in the original, it’s ‘ * * af Zh * * … ’, so after processing, it’s ‘ * * f Zh * * … ’. … So the final processed sentence is: ‘ Thqckbrwnfxjmpsrvthlzydg. * * f Zh * * … ’.” The same clauses and symbol span then recur. | A near-identical lexical unit repeats without further progress; the run fails to terminate naturally and reaches 16,331 tokens (the logged cap for this record). |
| Dataset | DS-Llama-8B | DS-Qwen-14B | QwQ-32B | Qwen3.6-27B | GLM-4.7-Flash |
|---|---|---|---|---|---|
| Normal trajectories | |||||
| GSM8K | 100.0% (D) | 100.0% (D) | 100.0% (D) | 100.0% (D) | 96.0% (D) |
| MMLU-Geor | 96.0% (D) | 96.0% (D) | 96.0% (D) | 100.0% (D) | 96.0% (D) |
| GPQA | 100.0% (R) | 100.0% (R) | 100.0% (R) | 100.0% (R) | 100.0% (R) |
| MMLU-Econometrics | 48.0% (D) | 68.0% (R) | 50.0% (R) | 68.0% (R) | 52.0% (R) |
| MMLU-World-History | 100.0% (R) | 100.0% (R) | 100.0% (R) | 100.0% (R) | 100.0% (R) |
| Model | Micro-F1 |
|---|---|
| DeepSeek-R1-Distill-Llama-8B | 92.0% |
| DeepSeek-R1-Distill-Qwen-14B | 89.9% |
| QwQ-32B | 84.3% |
| Qwen3.6-27B | 88.8% |
| GLM-4.7-Flash | 93.8% |
| Average | 89.8% |
| Model | |||
|---|---|---|---|
| DeepSeek-R1-Distill-Llama-8B | 638.58 | 1,836.56 | 1,357.37 |
| DeepSeek-R1-Distill-Qwen-14B | 668.54 | 1,528.44 | 1,184.48 |
| QwQ-32B | 1,384.90 | 1,845.13 | 1,661.04 |
| GLM-4.7-Flash | 987.12 | 1,991.77 | 1,589.91 |
| Qwen3.6-27B | 1,246.24 | 2,702.08 | 2,119.74 |
| Component | Version |
|---|---|
| Python | 3.12.9 |
| PyTorch | 2.7.0+cu126 |
| CUDA | 12.6 |
| cuDNN | 9.5.1 |
| Transformers | 4.52.3 |
| vLLM | 0.9.0 |
| Configuration identifier | Data sources |
|---|---|
| concise_reasoning | GSM8K, MMLU-Geor |
| productive_reasoning | GPQA, Econometrics, World_History |
| repetitive_reasoning | RECUR |
| repetitive_string | LoopLLM |
| repetitive_string | Joint |
| repetitive_string | Missing Premise (MiP) |
| Undefended | RADAR | AUSteer | n-gram | ||||
|---|---|---|---|---|---|---|---|
| Attack | Rate | Rate | (pp) | Rate | (pp) | Rate | (pp) |
| Recur | 56.0% | 44.0% | 48.0% | 0.0% | |||
| LoopLLM | 56.0% | 28.0% | 20.0% | 0.0% | |||
| MiP | 56.0% | 48.0% | 44.0% | 0.0% | |||
| Macro average | 56.0% | 40.0% | 37.3% | 0.0% | |||
| Method | Benign accuracy | (pp) |
|---|---|---|
| Undefended | 66.4% | – |
| RADAR | 65.6% | |
| AUSteer | 66.4% | |
| n-gram | 58.4% |
| Attack | Setting | Llama-8B | Qwen-14B | QwQ-32B | Qwen3.6 | GLM-4.7 | Overall |
|---|---|---|---|---|---|---|---|
| Recur | Undefended Loop | 56.0% | 12.0% | 8.0% | 4.0% | 0.0% | 16.0% |
| RADAR Loop | 44.0% | 8.0% | 4.0% | 4.0% | 0.0% | 12.0% | |
| LoopLLM | Undefended Loop | 56.0% | 36.0% | 0.0% | 0.0% | 80.0% | 34.4% |
| RADAR Loop | 28.0% | 0.0% | 0.0% | 0.0% | 36.0% | 12.8% | |
| MiP | Undefended Loop | 56.0% | 64.0% | 80.0% | 60.0% | 44.0% | 60.8% |
| RADAR Loop | 48.0% | 56.0% | 68.0% | 48.0% | 44.0% | 52.8% |
| Llama-8B | Qwen-14B | QwQ-32B | Qwen3.6-27B | GLM-4.7 | |
|---|---|---|---|---|---|
| 0.1 | 86.55% | 66.18% | 73.82% | 82.18% | 68.36% |
| 0.2 | 89.45% | 72.73% | 80.36% | 88.00% | 70.91% |
| 0.3 | 90.18% | 74.18% | 85.82% | 90.18% | 74.91% |
| 0.4 | 90.91% | 74.91% | 87.64% | 90.55% | 76.36% |
| 0.5 | 91.27% | 75.27% | 90.18% | 91.27% | 77.45% |
| 0.6 | 91.27% | 77.09% | 91.27% | 92.00% | 79.64% |
| Selection threshold | Evaluated | Mean layers | Mean / 8 |
|---|---|---|---|
| 0.4 | 80 | 1.095 | 13.7% |
| 0.6 | 80 | 1.527 | 19.1% |
| 0.7 | 80 | 1.838 | 23.0% |
| Layer index | |||
|---|---|---|---|
| 11 | 15.6% | 20.3% | 24.4% |
| 12 | 10.6% | 15.9% | 18.3% |
| 13 | 15.3% | 19.4% | 24.7% |
| 14 | 15.0% | 20.1% | 24.4% |
| 15 | 14.5% | 20.9% | 26.1% |
| 16 | 13.4% | 19.9% | 23.4% |
| Model | Shared subset | ||||||
|---|---|---|---|---|---|---|---|
| Llama-8B | – | – | – | – | – | – | – |
| Qwen-14B | – | – | – | – | – | – | – |
| QwQ-32B | – | – | – | – | – | – | – |
| Qwen3.6-27B | |||||||
| GLM-4.7 | – | – | – | – | – | – | – |