RATIO: Reasoning Analysis and Token-level Inference Optimization for Quantized Reasoning Models
Organizations: Shanghai Jiao Tong University
Abstract
Post-training quantization (PTQ) has become a widely adopted technique for reducing the memory footprint and inference cost of large language models (LLMs). However, recent studies reveal that when applied to reasoning models, PTQ not only degrades reasoning performance but also exacerbates overthinking, leading to longer reasoning trajectories. These issues may offset the efficiency gains expected from lower-precision inference. Existing approaches mainly rely on complex optimization procedures. More recent lightweight inference strategies instead use predefined overthinking markers, limiting their adaptability across quantized models. To address these issues, we propose Reasoning Analysis and Token-level Inference Optimization (RATIO), a framework that identifies model-specific overthinking tokens and assigns each a tailored penalty. RATIO first introduces Quantization-aware Reasoning Behavior Analysis (QRBA) to identify overthinking tokens by analyzing discrepancies between full-precision and quantized models. It then adopts Token-Specific Penalty Determination (TSPD), which leverages full-precision guidance to derive token-specific penalties without additional training. Extensive experiments show that RATIO achieves a better accuracy-efficiency trade-off than existing token-level interventions. Specifically, RATIO achieves up to 9.8 points accuracy improvement and reduces chain-of-thought (CoT) length by up to 51.3% compared with quantized baselines. The code will be available at https://github.com/steven-bao1/RATIO.
Figures & tables
| AIME | MATH | GSM8K | GPQA | HumanEval | Avg. | Acc | Len% | ||
| Acc (%) / Len (k) | Acc / Len | vs. baseline | |||||||
| BF16 | 20.83 / 25.11 | 85.60 / 6.06 | 84.69 / 2.81 | 36.36 / 9.97 | 73.17 / 6.50 | 60.13 / 10.09 | – | – | |
| GPTQ | 4.17 / 51.53 | 54.00 / 24.27 | 68.76 / 9.54 | 31.82 / 23.67 | 10.37 / 39.31 | 33.82 / 29.66 | – | – | |
| + Lotfi et al. (2026) | 4.17 / 47.53 | 58.20 / 18.48 | 70.81 / 6.22 | 27.27 / 21.14 | 20.73 / 32.16 | 36.24 / 25.11 | |||
| +RATIO | 5.00 / 42.55 | 61.40 / 14.06 | 71.49 / 3.11 | 30.81 / 20.64 | 18.29 / 29.90 | 37.40 / 22.05 | |||
| AWQ | 7.50 / 55.53 | 44.80 / 36.13 | 61.18 / 22.16 | 23.23 / 36.56 | 17.68 / 39.81 | 30.88 / 38.04 | – | – | |
| GPQA | AIME | GSM8K | MATH | HumanEval | Avg. | Acc | Len% | ||
|---|---|---|---|---|---|---|---|---|---|
| Acc (%) / Len (k) | Acc / Len | vs. baseline | |||||||
| AWQ | 23.23 / 36.56 | 7.50 / 55.53 | 61.18 / 22.16 | 44.80 / 36.13 | 17.68 / 39.81 | 30.88 / 38.04 | – | – | |
| QRBA ( ) | 29.80 / 32.37 | 4.17 / 48.73 | 69.29 / 9.17 | 59.40 / 20.21 | 34.76 / 29.86 | 39.48 / 28.07 | |||
| QRBA ( ) | 29.29 / 26.25 | 4.17 / 42.11 | 70.36 / 4.08 | 62.60 / 14.10 | 31.10 / 26.39 | 39.50 / 22.59 | |||
| TSPD | 20.71 / 31.74 | 7.50 / 47.74 | 66.19 / 10.20 | 59.00 / 20.21 | 33.54 / 28.25 | 37.39 / 27.63 | |||
| RATIO | 26.26 / 19.87 | 6.67 / 32.52 | 72.18 / 3.63 | 62.80 / 13.86 | 35.37 / 22.68 | 40.66 / 18.51 | |||
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| GPQA | AIME | GSM8K | MATH | HumanEval | Avg. | Acc | Len% | ||
| Acc (%) / Len (k) | Acc / Len | vs. baseline | |||||||
| FlatQuant | 34.34 / 25.68 | 8.33 / 37.20 | 74.37 / 4.06 | 65.40 / 16.47 | 43.90 / 14.64 | 45.27 / 19.61 | – | – | |
| + Lotfi et al. (2026) | 30.30 / 20.81 | 11.67 / 39.26 | 75.66 / 2.83 | 69.80 / 13.30 | 48.17 / 11.75 | 47.12 / 17.59 | |||
| Qwen-1.5B | +RATIO | 30.81 / 22.16 | 11.67 / 37.77 | 75.51 / 1.69 | 73.20 / 11.87 | 51.22 / 11.90 | 48.48 / 17.08 | ||
| Qwen-1.5B | Qwen-7B | Llama-8B | Qwen-14B | ||||
|---|---|---|---|---|---|---|---|
| Token | Token | Token | Token | ||||
| _but | 1.22 | _but | 1.06 | _back | 1.00 | _but | 1.00 |
| _what | 1.21 | _back | 0.58 | _try | 0.90 | _what | 1.04 |
| _try | 0.73 | _check | 0.48 | _well | 0.86 | _think | 0.71 |
| _again | 1.16 | _But | 1.15 | _check | 0.93 | _But | 1.23 |
| _well | 0.88 | _something | 1.09 | _But | 1.12 | _something | 1.57 |
| Token | Token | ||
|---|---|---|---|
| _but | 1.11 | _question | 1.11 |
| _what | 1.03 | But | 1.06 |
| _back | 0.83 | _However | 0.94 |
| _think | 0.81 | but | 1.00 |
| _But | 1.05 | _perhaps | 0.84 |
| _something | 1.00 | Wait | 1.02 |
| Case 1: Repetitive reasoning after reaching the correct result | |
|---|---|
| Prompt | From her science class study, Brady learned that each whale has 40 gallons of blood and that a shark has three times as much blood as a whale. How many gallons of blood do ten sharks have? |
| GPTQ-W3 | Each shark has gallons, so ten sharks have gallons. Wait, let me double-check that. Wait, that seems correct. Let me think if I interpreted the question correctly. Alternatively, perhaps I can think of the problem as ten sharks each having three times as much as a whale, so . Therefore, ten sharks have gallons. Alternatively, perhaps I can think of it as . Wait, 3 W. 3 W3 W3 W. |
| RATIO | A shark has three times as much blood as a whale. Therefore, each shark has gallons of blood. Ten such sharks have gallons of blood. The final answer is . |
| Case 2: Repetition followed by an incorrect reasoning trajectory | |
| Prompt | Compute in your head. |
| GPTQ-W3 | First, . Adding 99 gives , and adding 1 gives . Let me verify each step again. is 9,801, then , and adding 1 gives 9,901. Alternatively, maybe I can factor it as . That gives . Alternatively, if I think of as , that is . Wait, no, that is incorrect. So, 9999. 1099999999999999999999999999999999999999. |