cs.LGSep 30, 2026

RATIO: Reasoning Analysis and Token-level Inference Optimization for Quantized Reasoning Models

Authors: Chengzhu Bao, Xianglong Yan, Tianao Zhang, Jiaqi Chen, Shaoqiu Zhang, Yulun Zhang

Organizations: Shanghai Jiao Tong University

Abstract

Post-training quantization (PTQ) has become a widely adopted technique for reducing the memory footprint and inference cost of large language models (LLMs). However, recent studies reveal that when applied to reasoning models, PTQ not only degrades reasoning performance but also exacerbates overthinking, leading to longer reasoning trajectories. These issues may offset the efficiency gains expected from lower-precision inference. Existing approaches mainly rely on complex optimization procedures. More recent lightweight inference strategies instead use predefined overthinking markers, limiting their adaptability across quantized models. To address these issues, we propose Reasoning Analysis and Token-level Inference Optimization (RATIO), a framework that identifies model-specific overthinking tokens and assigns each a tailored penalty. RATIO first introduces Quantization-aware Reasoning Behavior Analysis (QRBA) to identify overthinking tokens by analyzing discrepancies between full-precision and quantized models. It then adopts Token-Specific Penalty Determination (TSPD), which leverages full-precision guidance to derive token-specific penalties without additional training. Extensive experiments show that RATIO achieves a better accuracy-efficiency trade-off than existing token-level interventions. Specifically, RATIO achieves up to 9.8 points accuracy improvement and reduces chain-of-thought (CoT) length by up to 51.3% compared with quantized baselines. The code will be available at https://github.com/steven-bao1/RATIO.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Quantization Inflates Reasoning: Token Inflation as a Hidden Cost of Low-Bit Reasoning Models

    Jun 24, 2026Xinyu Lian, Walid Krichene, Beichen Huang +4Large Language Model QuantizationQuantization-Aware Training

  2. Quantized Reasoning Models Think They Need to Think Longer, but They Do Not

    May 29, 2026Sanae Lotfi, Polina Kirichenko, Steven Li +1Large Reasoning ModelsChain-of-Thought Reasoning