cs.SESep 17, 2025

SEER: Self-Enhancing Chain-of-Thought Compression for Reasoning Models

Authors: Kerui Huang, Shuhan Liu, Xing Hu, Tongtong Xu, Lingfeng Bao, Xin Xia

Organizations: The State Key Laboratory of Blockchain and Data Security, Zhejiang University, China · State Key Laboratory for Novel Software Technology, Nanjing University, China

Abstract

Chain-of-Thought (CoT) prompting can substantially improve the reasoning ability of large language models (LLMs), but it often comes with high inference cost due to long and poorly controlled reasoning traces. This overhead is particularly problematic in software engineering tasks (e.g., code generation), where both latency and output reliability matter. To better understand this trade-off, we conduct an empirical study on widely used code generation benchmarks and observe that many modern reasoning models produce excessively verbose CoTs (often thousands of tokens), which frequently leads to truncation and unstable generation. Using a strict n-gram repetition detector, we find that most observed truncations are associated with degenerate looping behaviors. In addition, a HumanEval/129 case study shows that failed generations can be longer than successful ones, suggesting limited returns from overlong reasoning. Motivated by these findings, we propose SEER (Self-Enhancing Efficient Reasoning), a self-enhancing framework for adaptive CoT compression. SEER improves the conciseness of reasoning while preserving output quality, without relying on external compression tools. SEER refines self-generated CoT data via Best-of-N sampling to suppress looping and redundant traces, then applies a lightweight, data-driven filter to encourage concise yet correct reasoning. It then fine-tunes the model on the filtered data to internalize concise reasoning behaviors. Across four software engineering benchmarks on the evaluated DeepSeek-R1-Distill-Qwen-7B backbone, SEER reduces CoT length by 34.6% on average while improving task performance, with reduced truncation and fewer reasoning loops.

Explore similar work

CardsList
  1. Reconsidering Overthinking: Penalizing Internal and External Redundancy in CoT Reasoning

    Aug 4, 2025Taihang Zhen, Jialiang Hong, Kai Chen +12Large Reasoning ModelsChain-of-Thought Reasoning

  2. OS-Pruner: Pruning Chains-of-Thought of Reasoning Models via Optimal Stopping

    Jul 13, 2026Mohammed Ehab, Aymane El Gadarri, Vivek F. Farias +2Large Reasoning ModelsReasoning Chain

  3. SLAT: Segment-Level Adaptive Trimming for Efficient CoT Reasoning

    May 29, 2026Jian Yao, Xiongcai Luo, Ran Cheng +1Large Reasoning ModelsChain-of-Thought Reasoning