cs.LGJun 18, 2026

ADaPT: Token-Level Decoupling for Efficient Large Reasoning Models

Authors: Tingyun LiZishang JiangJinyi HanXinyi WangSihang JiangHan XiaZhaoqian DaiShuguang Ma+3 more

Organizations: School of Data Science, Fudan University · Shanghai Institute of Artificial Intelligence for Education, East China Normal University · College of Computer Science and Artificial Intelligence, Fudan University · Ant Group

Abstract

Large reasoning models rely on long chain-of-thought to achieve strong performance, but applying such reasoning uniformly incurs high computational cost. Existing efficiency-oriented methods attempt to shorten or mix reasoning strategies, yet often degrade reasoning capability. We identify the root cause as sequence-level coupling between efficiency incentives and correctness optimization, which implicitly penalizes long but correct reasoning trajectories. To address this issue, we propose Adaptive Dual-Process Thinking (ADaPT), a token-level dual-process framework that explicitly decouples efficiency and correctness signals during training. ADaPT introduces a mode-selection token to control fast and slow reasoning, applying efficiency-related rewards exclusively to this token to avoid penalizing correct long reasoning while encouraging efficiency when appropriate. Moreover, ADaPT enables precise and continuous control over the efficiency-performance trade-off at inference time: by adjusting the generation probability of the mode-selection token, a single trained model can smoothly move along the efficiency-performance Pareto frontier. Extensive experiments demonstrate that ADaPT significantly reduces inference cost while maintaining strong reasoning performance across multiple benchmarks.

Explore similar work

CardsList