cs.AIOct 8, 2026

When Should Agents Think? Adaptive Reasoning via Cross-Turn Estimation

Authors: Yiruo Cheng, Shen Huang, Xiaoshuai Song, Jiejun Tan, Guanting Dong, Pengjun Xie, Ji-Rong Wen, Zhicheng Dou

Organizations: Gaoling School of Artificial Intelligence, Renmin University of China · Alibaba Token Hub, Alibaba Group

Abstract

Large language model (LLM)-based agents have demonstrated strong capabilities on complex tasks. They typically perform reasoning before each action throughout an interaction trajectory. However, reasoning may not be necessary at every turn, as reasoning produced earlier can continue to support subsequent actions. A key challenge is therefore to determine when existing reasoning remains sufficient and when a new reasoning step is needed, without relying on costly generation-based verification. We find that decreases in the likelihood of subsequent reference actions after removing additional reasoning closely track whether those actions remain recoverable given earlier reasoning, providing an effective and lightweight signal for estimating cross-turn action support. Based on this observation, we propose Reasoning Adaptation through Cross-Turn Estimation (RACE), a training approach for adaptive agent reasoning. RACE introduces a Likelihood-Guided Progressive Reasoning Cover Detection (LoGiC) procedure that progressively identifies reasoning turns whose removal has limited impact on the current and subsequent reference actions. The resulting removal signals are incorporated into both supervised fine-tuning and agentic reinforcement learning, enabling the policy to learn when to reason and when to act directly. Extensive experiments on four representative agent benchmarks show that RACE substantially reduces reasoning cost while maintaining or improving task performance.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Adaptive Latent Agentic Reasoning

    Jun 1, 2026Dongwon Jung, Peng Shi, Yi Zhang +2Agentic ReasoningLLM Agents

  2. Agora: Enhancing LLM Agent Reasoning Via Auction-Based Task Allocation

    Jul 10, 2026Kaiji Zhou, Aleš Leonardis, Yue FengLLM Agent OrchestrationLLM Agent Routing

  3. Latent Reward Steering: An Adaptive Inference-Time Framework that Implicitly Promotes Cognitive Behaviors in Reasoning LLMs

    May 30, 2026Jiakang Li, Guanyu Zhu, Can Jin +8Reward ModelingLanguage Model Steering