Large language model (LLM)-based agents have demonstrated strong capabilities on complex tasks. They typically perform reasoning before each action throughout an interaction trajectory. However, reasoning may not be necessary at every turn, as reasoning produced earlier can continue to support subsequent actions. A key challenge is therefore to determine when existing reasoning remains sufficient and when a new reasoning step is needed, without relying on costly generation-based verification. We find that decreases in the likelihood of subsequent reference actions after removing additional reasoning closely track whether those actions remain recoverable given earlier reasoning, providing an effective and lightweight signal for estimating cross-turn action support. Based on this observation, we propose Reasoning Adaptation through Cross-Turn Estimation (RACE), a training approach for adaptive agent reasoning. RACE introduces a Likelihood-Guided Progressive Reasoning Cover Detection (LoGiC) procedure that progressively identifies reasoning turns whose removal has limited impact on the current and subsequent reference actions. The resulting removal signals are incorporated into both supervised fine-tuning and agentic reinforcement learning, enabling the policy to learn when to reason and when to act directly. Extensive experiments on four representative agent benchmarks show that RACE substantially reduces reasoning cost while maintaining or improving task performance.
Figures & tables
Figure 1: Analysis of cross-turn reasoning support. (a) Reference-action recovery at subsequent turns after removing later reasoning. (b) Decreases in the likelihood of subsequent reference actions closely correspond to the recovery rate, providing a lightweight proxy for action recoverability.
Figure 2: Overview of the Likelihood-Guided Progressive Reasoning Cover Detection (LoGiC). (1) Reference-action likelihoods are initialized on the full trajectory. Starting from the second turn, (2) LoGiC tests whether the reasoning can be removed, and (3) verifies the removal based on likelihood decrease, with each accepted removal updating the context used to evaluate subsequent turns.
Figure 3: Overview of our training approach RACE. RACE uses the LoGiC algorithm to identify reasoning steps that can be safely removed, and the resulting signals are used in both supervised fine-tuning and agentic reinforcement learning to learn adaptive reasoning behavior.
Method
ScienceWorld
WebShop
AppWorld
DeepSearch
Success ↑
Avg Tok. ↓
Success ↑
Avg Tok. ↓
Success ↑
Avg Tok. ↓
Success ↑
Avg Tok. ↓
ReAct-style Iterative Reasoning Agents
DeepSeek-V4-Pro
32.11
1,662
40.20
1,288
87.69
3,016
40.00
1,243
GLM-5.3
30.29
1,854
46.60
1,715
77.61
17,142
24.00
3,287
GPT-5.5 (Medium)
69.88
793
40.20
799
91.11
1,002
46.25
557
Claude-Opus-4.6 (Medium)
59.37
-
53.40
-
78.80
-
30.00
-
Table 1: Main results on four agent tasks, including ScienceWorld, WebShop, AppWorld, and DeepSearch. Avg Tok. denotes the average number of reasoning tokens generated per trajectory. For models with up to 32B parameters, the best results are in bold and the second are underlined .
Method
ScienceWorld
WebShop
AppWorld
DeepSearch
Success ↑
Avg Tok. ↓
Success ↑
Avg Tok. ↓
Success ↑
Avg Tok. ↓
Success ↑
Avg Tok. ↓
RACE-SFT-RL
73.28
1,403
53.40
535
65.13
1,900
35.00
822
w/o Reasoning Masking in RL
72.95
609
46.00
124
56.75
3,161
32.00
2,196
w/o LoGiC in RL ( RACE-SFT+GRPO )
75.54
2,573
51.40
2,670
67.18
8,574
35.75
3,507
w/o LoGiC-RL ( RACE-SFT )
43.05
1,966
45.40
1,232
63.08
3,330
28.25
1,856
w/o LoGiC ( SFT+GRPO )
74.55
2,073
52.00
2,702
65.30
3,545
35.75
2,163
Table 2: Ablation studies of RACE. The best results are in bold .
Figure 4: Comparison of task performance and reasoning efficiency with and without LoGiC.
Figure 5: Scaling analysis with Qwen3.5-2B, 4B, and 9B across ScienceWorld, WebShop, AppWorld, and DeepSearch, showing task performance and reasoning efficiency at different model scales under no additional training, RACE-SFT, and RACE-SFT-RL.
Benchmark
RACE-SFT
RACE-SFT-RL
Skip (%)
Tok./Think
Skip (%)
Tok./Think
ScienceWorld
37.22
132
69.65
273
WebShop
7.15
144
55.72
101
AppWorld
7.74
273
52.15
302
DeepSearch
1.79
259
5.77
117
Table 3: Analysis of adaptive reasoning behavior. Skip denotes the proportion of turns that directly execute actions without generating reasoning, while Tok./Think denotes the average reasoning tokens over reasoning turns.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Example from ScienceWorld Dataset
Task Instruction
Your task is to use chemistry to create violet paint. When you are done, focus on the violet paint.
Initial Observation
This room is called the art studio. In it, you see:
a substance called air
a large cupboard. The large cupboard door is closed.
Appendix
Table 4: A scientific reasoning example on ScienceWorld illustrating the adaptive reasoning behavior enabled by RACE across different interaction turns. […] indicates omitted content. Special symbols highlight reasoning ( blue box ), tool calls ( green box ), and tool results ( yellow box ).
Example from WebShop Dataset
Task Instruction
I like traditional, old and individually wrapped albert’s chocolate ice cubes 60 count tray chocolate, and price lower than 60.00 dollars
Model Output by RACE (9B)
<think> The user wants to search for a product. Let me construct a search query based on the instruction. I’ll search for “Albert’s Chocolate Ice Cubes 60 count tray traditional old individually wrapped”. </think>
Table 5: A web shopping example on WebShop illustrating the adaptive reasoning behavior enabled by RACE across different interaction turns, including product search, browsing, and selection. […] indicates omitted content. Special symbols highlight reasoning ( blue box ), tool calls ( green box ), tool results ( yellow box ), and clickable web elements ( gray box ).
Example from AppWorld Dataset
Task Instruction
Email the birth certificate found in my file system to my husband.
User Profile
My name is: Kristin White. My personal email is [email protected] and phone number is 6017026518.
Model Output by RACE (9B)
<think> I need to email the birth certificate found in my file system to my husband. Let me start by understanding my accounts and finding the birth certificate file, then identify my husband’s contact, and draft/send the email. First, I need credentials. Let me get supervisor profile and passwords, and maybe current phone/time. Let me call supervisor_show_account_passwords and supervisor_show_profile. Also, maybe phone login requires phone number and password. I have personal email and phone number given: email [email protected], phone 6017026518. Let me get credentials. Let me call supervisor APIs. </think>
Appendix
Table 6: An application use example on AppWorld illustrating the adaptive reasoning behavior enabled by RACE across different interaction turns. […] indicates omitted content. Special symbols highlight reasoning ( blue box ), tool calls ( green box ), and tool results ( yellow box ).
Which film has the director who was born first, Reaching For The Sun or Trifling Women?
Model Output by RACE (9B)
<think> We need answer question. Need identify directors of “Reaching For The Sun” and “Trifling Women”, their birth dates, compare. Need search. Use search tool. </think>
<tool call> search(query="Reaching For The Sun film director born") </tool call>
<tool response>
Appendix
Table 7: A deep search example on DeepSearch illustrating the adaptive reasoning behavior enabled by RACE across different interaction turns, including query decomposition, evidence collection, and answer synthesis. […] indicates omitted content. Special symbols highlight reasoning ( blue box ), tool calls ( green box ), and tool results ( yellow box ), and the final answer ( red box ).