Scaling test-time computation is a powerful way to improve language-model reasoning, and is particularly appealing for small reasoning models (sRMs) that are cheap to serve. However, is additional thinking always the right operation? By intervening at intermediate reasoning states across two model families and multiple scales, we find that self-refinement largely consolidates probability mass onto solutions already reachable from the current state, rather than making new ones reachable. These interventions reveal two failure regimes: execution bottlenecks, where the correct path is reachable and reflection can recover it, and knowledge bottlenecks, where relevant external information makes it reachable. Motivated by this distinction, we introduce FlyBy, a selective querying framework, and train 4B and 8B variants to reason first, diagnose what remains unresolved, and, at a knowledge bottleneck, query stronger models whose parametric knowledge extends beyond its own. Supervised fine-tuning bootstraps a multi-depth query action, and cost-aware reinforcement learning calibrates whether to query, what to ask, and how much to spend. On 1,158 hard problems across six benchmarks, FlyBy-4B achieves 45.96% pass@8, surpassing Qwen3-14B (41.64%) at 2.7 times lower serving cost, while also exceeding Qwen3-8B in pass@1 (16.85% vs. 15.31%). Scaling to FlyBy-8B further improves pass@8 to 51.81%.
Figures & tables
Figure 1 : Overview. Left: sRM failures exhibit two bottlenecks: (a) execution bottlenecks , where further reasoning can recover a reachable solution, and (b) knowledge bottlenecks , where scarce parametric knowledge requires external information. FlyBy addresses both bottlenecks via local reasoning and selective querying; see Fig. 11 for an example. Middle: FlyBy improves the performance-cost trade-off, outperforming same-scale models at lower cost and surpassing Qwen3-14B with FlyBy -4B; see Tab. 2 for details. Right: On problems unsolved by Qwen3-4B in 16 rollouts, FlyBy -4B reaches 28.7% pass@8, surpassing Qwen3-14B.
d
Backend
Max tokens
Input / Output
1
DeepSeek-V4-Flash
128
0.094 / 0.188
2
DeepSeek-V3.2
512
0.269 / 0.400
3
DeepSeek-V4-Pro
1,536
0.435 / 0.870
Table 1 : Multi-depth query tool. API prices are USD per 1M tokens.
Model
ArXivMath
GPQA-D
SuperGPQA
ChemBench
MedXpertQA
MMLU-Pro
Avg Perf.
Avg Cost ( ↓ )
Qwen3-4B
0 8.34
35.54
24.17
25.25
15.66
18.07
21.17
10.63
Qwen3-8B
10.69
49.40
35.30
45.99
26.26
37.75
34.23
15.38
Qwen3-14B
14.46
52.63
43.85
59.30
32.82
46.78
41.64
20.35
ForkingRL
0 8.81
34.86
23.75
27.90
16.98
17.00
21.55
11.47
Search-R1
11.17
36.38
27.14
38.48
15.53
26.17
25.81
0 8.35
Query Opening
15.70
44.47
39.01
50.73
23.45
30.99
34.06
0 8.56
Table 2 : Main results. Performance on hard reasoning problems across six benchmarks. We report pass@8 and inference cost per eight rollouts in milli-dollars (m,10^{-3}$ USD). Best and second-best results are bolded and underlined , respectively. Frontier-scale DeepSeek-V4-Pro is excluded from the ranking. For Search-R1 baseline, we exclude retrieval costs due to ambiguity.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Prefill throughput (tokens/s)
Decode throughput (tokens/s)
Qwen3-4B
66,199.94
5,166.24
Qwen3-8B
40,114.82
4,056.16
Qwen3-14B
22,564.44
2,823.72
Appendix
Table 3 : Measured inference throughput for Qwen3 models on a single NVIDIA H200 GPU. We report the prefill and decoding throughput used to estimate local inference cost in Eq. 8 .
Parameter
Value
Supervised Fine-Tuning (SFT)
Fine-tuning method
Full fine-tuning
Optimizer
AdamW
Learning rate
2×10−6
LR scheduler
Cosine
Warmup steps
8
Appendix
Table 4 : Hyperparameters for SFT and RL.
Benchmark
Task coverage
Evaluation subset
Eval pool
Hard
ArXivMath
Research-level mathematical reasoning from arXiv papers.
All questions from the December 2025–June 2026 monthly releases, pooled across months.
232
207
GPQA-Diamond
Graduate-level reasoning in biology, chemistry, and physics.
The complete Diamond subset.
198
80
SuperGPQA
Graduate-level knowledge and reasoning across academic disciplines.
A held-out sample of 500 questions from Science, Engineering, Medicine, and Agronomy, stratified by discipline and difficulty and disjoint from our training split.
500
252
MedXpertQA
Expert-level medical knowledge and clinical reasoning.
500 questions randomly sampled from the 2,450-question Text/test split.
500
410
MMLU-Pro
Broad academic knowledge and reasoning across 14 subject areas.
500 questions randomly sampled from the official test split.
500
129
ChemBench
Chemistry and materials science knowledge and reasoning.
All single-answer multiple-choice questions from Organic Chemistry (384), Materials Science (57), and Inorganic Chemistry (49), pooled across the three domains.
490
80
Appendix
Table 5 : Details of the evaluation benchmarks. We report the task coverage, evaluation subset, and number of questions before and after hard-subset selection. The hard subset contains questions answered correctly by Qwen3-4B in at most 4 of 16 rollouts and is fixed across all evaluated methods. All random sampling uses seed 42.
Model
ArXivMath
GPQA-D
SuperGPQA
ChemBench
MedXpertQA
MMLU-Pro
Avg Perf.
Avg Cost ( ↓ )
Qwen3-4B
2.26
7.89
6.13
5.31
3.46
3.97
4.84
1.33
Qwen3-8B
2.90
17.58
16.29
25.86
10.95
18.31
15.31
1.92
Qwen3-14B
4.50
26.02
25.22
37.66
13.61
25.15
22.03
2.54
ForkingRL
1.63
7.58
5.46
5.39
3.67
3.39
4.52
1.43
Search-R1
3.89
10.08
8.85
10.55
3.63
6.35
7.22
1.04
Query Opening
5.01
15.78
15.53
26.56
8.63
12.60
14.02
1.07
Appendix
Table 6 : Pass@1 results. Performance on hard reasoning problems across six benchmarks. We report pass@1 and inference cost per rollout in milli-dollars (m,10^{-3}$ USD). Best and second-best results are bolded and underlined , respectively. Frontier-scale DeepSeek-V4-Pro is excluded from the ranking. For Search-R1 baseline, we exclude retrieval costs due to ambiguity.
Model
Query Backend
ArXivMath
GPQA-D
SuperGPQA
ChemBench
MedXpertQA
MMLU-Pro
Avg Perf.
Avg Cost ( ↓ )
Qwen3-14B
-
14.46
52.63
43.85
59.30
32.82
46.78
41.64
20.35
FlyBy -4B
DeepSeek
21.78
60.01
45.65
70.73
35.37
42.21
45.96
7.42
GPT-5.6 Luna
22.39
55.22
43.53
64.53
35.59
38.02
43.21
8.86
HY3
22.19
51.04
43.62
68.06
34.26
44.61
43.96
8.01
Appendix
Table 7 : Generalization across query backends. We transfer the learned querying policy to alternative query backends without retraining. DeepSeek denotes the training-time backend configuration, which uses different models across query depths, while GPT-5.6 Luna and HY3 replace this backend only at inference time. We report pass@8 and inference cost over 8 rollouts in milli-dollars (m,10^{-3}$ USD). Best and second-best results are bolded and underlined , respectively.
Input
Accuracy
Δ
Problem only
10.20
–
Problem + observation
12.16
+1.96
Appendix
Table 8 : Analysis of answer leakage from tool observations. We evaluate Qwen3-4B with thinking disabled.
Large Reasoning Models (LRMs) improve performance by allocating additional inference-time compute to generate extended chain-of-thought reasoning. However, recent studies reveal that sequential test-time scaling often yields diminishing or even negative returns, as longer traces exhibit increased uncertainty, error compounding, and drift from the original problem. We propose ThinkRetrieve, a test-time scaling framework that augments the reasoning traces of LRMs with dynamically retrieved solved examples at each reasoning step. Given an external corpus of problems paired with step-by-step solutions, ThinkRetrieve retrieves relevant exemplars at each intermediate step and injects them directly into the thinking trace, providing the model with guidance on how to reason rather than merely what facts are relevant. Experiments across five reasoning models (1.5B--8B parameters) on GSM-8K, MATH-500, AIME 2025, and SciQ demonstrate that ThinkRetrieve consistently improves accuracy over standard test-time scaling, with relative gains of up to 60% on AIME 2025.
Large reasoning models (LRMs) improve problem solving through extended reasoning, but often misallocate test-time compute. Existing efficiency methods reduce cost by compressing reasoning traces or conditioning budget on perceived difficulty, yet largely overlook solvability. As a result, they may spend large budgets on queries beyond the model's capability while compressing hard-but-solvable queries that require deeper reasoning. In this work, we formulate adaptive reasoning as a computational investment under uncertainty, where budget should follow the expected return of reasoning rather than perceived difficulty alone. To instantiate this principle, we propose Budget-Efficient Thinking (BET), a two-stage framework that combines behavioral cold-start with GRPO under an investment-cost-aware reward. By aligning solve-or-fold decisions with rollout-derived solvability, BET learns three behaviors: (1) short solve, answering easy queries concisely; (2) nice fold, abstaining early when continued reasoning has near-zero expected return; and (3) hero call, preserving sufficient compute for hard-but-solvable queries. Across seven benchmarks and three base models, BET reduces reasoning tokens by ~55% on average while achieving overall performance improvements, and transfers zero-shot from mathematical reasoning to scientific QA and logical reasoning with comparable efficiency gains.
Zhaomeng Zhou, Lan Zhang, Junyang Wang +2
University of Science and Technology of China · The Chinese University of Hong Kong
Recent advances in large language models (LLMs) have demonstrated that reinforcement fine-tuning of pretrained base models can lead to significant gains in reasoning performance at inference time. In this work, we theoretically analyze why reinforcement fine-tuning induces better reasoning ability than purely supervised fine-tuning (SFT) methods. We model chain-of-thought (CoT) reasoning as a pathfinding problem on graphs and compare the popular method of reinforcement learning with verifiable rewards (RLVR) against traditional SFT. We prove that SFT, when trained on golden shortest paths without negative examples, fails to learn how to efficiently backtrack. In contrast, an RLVR-trained model can learn how to efficiently backtrack from dead ends using only outcome reward. This leads to an exponential separation in inference-time compute between the two methods, and demonstrates that RLVR leads the model to learn the location of difficult decisions in a reasoning chain, ultimately allowing for better allocation of inference-time compute. Finally, we show that the reasoning traces of an RLVR model can be distilled to train a base model to backtrack efficiently as well.
Stanley Wei, Juno Kim
1Princeton University · University of California, Berkeley.