Scaling test-time computation is a powerful way to improve language-model reasoning, and is particularly appealing for small reasoning models (sRMs) that are cheap to serve. However, is additional thinking always the right operation? By intervening at intermediate reasoning states across two model families and multiple scales, we find that self-refinement largely consolidates probability mass onto solutions already reachable from the current state, rather than making new ones reachable. These interventions reveal two failure regimes: execution bottlenecks, where the correct path is reachable and reflection can recover it, and knowledge bottlenecks, where relevant external information makes it reachable. Motivated by this distinction, we introduce FlyBy, a selective querying framework, and train 4B and 8B variants to reason first, diagnose what remains unresolved, and, at a knowledge bottleneck, query stronger models whose parametric knowledge extends beyond its own. Supervised fine-tuning bootstraps a multi-depth query action, and cost-aware reinforcement learning calibrates whether to query, what to ask, and how much to spend. On 1,158 hard problems across six benchmarks, FlyBy-4B achieves 45.96% pass@8, surpassing Qwen3-14B (41.64%) at 2.7 times lower serving cost, while also exceeding Qwen3-8B in pass@1 (16.85% vs. 15.31%). Scaling to FlyBy-8B further improves pass@8 to 51.81%.
Figures & tables
Figure 1 : Overview. Left: sRM failures exhibit two bottlenecks: (a) execution bottlenecks , where further reasoning can recover a reachable solution, and (b) knowledge bottlenecks , where scarce parametric knowledge requires external information. FlyBy addresses both bottlenecks via local reasoning and selective querying; see Fig. 11 for an example. Middle: FlyBy improves the performance-cost trade-off, outperforming same-scale models at lower cost and surpassing Qwen3-14B with FlyBy -4B; see Tab. 2 for details. Right: On problems unsolved by Qwen3-4B in 16 rollouts, FlyBy -4B reaches 28.7% pass@8, surpassing Qwen3-14B.
d
Backend
Max tokens
Input / Output
1
DeepSeek-V4-Flash
128
0.094 / 0.188
2
DeepSeek-V3.2
512
0.269 / 0.400
3
DeepSeek-V4-Pro
1,536
0.435 / 0.870
Table 1 : Multi-depth query tool. API prices are USD per 1M tokens.
Model
ArXivMath
GPQA-D
SuperGPQA
ChemBench
MedXpertQA
MMLU-Pro
Avg Perf.
Avg Cost ( ↓ )
Qwen3-4B
0 8.34
35.54
24.17
25.25
15.66
18.07
21.17
10.63
Qwen3-8B
10.69
49.40
35.30
45.99
26.26
37.75
34.23
15.38
Qwen3-14B
14.46
52.63
43.85
59.30
32.82
46.78
41.64
20.35
ForkingRL
0 8.81
34.86
23.75
27.90
16.98
17.00
21.55
11.47
Search-R1
11.17
36.38
27.14
38.48
15.53
26.17
25.81
0 8.35
Query Opening
15.70
44.47
39.01
50.73
23.45
30.99
34.06
0 8.56
Table 2 : Main results. Performance on hard reasoning problems across six benchmarks. We report pass@8 and inference cost per eight rollouts in milli-dollars (m,10^{-3}$ USD). Best and second-best results are bolded and underlined , respectively. Frontier-scale DeepSeek-V4-Pro is excluded from the ranking. For Search-R1 baseline, we exclude retrieval costs due to ambiguity.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Prefill throughput (tokens/s)
Decode throughput (tokens/s)
Qwen3-4B
66,199.94
5,166.24
Qwen3-8B
40,114.82
4,056.16
Qwen3-14B
22,564.44
2,823.72
Appendix
Table 3 : Measured inference throughput for Qwen3 models on a single NVIDIA H200 GPU. We report the prefill and decoding throughput used to estimate local inference cost in Eq. 8 .
Parameter
Value
Supervised Fine-Tuning (SFT)
Fine-tuning method
Full fine-tuning
Optimizer
AdamW
Learning rate
2×10−6
LR scheduler
Cosine
Warmup steps
8
Appendix
Table 4 : Hyperparameters for SFT and RL.
Benchmark
Task coverage
Evaluation subset
Eval pool
Hard
ArXivMath
Research-level mathematical reasoning from arXiv papers.
All questions from the December 2025–June 2026 monthly releases, pooled across months.
232
207
GPQA-Diamond
Graduate-level reasoning in biology, chemistry, and physics.
The complete Diamond subset.
198
80
SuperGPQA
Graduate-level knowledge and reasoning across academic disciplines.
A held-out sample of 500 questions from Science, Engineering, Medicine, and Agronomy, stratified by discipline and difficulty and disjoint from our training split.
500
252
MedXpertQA
Expert-level medical knowledge and clinical reasoning.
500 questions randomly sampled from the 2,450-question Text/test split.
500
410
MMLU-Pro
Broad academic knowledge and reasoning across 14 subject areas.
500 questions randomly sampled from the official test split.
500
129
ChemBench
Chemistry and materials science knowledge and reasoning.
All single-answer multiple-choice questions from Organic Chemistry (384), Materials Science (57), and Inorganic Chemistry (49), pooled across the three domains.
490
80
Appendix
Table 5 : Details of the evaluation benchmarks. We report the task coverage, evaluation subset, and number of questions before and after hard-subset selection. The hard subset contains questions answered correctly by Qwen3-4B in at most 4 of 16 rollouts and is fixed across all evaluated methods. All random sampling uses seed 42.
Model
ArXivMath
GPQA-D
SuperGPQA
ChemBench
MedXpertQA
MMLU-Pro
Avg Perf.
Avg Cost ( ↓ )
Qwen3-4B
2.26
7.89
6.13
5.31
3.46
3.97
4.84
1.33
Qwen3-8B
2.90
17.58
16.29
25.86
10.95
18.31
15.31
1.92
Qwen3-14B
4.50
26.02
25.22
37.66
13.61
25.15
22.03
2.54
ForkingRL
1.63
7.58
5.46
5.39
3.67
3.39
4.52
1.43
Search-R1
3.89
10.08
8.85
10.55
3.63
6.35
7.22
1.04
Query Opening
5.01
15.78
15.53
26.56
8.63
12.60
14.02
1.07
Appendix
Table 6 : Pass@1 results. Performance on hard reasoning problems across six benchmarks. We report pass@1 and inference cost per rollout in milli-dollars (m,10^{-3}$ USD). Best and second-best results are bolded and underlined , respectively. Frontier-scale DeepSeek-V4-Pro is excluded from the ranking. For Search-R1 baseline, we exclude retrieval costs due to ambiguity.
Model
Query Backend
ArXivMath
GPQA-D
SuperGPQA
ChemBench
MedXpertQA
MMLU-Pro
Avg Perf.
Avg Cost ( ↓ )
Qwen3-14B
-
14.46
52.63
43.85
59.30
32.82
46.78
41.64
20.35
FlyBy -4B
DeepSeek
21.78
60.01
45.65
70.73
35.37
42.21
45.96
7.42
GPT-5.6 Luna
22.39
55.22
43.53
64.53
35.59
38.02
43.21
8.86
HY3
22.19
51.04
43.62
68.06
34.26
44.61
43.96
8.01
Appendix
Table 7 : Generalization across query backends. We transfer the learned querying policy to alternative query backends without retraining. DeepSeek denotes the training-time backend configuration, which uses different models across query depths, while GPT-5.6 Luna and HY3 replace this backend only at inference time. We report pass@8 and inference cost over 8 rollouts in milli-dollars (m,10^{-3}$ USD). Best and second-best results are bolded and underlined , respectively.
Input
Accuracy
Δ
Problem only
10.20
–
Problem + observation
12.16
+1.96
Appendix
Table 8 : Analysis of answer leakage from tool observations. We evaluate Qwen3-4B with thinking disabled.