Inference scaling has been shown to improve large language model (LLM) performance, and this principle naturally extends to autonomous LLM agents through increased search budgets, which we refer to as search scaling. Although prior work has characterized the mechanisms, scaling behavior, and performance limits of LLM inference scaling, much less is known about these questions in autonomous research. Therefore, we investigate how search scaling affects research performance and what mechanisms drive these gains using 50 quantitative factor-mining tasks grounded in financial research reports. Each task requires an agent to carry out an end-to-end research loop, from interpreting a hypothesis and implementing it in code to evaluating and iteratively refining the resulting factor. Across nine models, we examine how model capability, search depth, and search organization shape factor quality by tracing performance across varying budgets, transferring intermediate research states between models, and comparing different search strategies. We find that (1) initial performance is more strongly associated with model capability, while deeper search can narrow cross-model gaps; (2) model grafting shows that the early research state materially shapes final performance; and (3) parallel search outperforms sequential search under the same iteration budget, consistent with benefits from broader coverage of the search space. Further trajectory analysis shows that higher-performing models more effectively diagnose failures, revise search directions, and preserve the intended economic hypothesis when selecting candidates. These findings suggest that future progress in autonomous research will require stronger models together with adaptive policies for deploying test-time computation throughout the research process.
Figures & tables
Score
Sharpe
Monotonicity
Alignment
Cost per task (RMB)
Model
b=3
b=6
b=10
b=3
b=6
b=10
b=3
b=6
b=10
b=3
b=6
b=10
b=3
b=6
b=10
Claude Opus 5
75.45
87.69
92.50
2.008
2.334
2.661
0.952
0.964
0.989
4.04
4.04
3.70
53.03
68.71
116.94
GLM-5.3
67.15
83.12
92.65
1.787
2.212
2.568
0.948
0.970
0.978
4.24
4.24
3.84
17.14
25.04
39.61
DeepSeek-V4.1-Flash
65.35
73.51
83.44
1.739
1.956
2.266
0.942
0.969
0.983
4.32
4.32
3.92
1.43
2.51
3.25
DeepSeek-V4-Pro
53.68
60.59
71.40
1.429
1.612
2.054
0.940
0.965
0.963
4.20
4.20
3.70
17.70
31.85
56.37
Doubao Seed 2.1 Pro
49.58
59.95
67.97
1.319
1.595
1.827
0.953
0.961
0.967
4.10
4.10
3.96
7.36
9.29
14.31
Table 1: Results across research depths.
Research schedule
Score
Sharpe
Monotonicity
Alignment
Turnover
IR
Qwen3.7-Plus (3) → GLM-5.3 (7)
72.28
2.161
0.976
3.56
8.05%
2.125
GLM-5.3 (3) → Qwen3.7-Plus (7)
76.41
2.210
0.970
3.68
7.26%
2.180
Table 2: Model grafting under a ten-iteration total budget.
Score
Sharpe
Model
Parallel
Sequential
Parallel
Sequential
DeepSeek-V4.1-Flash
89.80
83.44
2.390
2.266
Doubao Seed 2.1 Pro
73.00
67.97
1.943
1.827
MiniMax-M3
61.57
55.99
1.638
1.611
Table 3: Sequential and parallel search under a ten-iteration total budget.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Signal family
1 min
5 min
30 min
Daily
Total
Volatility
2
-
-
3
5
Momentum
3
-
1
2
6
Liquidity
3
-
-
3
6
Microstructure
2
1
-
1
4
Prospect theory
1
3
-
2
6
Beta
2
-
-
1
3
Appendix
Table A.1: Technical tasks by signal family and input frequency.
Signal family
5 min
Daily
Fundamentals only
Total
Valuation
-
8
-
8
Profitability
-
4
-
4
Earnings surprise
1
1
-
2
Growth
-
1
1
2
Dividends
-
2
-
2
Quality
-
1
-
1
Appendix
Table A.2: Fundamental tasks by signal family and input frequency.
Setting
Value
Research tasks
50 fixed objectives: 20 fundamental, 30 technical
Score aggregation
All 50 tasks; failed or invalid tasks contribute zero scoring inputs
Research iterations
One continuous 10-iteration trajectory per model–task pair; a checkpoint after every iteration
Independent repetitions
5 complete runs per model; reported Scores average the 5 run-level Scores
Reasoning effort
High where supported
Session timeout
8 hours
Appendix
Table D.1: Main experimental settings.
Model
AAII
Reasoning Effort
Claude Opus 5
48
high
GLM-5.3
45
max
DeepSeek-V4.1-Flash
39
max
DeepSeek-V4-Pro
36
max
GLM-5.2
34
max
MiniMax-M3
29
reasoning
Appendix
Table D.2: External capability values used in the empirical scaling fit.
No.
Task ID
Sharpe
Monotonicity
Alignment
b=3
b=6
b=10
b=3
b=6
b=10
b=3
b=6
b=10
1
accr_excess_gap
2.688
2.688
2.736
1.000
1.000
0.952
4
4
4
2
anchor_fit_value
1.791
2.170
2.547
1.000
1.000
1.000
4
4
3
3
aog_vwap30_deq80
1.538
1.538
1.538
0.939
0.939
0.903
4
4
3
4
cohort_excess_pulse_rev
2.461
2.461
3.905
1.000
1.000
1.000
4
4
3
5
droa_early_pit
3.935
3.935
3.935
0.988
0.988
0.988
5
5
5
Appendix
Table E.1: Task-level metrics for Claude Opus 5 at iterations 3, 6, and 10.
No.
Task ID
Sharpe
Monotonicity
Alignment
b=3
b=6
b=10
b=3
b=6
b=10
b=3
b=6
b=10
1
accr_excess_gap
2.516
2.516
2.622
0.952
0.952
0.952
4
4
5
2
anchor_fit_value
2.546
3.623
3.623
0.988
0.988
0.988
5
5
4
3
aog_vwap30_deq80
0.913
1.700
2.353
0.891
0.988
1.000
4
4
4
4
cohort_excess_pulse_rev
1.387
2.653
3.034
1.000
1.000
1.000
4
4
4
5
droa_early_pit
2.790
4.291
4.341
1.000
1.000
0.988
4
4
4
Appendix
Table E.2: Task-level metrics for GLM-5.3 at iterations 3, 6, and 10.
No.
Task ID
Sharpe
Monotonicity
Alignment
b=3
b=6
b=10
b=3
b=6
b=10
b=3
b=6
b=10
1
accr_excess_gap
2.298
2.298
2.298
0.964
0.964
0.964
4
4
4
2
anchor_fit_value
2.084
3.124
3.124
1.000
1.000
0.988
4
4
3
3
aog_vwap30_deq80
0.789
0.789
1.112
0.964
0.964
0.988
4
4
4
4
cohort_excess_pulse_rev
1.397
1.397
2.753
1.000
1.000
1.000
4
4
3
5
droa_early_pit
4.162
5.063
5.063
0.988
1.000
1.000
4
4
4
Appendix
Table E.3: Task-level metrics for DeepSeek-V4.1-Flash at iterations 3, 6, and 10.
No.
Task ID
Sharpe
Monotonicity
Alignment
b=3
b=6
b=10
b=3
b=6
b=10
b=3
b=6
b=10
1
accr_excess_gap
1.721
1.887
2.148
0.988
0.988
0.988
4
4
4
2
anchor_fit_value
2.137
2.137
1.516
1.000
1.000
1.000
4
4
4
3
aog_vwap30_deq80
0.601
0.601
0.149
0.842
0.842
0.333
5
5
2
4
cohort_excess_pulse_rev
1.362
1.362
2.234
0.964
0.964
1.000
4
4
4
5
droa_early_pit
1.456
1.628
4.406
0.964
1.000
1.000
5
5
4
Appendix
Table E.4: Task-level metrics for DeepSeek-V4-Pro at iterations 3, 6, and 10.
No.
Task ID
Sharpe
Monotonicity
Alignment
b=3
b=6
b=10
b=3
b=6
b=10
b=3
b=6
b=10
1
accr_excess_gap
1.953
2.313
2.340
0.952
0.964
0.964
4
4
3
2
anchor_fit_value
1.333
1.629
2.015
1.000
1.000
1.000
4
4
3
3
aog_vwap30_deq80
3.087
3.090
3.235
1.000
1.000
1.000
5
5
4
4
cohort_excess_pulse_rev
1.278
1.816
2.074
0.976
1.000
1.000
5
5
4
5
droa_early_pit
3.026
3.192
3.361
1.000
1.000
1.000
3
3
4
Appendix
Table E.5: Task-level metrics for Doubao Seed 2.1 Pro at iterations 3, 6, and 10.
No.
Task ID
Sharpe
Monotonicity
Alignment
b=3
b=6
b=10
b=3
b=6
b=10
b=3
b=6
b=10
1
accr_excess_gap
2.168
2.168
1.266
0.976
0.976
0.976
4
4
4
2
anchor_fit_value
1.014
1.385
1.588
0.988
1.000
1.000
4
4
4
3
aog_vwap30_deq80
1.609
1.609
0.670
0.927
0.927
0.939
4
4
4
4
cohort_excess_pulse_rev
1.448
1.448
1.519
0.988
0.988
1.000
4
4
4
5
droa_early_pit
2.095
2.095
2.821
1.000
1.000
0.976
3
3
3
Appendix
Table E.6: Task-level metrics for GLM-5.2 at iterations 3, 6, and 10.
No.
Task ID
Sharpe
Monotonicity
Alignment
b=3
b=6
b=10
b=3
b=6
b=10
b=3
b=6
b=10
1
accr_excess_gap
1.256
1.262
1.423
0.915
0.976
0.976
5
2
2
2
anchor_fit_value
1.944
1.993
1.772
1.000
1.000
1.000
4
4
4
3
aog_vwap30_deq80
-0.035
-3.320
-0.265
0.745
0.879
0.915
5
4
4
4
cohort_excess_pulse_rev
1.267
1.431
1.398
0.976
0.964
0.988
4
5
4
5
droa_early_pit
0.887
0.810
1.148
0.976
0.988
0.939
5
3
4
Appendix
Table E.7: Task-level metrics for MiniMax-M3 at iterations 3, 6, and 10.
No.
Task ID
Sharpe
Monotonicity
Alignment
b=3
b=6
b=10
b=3
b=6
b=10
b=3
b=6
b=10
1
accr_excess_gap
0.016
0.745
2.615
0.491
0.539
1.000
3
3
0
2
anchor_fit_value
1.502
1.502
1.689
0.988
0.988
0.988
2
2
3
3
aog_vwap30_deq80
0.710
0.710
0.710
0.842
0.842
0.842
4
4
4
4
cohort_excess_pulse_rev
1.658
1.658
2.404
0.988
1.000
1.000
4
4
3
5
droa_early_pit
1.865
2.543
2.851
0.964
0.988
1.000
3
3
2
Appendix
Table E.8: Task-level metrics for Qwen3.7-Plus at iterations 3, 6, and 10.
No.
Task ID
Sharpe
Monotonicity
Alignment
b=3
b=6
b=10
b=3
b=6
b=10
b=3
b=6
b=10
1
accr_excess_gap
-0.700
-0.524
-0.524
0.250
0.273
0.273
4
4
3
2
anchor_fit_value
1.741
1.741
-1.628
1.000
1.000
0.661
3
3
4
3
aog_vwap30_deq80
0.334
0.334
-0.092
0.709
0.915
1.000
5
5
2
4
cohort_excess_pulse_rev
1.360
1.360
1.744
0.976
0.976
1.000
5
5
4
5
droa_early_pit
1.548
1.548
4.489
0.988
0.988
1.000
3
3
4
Appendix
Table E.9: Task-level metrics for DeepSeek-V3.1 at iterations 3, 6, and 10.
Reinforcement Learning (RL) has emerged as a powerful training paradigm for LLM-based agents. However, scaling agentic RL for deep research remains constrained by two coupled challenges: hand-crafted synthetic data fails to elicit genuine real-world search capabilities, and real-world search dependency during RL training introduces instability and prohibitive cost, which limits the scalability of Agentic RL. LiteResearcher is a training framework that makes Agentic RL scalable: by constructing a lite virtual world that mirrors real-world search dynamics, we enable a continuously improving training recipe that empowers a tiny search agent to outperform large-scale open-source and commercial models (e.g., Tongyi DeepResearch and Claude-4.5 Sonnet). Specifically, on common benchmarks such as GAIA and Xbench, our LiteResearcher-4B achieves open-source state-of-the-art results of 71.3% and 78.0% respectively, demonstrating that scalable RL training is a key enabler for Deep Research Agents.
Wanli Li, Bince Qu, Bo Pan +5
Zhejiang University · Simplex AI · The Hong Kong Polytechnic University
Deep search capabilities have become an indispensable competency for frontier Large Language Model (LLM) agents, yet their development remains dominated by industrial giants. The typical industry recipe involves a highly resource-intensive pipeline spanning pre-training, continual pre-training (CPT), supervised fine-tuning (SFT), and reinforcement learning (RL). In this report, we show that when fueled with informative and high-difficulty trajectories, a simple SFT approach could be surprisingly powerful for training frontier search agents. By introducing three simple data synthesis modifications: scaling knowledge graph size for richer exploration, expanding the tool set size for broader functionality, and strict low-step filtering, we establish a stronger baseline. Trained on merely 10.6k data points, our OpenSeeker-v2 achieves state-of-the-art performance across 4 benchmarks (30B-sized agents with ReAct paradigm): 46.0% on BrowseComp, 58.1% on BrowseComp-ZH, 34.6% on Humanity's Last Exam, and 78.0% on xbench, surpassing even Tongyi DeepResearch trained with heavy CPT+SFT+RL pipeline, which achieves 43.4%, 46.7%, 32.9%, and 75.0%, respectively. Notably, OpenSeeker-v2 represents the first state-of-the-art search agent within its model scale and paradigm to be developed by a purely academic team using only SFT. We are excited to open-source the OpenSeeker-v2 model weights and share our simple yet effective findings to make frontier search agent research more accessible to the community.
Large language model based search agents increasingly adopt multi-agent architectures in which a main agent decomposes a complex question into sub-queries and dispatches them to parallel sub-agents. However, existing systems instantiate all roles from a single model of identical scale, leaving open how model capacity should be distributed across roles. We factorize hierarchical search into three roles: a delegation role responsible for task decomposition, an execution role responsible for retrieval and evidence extraction, and an answer generation role held fixed as a confound control. We then conduct controlled capacity sweeps along the delegation and execution axes on five multi-hop QA benchmarks. The experiments yield three findings. First, role factorization consistently outperforms a single-agent baseline, improving exact match from 4.5 to 8.6 points across six model scales. Second, capacity sensitivity is asymmetric: scaling the delegation backbone improves EM by ~11 points, whereas scaling the execution sub-agent moves EM by only ~2.6 points, identifying decomposition as the capability bottleneck. Third, a 1.7B-parameter executor trained via quality-filtered trajectory distillation matches a frontier sub-agent in accuracy while consuming 37% fewer sub-agent tokens, advancing the Pareto frontier. These results suggest a concrete recipe for building hierarchical search agents: concentrate capacity at delegation and downsize execution without sacrificing accuracy. Our code is available at https://github.com/QinnanCai0115/role-factorized-search.
Qinnan Cai, Yibo Zhao, Xiang Li
School of Data Science and Engineering East China Normal University