Inference scaling has been shown to improve large language model (LLM) performance, and this principle naturally extends to autonomous LLM agents through increased search budgets, which we refer to as search scaling. Although prior work has characterized the mechanisms, scaling behavior, and performance limits of LLM inference scaling, much less is known about these questions in autonomous research. Therefore, we investigate how search scaling affects research performance and what mechanisms drive these gains using 50 quantitative factor-mining tasks grounded in financial research reports. Each task requires an agent to carry out an end-to-end research loop, from interpreting a hypothesis and implementing it in code to evaluating and iteratively refining the resulting factor. Across nine models, we examine how model capability, search depth, and search organization shape factor quality by tracing performance across varying budgets, transferring intermediate research states between models, and comparing different search strategies. We find that (1) initial performance is more strongly associated with model capability, while deeper search can narrow cross-model gaps; (2) model grafting shows that the early research state materially shapes final performance; and (3) parallel search outperforms sequential search under the same iteration budget, consistent with benefits from broader coverage of the search space. Further trajectory analysis shows that higher-performing models more effectively diagnose failures, revise search directions, and preserve the intended economic hypothesis when selecting candidates. These findings suggest that future progress in autonomous research will require stronger models together with adaptive policies for deploying test-time computation throughout the research process.
Figures & tables
Score
Sharpe
Monotonicity
Alignment
Cost per task (RMB)
Model
b=3
b=6
b=10
b=3
b=6
b=10
b=3
b=6
b=10
b=3
b=6
b=10
b=3
b=6
b=10
Claude Opus 5
75.45
87.69
92.50
2.008
2.334
2.661
0.952
0.964
0.989
4.04
4.04
3.70
53.03
68.71
116.94
GLM-5.3
67.15
83.12
92.65
1.787
2.212
2.568
0.948
0.970
0.978
4.24
4.24
3.84
17.14
25.04
39.61
DeepSeek-V4.1-Flash
65.35
73.51
83.44
1.739
1.956
2.266
0.942
0.969
0.983
4.32
4.32
3.92
1.43
2.51
3.25
DeepSeek-V4-Pro
53.68
60.59
71.40
1.429
1.612
2.054
0.940
0.965
0.963
4.20
4.20
3.70
17.70
31.85
56.37
Doubao Seed 2.1 Pro
49.58
59.95
67.97
1.319
1.595
1.827
0.953
0.961
0.967
4.10
4.10
3.96
7.36
9.29
14.31
Table 1: Results across research depths.
Research schedule
Score
Sharpe
Monotonicity
Alignment
Turnover
IR
Qwen3.7-Plus (3) → GLM-5.3 (7)
72.28
2.161
0.976
3.56
8.05%
2.125
GLM-5.3 (3) → Qwen3.7-Plus (7)
76.41
2.210
0.970
3.68
7.26%
2.180
Table 2: Model grafting under a ten-iteration total budget.
Score
Sharpe
Model
Parallel
Sequential
Parallel
Sequential
DeepSeek-V4.1-Flash
89.80
83.44
2.390
2.266
Doubao Seed 2.1 Pro
73.00
67.97
1.943
1.827
MiniMax-M3
61.57
55.99
1.638
1.611
Table 3: Sequential and parallel search under a ten-iteration total budget.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Signal family
1 min
5 min
30 min
Daily
Total
Volatility
2
-
-
3
5
Momentum
3
-
1
2
6
Liquidity
3
-
-
3
6
Microstructure
2
1
-
1
4
Prospect theory
1
3
-
2
6
Beta
2
-
-
1
3
Appendix
Table A.1: Technical tasks by signal family and input frequency.
Signal family
5 min
Daily
Fundamentals only
Total
Valuation
-
8
-
8
Profitability
-
4
-
4
Earnings surprise
1
1
-
2
Growth
-
1
1
2
Dividends
-
2
-
2
Quality
-
1
-
1
Appendix
Table A.2: Fundamental tasks by signal family and input frequency.
Setting
Value
Research tasks
50 fixed objectives: 20 fundamental, 30 technical
Score aggregation
All 50 tasks; failed or invalid tasks contribute zero scoring inputs
Research iterations
One continuous 10-iteration trajectory per model–task pair; a checkpoint after every iteration
Independent repetitions
5 complete runs per model; reported Scores average the 5 run-level Scores
Reasoning effort
High where supported
Session timeout
8 hours
Appendix
Table D.1: Main experimental settings.
Model
AAII
Reasoning Effort
Claude Opus 5
48
high
GLM-5.3
45
max
DeepSeek-V4.1-Flash
39
max
DeepSeek-V4-Pro
36
max
GLM-5.2
34
max
MiniMax-M3
29
reasoning
Appendix
Table D.2: External capability values used in the empirical scaling fit.
No.
Task ID
Sharpe
Monotonicity
Alignment
b=3
b=6
b=10
b=3
b=6
b=10
b=3
b=6
b=10
1
accr_excess_gap
2.688
2.688
2.736
1.000
1.000
0.952
4
4
4
2
anchor_fit_value
1.791
2.170
2.547
1.000
1.000
1.000
4
4
3
3
aog_vwap30_deq80
1.538
1.538
1.538
0.939
0.939
0.903
4
4
3
4
cohort_excess_pulse_rev
2.461
2.461
3.905
1.000
1.000
1.000
4
4
3
5
droa_early_pit
3.935
3.935
3.935
0.988
0.988
0.988
5
5
5
Appendix
Table E.1: Task-level metrics for Claude Opus 5 at iterations 3, 6, and 10.
No.
Task ID
Sharpe
Monotonicity
Alignment
b=3
b=6
b=10
b=3
b=6
b=10
b=3
b=6
b=10
1
accr_excess_gap
2.516
2.516
2.622
0.952
0.952
0.952
4
4
5
2
anchor_fit_value
2.546
3.623
3.623
0.988
0.988
0.988
5
5
4
3
aog_vwap30_deq80
0.913
1.700
2.353
0.891
0.988
1.000
4
4
4
4
cohort_excess_pulse_rev
1.387
2.653
3.034
1.000
1.000
1.000
4
4
4
5
droa_early_pit
2.790
4.291
4.341
1.000
1.000
0.988
4
4
4
Appendix
Table E.2: Task-level metrics for GLM-5.3 at iterations 3, 6, and 10.
No.
Task ID
Sharpe
Monotonicity
Alignment
b=3
b=6
b=10
b=3
b=6
b=10
b=3
b=6
b=10
1
accr_excess_gap
2.298
2.298
2.298
0.964
0.964
0.964
4
4
4
2
anchor_fit_value
2.084
3.124
3.124
1.000
1.000
0.988
4
4
3
3
aog_vwap30_deq80
0.789
0.789
1.112
0.964
0.964
0.988
4
4
4
4
cohort_excess_pulse_rev
1.397
1.397
2.753
1.000
1.000
1.000
4
4
3
5
droa_early_pit
4.162
5.063
5.063
0.988
1.000
1.000
4
4
4
Appendix
Table E.3: Task-level metrics for DeepSeek-V4.1-Flash at iterations 3, 6, and 10.
No.
Task ID
Sharpe
Monotonicity
Alignment
b=3
b=6
b=10
b=3
b=6
b=10
b=3
b=6
b=10
1
accr_excess_gap
1.721
1.887
2.148
0.988
0.988
0.988
4
4
4
2
anchor_fit_value
2.137
2.137
1.516
1.000
1.000
1.000
4
4
4
3
aog_vwap30_deq80
0.601
0.601
0.149
0.842
0.842
0.333
5
5
2
4
cohort_excess_pulse_rev
1.362
1.362
2.234
0.964
0.964
1.000
4
4
4
5
droa_early_pit
1.456
1.628
4.406
0.964
1.000
1.000
5
5
4
Appendix
Table E.4: Task-level metrics for DeepSeek-V4-Pro at iterations 3, 6, and 10.
No.
Task ID
Sharpe
Monotonicity
Alignment
b=3
b=6
b=10
b=3
b=6
b=10
b=3
b=6
b=10
1
accr_excess_gap
1.953
2.313
2.340
0.952
0.964
0.964
4
4
3
2
anchor_fit_value
1.333
1.629
2.015
1.000
1.000
1.000
4
4
3
3
aog_vwap30_deq80
3.087
3.090
3.235
1.000
1.000
1.000
5
5
4
4
cohort_excess_pulse_rev
1.278
1.816
2.074
0.976
1.000
1.000
5
5
4
5
droa_early_pit
3.026
3.192
3.361
1.000
1.000
1.000
3
3
4
Appendix
Table E.5: Task-level metrics for Doubao Seed 2.1 Pro at iterations 3, 6, and 10.
No.
Task ID
Sharpe
Monotonicity
Alignment
b=3
b=6
b=10
b=3
b=6
b=10
b=3
b=6
b=10
1
accr_excess_gap
2.168
2.168
1.266
0.976
0.976
0.976
4
4
4
2
anchor_fit_value
1.014
1.385
1.588
0.988
1.000
1.000
4
4
4
3
aog_vwap30_deq80
1.609
1.609
0.670
0.927
0.927
0.939
4
4
4
4
cohort_excess_pulse_rev
1.448
1.448
1.519
0.988
0.988
1.000
4
4
4
5
droa_early_pit
2.095
2.095
2.821
1.000
1.000
0.976
3
3
3
Appendix
Table E.6: Task-level metrics for GLM-5.2 at iterations 3, 6, and 10.
No.
Task ID
Sharpe
Monotonicity
Alignment
b=3
b=6
b=10
b=3
b=6
b=10
b=3
b=6
b=10
1
accr_excess_gap
1.256
1.262
1.423
0.915
0.976
0.976
5
2
2
2
anchor_fit_value
1.944
1.993
1.772
1.000
1.000
1.000
4
4
4
3
aog_vwap30_deq80
-0.035
-3.320
-0.265
0.745
0.879
0.915
5
4
4
4
cohort_excess_pulse_rev
1.267
1.431
1.398
0.976
0.964
0.988
4
5
4
5
droa_early_pit
0.887
0.810
1.148
0.976
0.988
0.939
5
3
4
Appendix
Table E.7: Task-level metrics for MiniMax-M3 at iterations 3, 6, and 10.
No.
Task ID
Sharpe
Monotonicity
Alignment
b=3
b=6
b=10
b=3
b=6
b=10
b=3
b=6
b=10
1
accr_excess_gap
0.016
0.745
2.615
0.491
0.539
1.000
3
3
0
2
anchor_fit_value
1.502
1.502
1.689
0.988
0.988
0.988
2
2
3
3
aog_vwap30_deq80
0.710
0.710
0.710
0.842
0.842
0.842
4
4
4
4
cohort_excess_pulse_rev
1.658
1.658
2.404
0.988
1.000
1.000
4
4
3
5
droa_early_pit
1.865
2.543
2.851
0.964
0.988
1.000
3
3
2
Appendix
Table E.8: Task-level metrics for Qwen3.7-Plus at iterations 3, 6, and 10.
No.
Task ID
Sharpe
Monotonicity
Alignment
b=3
b=6
b=10
b=3
b=6
b=10
b=3
b=6
b=10
1
accr_excess_gap
-0.700
-0.524
-0.524
0.250
0.273
0.273
4
4
3
2
anchor_fit_value
1.741
1.741
-1.628
1.000
1.000
0.661
3
3
4
3
aog_vwap30_deq80
0.334
0.334
-0.092
0.709
0.915
1.000
5
5
2
4
cohort_excess_pulse_rev
1.360
1.360
1.744
0.976
0.976
1.000
5
5
4
5
droa_early_pit
1.548
1.548
4.489
0.988
0.988
1.000
3
3
4
Appendix
Table E.9: Task-level metrics for DeepSeek-V3.1 at iterations 3, 6, and 10.