cs.AISep 28, 2026

From Search to Research: Exploring Search Scaling in Autonomous Quantitative Factor Mining

Authors: Kangcheng Deng, Hui Cai, Jiacheng Lu, Chester Zhongshu Qian, Rui Sun, Beidi Luan, Jing Li, Daxin Jiang, +1 more

Organizations: StepFun · Shanghai Jiao Tong University · University of California, Los Angeles · FinStep

Abstract

Inference scaling has been shown to improve large language model (LLM) performance, and this principle naturally extends to autonomous LLM agents through increased search budgets, which we refer to as search scaling. Although prior work has characterized the mechanisms, scaling behavior, and performance limits of LLM inference scaling, much less is known about these questions in autonomous research. Therefore, we investigate how search scaling affects research performance and what mechanisms drive these gains using 50 quantitative factor-mining tasks grounded in financial research reports. Each task requires an agent to carry out an end-to-end research loop, from interpreting a hypothesis and implementing it in code to evaluating and iteratively refining the resulting factor. Across nine models, we examine how model capability, search depth, and search organization shape factor quality by tracing performance across varying budgets, transferring intermediate research states between models, and comparing different search strategies. We find that (1) initial performance is more strongly associated with model capability, while deeper search can narrow cross-model gaps; (2) model grafting shows that the early research state materially shapes final performance; and (3) parallel search outperforms sequential search under the same iteration budget, consistent with benefits from broader coverage of the search space. Further trajectory analysis shows that higher-performing models more effectively diagnose failures, revise search directions, and preserve the intended economic hypothesis when selecting candidates. These findings suggest that future progress in autonomous research will require stronger models together with adaptive policies for deploying test-time computation throughout the research process.

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent

    Apr 20, 2026Wanli Li, Bince Qu, Bo Pan +5Agentic Reinforcement LearningDeep Research

  2. Think Big, Search Small: Where Capacity Matters in Hierarchical Search Agents?

    Jul 8, 2026Qinnan Cai, Yibo Zhao, Xiang LiSearch AgentsMulti-Hop Reasoning