stat.MLSep 28, 2026

AlphaPareto: Formulaic Alpha Discovery with LLM-Guided Multi-Objective Reinforcement Learning

Authors: Yingbo Zhao, Zeyu Yang, Zhoufan Zhu

Organizations: School of Economics Xiamen University Fujian, China · Paula and Gregory Chow Institute for Studies in Economics Xiamen University Fujian, China · School of Economics & Wang Yanan Institute for Studies in Economics Xiamen University Fujian, China

Abstract

Formulaic alpha discovery is a core challenge in quantitative trading, as identifying alphas that work well together remains difficult. Recent reinforcement learning (RL) methods formulate this task as a Markov decision process (MDP), but two important issues remain unresolved. First, as the alpha pool evolves, the reward function changes accordingly, making the MDP inherently non-stationary. Second, most existing methods optimize a single objective, typically predictive power, while ignoring other important properties of a high-quality alpha pool. Motivated by these challenges, we propose AlphaPareto, an RL method for formulaic alpha discovery. To address non-stationarity, AlphaPareto augments the state to include both the alpha under construction and the current alpha pool, and applies a large language model (LLM) to encode the pool. This design allows the agent to adapt to the evolving search environment. To overcome the limitation of single-objective reward design, AlphaPareto replaces the scalar reward with a multi-objective vector-valued reward that simultaneously captures predictive power, temporal stability, perturbation robustness, and diversity, and optimizes these objectives through a Pareto-regularized learning procedure. Empirical applications to real-world datasets show that our AlphaPareto method outperforms its competitors.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery

    May 14, 2026Lingzhe Zhang, Tong Jia, Yunpeng Zhai +5Exploratory Factor AnalysisReinforcement Fine-Tuning

  2. AlphaRJM: Reward-Jump Memory for Stochastic Return-Guided Alpha Discovery

    Sep 8, 2026Sayan Dhan, Selvaraju NatarajanSoft Actor-CriticStochastic

  3. Towards Autonomous Formulaic Alpha Discovery: An Evolutionary Computation Perspective

    Aug 3, 2026Xinwei Yu, Yiyang Fu, Mingcheng Fan +3Evolutionary OptimizationModel Discovery