AlphaPareto: Formulaic Alpha Discovery with LLM-Guided Multi-Objective Reinforcement Learning
Organizations: School of Economics Xiamen University Fujian, China · Paula and Gregory Chow Institute for Studies in Economics Xiamen University Fujian, China · School of Economics & Wang Yanan Institute for Studies in Economics Xiamen University Fujian, China
Abstract
Formulaic alpha discovery is a core challenge in quantitative trading, as identifying alphas that work well together remains difficult. Recent reinforcement learning (RL) methods formulate this task as a Markov decision process (MDP), but two important issues remain unresolved. First, as the alpha pool evolves, the reward function changes accordingly, making the MDP inherently non-stationary. Second, most existing methods optimize a single objective, typically predictive power, while ignoring other important properties of a high-quality alpha pool. Motivated by these challenges, we propose AlphaPareto, an RL method for formulaic alpha discovery. To address non-stationarity, AlphaPareto augments the state to include both the alpha under construction and the current alpha pool, and applies a large language model (LLM) to encode the pool. This design allows the agent to adapt to the evolving search environment. To overcome the limitation of single-objective reward design, AlphaPareto replaces the scalar reward with a multi-objective vector-valued reward that simultaneously captures predictive power, temporal stability, perturbation robustness, and diversity, and optimizes these objectives through a Pareto-regularized learning procedure. Empirical applications to real-world datasets show that our AlphaPareto method outperforms its competitors.
Figures & tables
| CSI300 | CSI800 | Market | ||||||
|---|---|---|---|---|---|---|---|---|
| Method | Mean | Std | Mean | Std | Mean | Std | ||
| Alpha158 | 2.99% | - | 4.77% | - | 4.04% | - | ||
| MLP | 1.51% | (2.33%) | 4.63% | (0.90%) | 7.00% | (0.18%) | ||
| GP | 1.09% | (0.82%) | 2.74% | (1.40%) | 7.52% | (1.53%) | ||
| AlphaAgent | 0.51% | (0.34%) | 0.48% | (0.46%) | 1.10% | (0.59%) | ||
| R&D-Agent-Quant | 2.39% | (0.96%) | 2.45% | (0.73%) | 1.48% | (0.91%) | ||
| Model | AV ( ) | MDD ( ) | IR ( ) | SR ( ) |
|---|---|---|---|---|
| Alpha158 | 33.17% | -18.71% | 1.08 | 1.34 |
| MLP | 47.09% | -25.21% | 1.57 | 1.58 |
| GP | 24.49% | -13.29% | 0.77 | 1.18 |
| AlphaGen | 40.03% | -15.91% | 1.65 | 1.61 |
| AlphaAgent | 35.07% | -29.94% | 0.71 | 0.89 |
| R&D-Agent-Quant | 37.30% | -30.18% | 0.76 | 0.96 |
| CSI300 | CSI800 | Market | ||||||
|---|---|---|---|---|---|---|---|---|
| Model | Mean | Std | Mean | Std | Mean | Std | ||
| DeepSeek-1.5B | 2.65% | (0.99%) | 5.39% | (0.44%) | 9.33% | (0.92%) | ||
| Qwen-Embedding-4B | 3.92% | (0.46%) | 5.70% | (0.87%) | 10.10% | (1.01%) | ||
| DeepSeek-7B | 3.38% | (0.56%) | 5.50% | (1.23%) | 9.50% | (0.56%) | ||
| DeepSeek-14B | 3.71% | (1.73%) | 5.59% | (0.46%) | 9.74% | (0.64%) | ||
| DeepSeek-32B | 3.26% | (1.43%) | 4.76% | (0.52%) | 9.88% | (1.13%) | ||
| Included Components | CSI300 | CSI800 | Market | ||||||
|---|---|---|---|---|---|---|---|---|---|
| LLM | MORL | Mean | Std | Mean | Std | Mean | Std | ||
| 3.69% | (0.66%) | 5.07% | (0.63%) | 8.44% | (1.09%) | ||||
| 2.37% | (0.68%) | 5.04% | (0.42%) | 9.68% | (0.83%) | ||||
| 3.99% | (1.27%) | 5.54% | (1.03%) | 9.63% | (0.80%) | ||||
| 3.92% | (0.46%) | 5.70% | (0.87%) | 10.10% | (1.01%) | ||||
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Tokens | Description |
|---|---|
| Features | |
| Open/High/Low/Close/Vwap/Volume | Opening/high/low/closing/vwap price or volume of stock at time . |
| Constant | Number from . |
| Time delta | Integer from , which is used in the time-series operators. |
| Time-series Operators | |
| Return the value of , where is the time delta and is a feature of | |
| Chosen model | Embedding dimension |
|---|---|
| DeepSeek-1.5B | 1536 |
| Qwen-Embedding-4B | 2560 |
| DeepSeek-7B | 3584 |
| DeepSeek-14B | 5120 |
| DeepSeek-32B | 5120 |
| Hyperparameter | Values |
|---|---|
| Instruments | CSI300 / CSI800 / Market |
| (Alpha Pool Size) | 10, 20 |
| Leading days | 20 |
| Batch size | 128 |
| Optimizer | Adam |
| Learning rate |
| CSI300 | CSI800 | Market | ||||||
|---|---|---|---|---|---|---|---|---|
| Method | -value | -value | -value | |||||
| Alpha158 | 4.55 | (0.52%) | 2.40 | (3.73%) | 13.39 | (0.01%) | ||
| MLP | 2.73 | (2.62%) | 7.05 | (0.11%) | 8.00 | (0.07%) | ||
| GP | 16.11 | (0.00%) | 9.58 | (0.03%) | 10.87 | (0.02%) | ||
| AlphaAgent | 55.88 | (0.00%) | 25.41 | (0.00%) | 44.84 | (0.00%) | ||
| R&D-Agent-Quant | 6.59 | (0.14%) | 34.67 | (0.00%) | 116.93 | (0.00%) | ||
| CSI800 | ||
|---|---|---|
| Pool representation | Mean | Std |
| None (MORL only) | 5.54% | (1.03%) |
| Matched noise | 4.22% | (1.62%) |
| LLM embedding (AlphaPareto) | 5.70% | (0.87%) |
| S&P 500 | ||
|---|---|---|
| Method | Mean | Std |
| Alpha158 | 3.92% | - |
| MLP | 5.43% | (0.31%) |
| AlphaAgent | -0.01% | (0.56%) |
| R&D-Agent-Quant | 1.29% | (0.49%) |
| AlphaGen | 7.03% | (0.30%) |
| Experiment | Runs | Hardware per run | Time per run | Total compute |
|---|---|---|---|---|
| Main comparison | 195 | GPU | 8h | 1560h |
| LLM scaling study | 150 | GPU | 8h | 1200h |
| Ablation study | 360 | GPU | 7h | 2520h |