Self-evolving agents aim to turn research feedback into reusable skills, tools, and research rules. Whether these accumulated capabilities continue to improve later research requires controlled evaluation. Long-horizon alpha discovery provides a state-dependent setting: once a new factor enters the portfolio, the predictive information already covered changes, so the value of the same candidate or experience may change over time. We introduce EverMine, an empirical framework for studying self-evolving research capabilities in long-horizon alpha discovery. EverMine decomposes the research state into history (Hist), the current factor portfolio (Frontier), and reusable capabilities (Cap). Under matched resource limits, we compare complete runs with fixed or evolving Cap, and replace Cap while holding Hist and Frontier fixed to estimate the conditional value of accumulated capabilities. We also combine full trajectories with historical-state replay to examine how experience-based decisions affect candidate selection and portfolio outcomes. Across 18 long-horizon trajectories, end-to-end comparisons show no consistent gain from Cap evolution. Across 48 continuation branches from shared Hist and Frontier states, accumulated Cap also does not consistently outperform the initial Cap. Parameter tuning of existing factor structures can still improve the portfolio. In an exploratory replay of two screening batches from one Evolving trajectory, some screened-out candidates have positive marginal value at the original state, yet submitting all screened-out candidates sequentially slightly lowers final portfolio IC in both batches. These results show that candidate value depends on the evolving portfolio and submission order, and motivate evaluating self-evolving research capabilities through end-to-end outcomes, conditional capability value, and the consequences of experience-based decisions.
Figures & tables
Model
nF
nE
Mean Fixed IC
Mean Evolving IC
Mean IC diff. Δ
95% interval
Development
Qwen3.8-27B
6
6
60.902
61.643
+0.741
[-1.181, +2.681]
DeepSeek-V4.1-Flash
3
3
68.166
68.071
-0.095
[-6.819, +6.624]
OOS
Qwen3.8-27B
6
6
88.324
87.932
-0.391
[-4.158, +3.180]
DeepSeek-V4.1-Flash
3
3
100.779
99.141
-1.637
[-18.227, +13.089]
Table 1: Complete end-to-end endpoint results. IC values and intervals are multiplied by 1000 . nF and nE are the numbers of independent trajectories. Differences are Evolving minus Fixed, with 95% bootstrap intervals.
Model
Checkpoint
Parent runs
Paired blocks
ΔQ(Cap0)
ΔQ(Capt)
τcap
95% interval
Development
Qwen3.8-27B
25%
6
12
+2.034
+1.628
-0.405
[-1.097, +0.119]
Qwen3.8-27B
75%
6
12
+0.500
+0.542
+0.042
[-0.494, +0.573]
OOS
Qwen3.8-27B
25%
6
12
+3.005
+1.749
-1.256
[-3.107, +0.197]
Qwen3.8-27B
75%
6
12
+0.718
+0.626
-0.092
[-2.048, +1.574]
Table 2: Complete results for capability-state replacement. IC gains and intervals are multiplied by 1000 . Starting groups are paired blocks, and parent trajectories are the independent statistical units. The effect is ΔQ(Capt)−ΔQ(Cap0) , with 95% bootstrap intervals.
A. Known net contributions after the pool first reaches capacity
Model
Cap
New structures
Parameter tuning of existing structures
Exact resubmissions of previously recorded factors
Runs with positive tuning contribution
Qwen
Fixed
+0.932
+0.588
−0.049
4/6
Qwen
Evolving
+1.657
+0.567
−0.022
6/6
DeepSeek
Fixed
+5.910
+2.009
+0.076
3/3
DeepSeek
Evolving
+6.736
+0.996
+0.050
3/3
B. Screened-out factors: original-state gains and sequential-submission outcomes
Table 3: Research choices and realized portfolio gains. IC values are multiplied by 1000 . A: submissions after the pool first reaches capacity are classified as new structures, parameter tuning of existing structures, or repeated submissions of previously recorded factors. Net gains are summed within each trajectory and then averaged equally within each group. Four repeated-submission outcomes are missing and are reported using known contributions. B: two complete screening batches from the same Qwen Evolving trajectory. “Positive on original portfolio” counts screened-out factors with ΔIC>10−6 when evaluated independently at the original state. Additional factors are then submitted in a fixed order, and the final portfolio is reported.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Dimension
Reusable judgments and operations
Intended change in research behavior
Market mechanism
Form competing explanations from participant behavior, constraints, and information diffusion
Move from raw formula variation toward competing market explanations that can be distinguished by evidence
Measurement and representation
Check what fields, reference frames, time scales, and transformations actually measure
Correct proxies, scales, and representations so that expression changes are not mistaken for new mechanisms
Current portfolio gaps
Use members, weights, and feedback to identify missing information in the current portfolio
Select research questions that may add information not yet covered by the portfolio
Test design
Design controls and counterexamples that distinguish competing explanations
Use limited experiments to obtain discriminative evidence so that both successes and failures update judgments
Belief updating
Revise the scope of a judgment using supporting evidence and counterexamples
Retain, narrow, revise, disable, or re-enable existing judgments
Attention and budget
Allocate effort among exploitation, exploration, measurement correction, and historical review
Avoid prolonged spending on low-information branches while preserving useful depth
Appendix
Table 4: Capability dimensions in the Evolving condition and their intended changes to research behavior.
A. Sample Size and Dynamic Universe
Period
10-min bars
Valid assets/bar
Unique assets
Nominal asset–bars
Development
52,560
58/60/60
163
3,153,600
OOS test
26,064
59/60/60
121
1,563,840
Total
78,624
58/60/60
199
4,717,440
Appendix
Table 5: Binance Spot dynamic-U60 dataset statistics. Valid assets per bar are reported as minimum/median/maximum.
Model
Condition
20%
50%
80%
Qwen3.8-27B
Fixed
0.06013 (6)
0.06066 (6)
0.06090 (5)
Qwen3.8-27B
Evolving
0.05898 (6)
0.06039 (6)
0.06149 (6)
DeepSeek-V4.1-Flash
Fixed
0.06319 (3)
0.06431 (3)
0.06618 (3)
DeepSeek-V4.1-Flash
Evolving
0.06216 (3)
0.06382 (3)
0.06442 (3)
Appendix
Table 6: Development-period factor-pool IC at selected joint-budget checkpoints. Parentheses show the number of trajectories reaching each checkpoint.
We study recursive self-improvement at the level of quantitative-investment research: whether an autonomous system can use evidence from earlier experiments to improve the hypotheses and candidates proposed in later iterations. We present AQuA, which comprises two separate language-model-driven research systems: one for symbolic factor discovery and one for trainable model development. Each system records experimental results and uses them to guide subsequent proposals. Each operates in a fixed sandbox, which fixes the data splits, feature and label definitions, and evaluator while allowing the model to act only through constrained factor expressions or configuration diffs. The factor system, a manager-mediated multi-agent pipeline, discovers and combines factors into a signal that reaches a combined validation information coefficient of about 0.190 on a crypto universe. The model system, a config-driven loop over a hybrid time-series architecture, reaches a per-stock information coefficient of +0.0843 on US equities and converts it into a threshold long/short strategy with a held-out Sharpe of up to +2.50 at a two-leg cost. The strategy is positive in every year from 2021 to 2025.
Jiacheng Guo, Suozhi Huang, Yunlong Gao +5
Princeton University · Ant Group · Stanford University
Large language model agents can discover alphas, yet current methods have three weaknesses. The search cannot adapt during the run, automation usually ends at alpha generation while library selection and model choice stay manual, and alpha discovery can read the test window through loop feedback or code problems. We present AutoScientist-Quant, a self evolving search process that regards quantitative research as one budgeted search problem. A single controller conditions every decision on the remaining budget, choosing at each round whether to improve, combine, pivot, or stop, which node to expand, how many alphas to generate, and how to retrieve past trajectories from the shared memory. The same core then selects from the library and tunes the model, closing the loop from hypothesis to deployable strategy. We also review the evaluation pipeline reused from prior work, fix two lookahead problems, and keep the feedback window disjoint from the held out test window, so every comparison tests true generalization. On CSI universes, the framework attains the best value of nearly every metric in every setting, and these conclusions hold across several backbones and markets.
Recent advances in LLM agents enable systems that autonomously refine workflows, accumulate reusable skills, self-train their underlying models, and maintain persistent memory. However, we show that such self-evolution is often non-monotonic: adapting to new task distributions can progressively degrade previously acquired capabilities across all major evolution channels. We identify this phenomenon as \emph{capability erosion under self-evolution} and show that it consistently emerges across workflow, skill, model, and memory evolution. To mitigate this issue, we propose \emph{Capability-Preserving Evolution} (CPE), a general stabilization principle that constrains destructive capability drift during continual adaptation. Across all four evolution dimensions, CPE consistently improves retained capability stability while preserving adaptation performance. For example, in workflow evolution, CPE improves retained simple-task performance from 41.8% to 52.8% under GPT-5.1 optimization while simultaneously achieving stronger complex-task adaptation. Our findings suggest that stable long-horizon self-evolving agents require not only acquiring new capabilities, but also explicitly preserving previously learned ones during continual adaptation.