Self-evolving agents aim to turn research feedback into reusable skills, tools, and research rules. Whether these accumulated capabilities continue to improve later research requires controlled evaluation. Long-horizon alpha discovery provides a state-dependent setting: once a new factor enters the portfolio, the predictive information already covered changes, so the value of the same candidate or experience may change over time. We introduce EverMine, an empirical framework for studying self-evolving research capabilities in long-horizon alpha discovery. EverMine decomposes the research state into history (Hist), the current factor portfolio (Frontier), and reusable capabilities (Cap). Under matched resource limits, we compare complete runs with fixed or evolving Cap, and replace Cap while holding Hist and Frontier fixed to estimate the conditional value of accumulated capabilities. We also combine full trajectories with historical-state replay to examine how experience-based decisions affect candidate selection and portfolio outcomes. Across 18 long-horizon trajectories, end-to-end comparisons show no consistent gain from Cap evolution. Across 48 continuation branches from shared Hist and Frontier states, accumulated Cap also does not consistently outperform the initial Cap. Parameter tuning of existing factor structures can still improve the portfolio. In an exploratory replay of two screening batches from one Evolving trajectory, some screened-out candidates have positive marginal value at the original state, yet submitting all screened-out candidates sequentially slightly lowers final portfolio IC in both batches. These results show that candidate value depends on the evolving portfolio and submission order, and motivate evaluating self-evolving research capabilities through end-to-end outcomes, conditional capability value, and the consequences of experience-based decisions.
Figures & tables
Model
nF
nE
Mean Fixed IC
Mean Evolving IC
Mean IC diff. Δ
95% interval
Development
Qwen3.8-27B
6
6
60.902
61.643
+0.741
[-1.181, +2.681]
DeepSeek-V4.1-Flash
3
3
68.166
68.071
-0.095
[-6.819, +6.624]
OOS
Qwen3.8-27B
6
6
88.324
87.932
-0.391
[-4.158, +3.180]
DeepSeek-V4.1-Flash
3
3
100.779
99.141
-1.637
[-18.227, +13.089]
Table 1: Complete end-to-end endpoint results. IC values and intervals are multiplied by 1000 . nF and nE are the numbers of independent trajectories. Differences are Evolving minus Fixed, with 95% bootstrap intervals.
Model
Checkpoint
Parent runs
Paired blocks
ΔQ(Cap0)
ΔQ(Capt)
τcap
95% interval
Development
Qwen3.8-27B
25%
6
12
+2.034
+1.628
-0.405
[-1.097, +0.119]
Qwen3.8-27B
75%
6
12
+0.500
+0.542
+0.042
[-0.494, +0.573]
OOS
Qwen3.8-27B
25%
6
12
+3.005
+1.749
-1.256
[-3.107, +0.197]
Qwen3.8-27B
75%
6
12
+0.718
+0.626
-0.092
[-2.048, +1.574]
Table 2: Complete results for capability-state replacement. IC gains and intervals are multiplied by 1000 . Starting groups are paired blocks, and parent trajectories are the independent statistical units. The effect is ΔQ(Capt)−ΔQ(Cap0) , with 95% bootstrap intervals.
A. Known net contributions after the pool first reaches capacity
Model
Cap
New structures
Parameter tuning of existing structures
Exact resubmissions of previously recorded factors
Runs with positive tuning contribution
Qwen
Fixed
+0.932
+0.588
−0.049
4/6
Qwen
Evolving
+1.657
+0.567
−0.022
6/6
DeepSeek
Fixed
+5.910
+2.009
+0.076
3/3
DeepSeek
Evolving
+6.736
+0.996
+0.050
3/3
B. Screened-out factors: original-state gains and sequential-submission outcomes
Table 3: Research choices and realized portfolio gains. IC values are multiplied by 1000 . A: submissions after the pool first reaches capacity are classified as new structures, parameter tuning of existing structures, or repeated submissions of previously recorded factors. Net gains are summed within each trajectory and then averaged equally within each group. Four repeated-submission outcomes are missing and are reported using known contributions. B: two complete screening batches from the same Qwen Evolving trajectory. “Positive on original portfolio” counts screened-out factors with ΔIC>10−6 when evaluated independently at the original state. Additional factors are then submitted in a fixed order, and the final portfolio is reported.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Dimension
Reusable judgments and operations
Intended change in research behavior
Market mechanism
Form competing explanations from participant behavior, constraints, and information diffusion
Move from raw formula variation toward competing market explanations that can be distinguished by evidence
Measurement and representation
Check what fields, reference frames, time scales, and transformations actually measure
Correct proxies, scales, and representations so that expression changes are not mistaken for new mechanisms
Current portfolio gaps
Use members, weights, and feedback to identify missing information in the current portfolio
Select research questions that may add information not yet covered by the portfolio
Test design
Design controls and counterexamples that distinguish competing explanations
Use limited experiments to obtain discriminative evidence so that both successes and failures update judgments
Belief updating
Revise the scope of a judgment using supporting evidence and counterexamples
Retain, narrow, revise, disable, or re-enable existing judgments
Attention and budget
Allocate effort among exploitation, exploration, measurement correction, and historical review
Avoid prolonged spending on low-information branches while preserving useful depth
Appendix
Table 4: Capability dimensions in the Evolving condition and their intended changes to research behavior.
A. Sample Size and Dynamic Universe
Period
10-min bars
Valid assets/bar
Unique assets
Nominal asset–bars
Development
52,560
58/60/60
163
3,153,600
OOS test
26,064
59/60/60
121
1,563,840
Total
78,624
58/60/60
199
4,717,440
Appendix
Table 5: Binance Spot dynamic-U60 dataset statistics. Valid assets per bar are reported as minimum/median/maximum.
Model
Condition
20%
50%
80%
Qwen3.8-27B
Fixed
0.06013 (6)
0.06066 (6)
0.06090 (5)
Qwen3.8-27B
Evolving
0.05898 (6)
0.06039 (6)
0.06149 (6)
DeepSeek-V4.1-Flash
Fixed
0.06319 (3)
0.06431 (3)
0.06618 (3)
DeepSeek-V4.1-Flash
Evolving
0.06216 (3)
0.06382 (3)
0.06442 (3)
Appendix
Table 6: Development-period factor-pool IC at selected joint-budget checkpoints. Parentheses show the number of trajectories reaching each checkpoint.