In memetic symbolic regression, candidate structures are compared after coefficient calibration, so the calibration protocol itself contributes to evolutionary selection. Standard centralized calibration evaluates each structure at its pooled-sample optimum, ignoring how stable this calibration is under covariate shifts, and can thus favor structures whose fit relies on sample-specific coefficients. We propose Dirichlet-Sinkhorn Constant Averaging (DSCA), a calibration strategy that partitions the optimization data into equally sized subsets with different covariate distributions, calibrates each candidate independently on every partition, and evaluates it at the mean of the resulting parameters. We show that the excess loss of DSCA relative to centralized calibration vanishes at the population level for correctly specified, identifiable expressions, whereas under misspecification it persists when partition-specific calibrations do not aggregate to the pooled optimum. On synthetic benchmarks and ten real-world datasets, DSCA improves functional recovery and the accuracy-complexity trade-off over centralized Broyden-Fletcher-Goldfarb-Shanno and Levenberg-Marquardt calibration, under selection by negative log-likelihood and by the Akaike and Bayesian information criteria. Mechanism analyses associate the DSCA excess loss with the generalization gap and show that the effect is not reproduced by repeated centralized fitting. These results indicate that controlled heterogeneous calibration provides a complementary source of selection pressure in symbolic-regression search.
Figures & tables
Parameter
Value/Setting
Population size ( μ )
100
Offspring size ( λ )
100
Number of generations
100
Initialization method
Random with decay
Parsimony
0.8
Parsimony decay
0.85
TABLE I: Evolutionary and numerical optimization hyperparameters.
Function
Generating Function
Distribution
d
Reference
f1(x)
8x12+8x23−15
xi∼U(−3,3)
2
Jin-2 [ 6 , 39 , 40 ]
f2(x)
2.5x14−1.3x13+0.5x22−1.7x2
xi∼U(−3,3)
2
Jin-1 [ 6 , 39 , 40 ]
f3(x)
1.35x1x2+5.5sin((x1−1)(x2−1))
xi∼U(−3,3)
2
Jin-6 [ 6 , 39 , 40 ]
f4(x)
1.2+(x2−2.5)2e−(x1−1)2
xi∼U(0.3,4)
2
Vladislavleva-1 [ 41 ]
f5(x)
(x1−10)x2230(x1−1)(x3−1)
x1,x3∼U(0.05,2),x2∼U(1,2)
3
Vladislavleva-5 [ 41 ]
TABLE II: Generating functions and variable distributions for the synthetic benchmark track.
Dataset
d
B
R
total
test
Fri_c1
25
0
25
1000
250
Fri_c2
5
0
5
1000
250
Pollen
4
0
4
3790
948
Puma8NH
8
0
8
8192
2048
Echo M.
9
3
6
17233
4309
Houses
8
0
8
20330
5083
TABLE III: Characteristics of the selected PMLB datasets.
NLL
AIC
BIC
N
Noise
BFGS
LM
BFGS
LM
BFGS
LM
100
Low
3/1/1
4/1/0
3/1/1
3/1/1
3/2/0
1/3/1
High
2/3/0
2/3/0
2/3/0
2/3/0
4/1/0
2/2/1
200
Low
3/2/0
3/2/0
2/3/0
3/2/0
2/3/0
2/2/1
High
3/2/0
3/1/1
3/2/0
3/2/0
3/2/0
3/1/1
TABLE IV: Win/tie/loss counts of DSCA against each baseline over the synthetic datasets, by sample size, noise level, and selection criterion.
100 Samples
200 Samples
Algorithm
NLL
AIC
BIC
NLL
AIC
BIC
DSCA (our)
26
31
43
43
49
55
BFGS
14
14
30
28
31
48
LM
5
13
37
24
30
49
TABLE V: Aggregate success counts on synthetic datasets by sample size and selection criterion.
Fig. 1: Bayesian Wilcoxon signed-rank test results evaluating predictive accuracy (top panels) and model complexity (bottom panels) against the baselines.
Algorithm
BFGS
DSCA (our)
Criterion
N samples
NLL
100
0.26 ± 0.57
0.85 ± 0.74
200
0.28 ± 0.60
0.40 ± 0.63
500
0.02 ± 0.04
0.14 ± 0.40
AIC
100
0.30 ± 0.43
0.78 ± 0.66
200
0.36 ± 0.69
0.54 ± 0.67
TABLE VI: Predictive accuracy ( R2 ) differences ( Δa ) relative to the Levenberg-Marquardt baseline (Mean ± Std).
Algorithm
BFGS
DSCA (our)
Criterion
N samples
NLL
100
-0.2 ± 1.1
-15.8 ± 10.8
200
-0.6 ± 2.3
-7.5 ± 6.2
500
-1.2 ± 2.9
-4.0 ± 2.8
AIC
100
-0.8 ± 2.8
-23.3 ± 11.4
200
-2.1 ± 4.9
-15.1 ± 13.3
TABLE VII: Model complexity differences ( Δa ) relative to the Levenberg-Marquardt baseline (Mean ± Std).
Fig. 2: Algorithmic training time (in seconds) across different selection criteria and sample size regimes.
Fig. 3: Generalization gap of the pooled final populations, by the optimizer that produced each structure (200-sample regime, centralized re-fit).
Fig. 4: Excess loss ΔN′ against the generalization gap for the pooled final populations, colored by the optimizer that produced the structure.
Fig. 5: Ablation study of the predictive accuracy ( R2 ) regret gap under different partitioning strategies.
Fig. 6: Ablation study of the model complexity regret gap under different partitioning strategies.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Fig. 7: Success rates on synthetic datasets ( N=100 ) under low and high noise regimes.
Fig. 8: Success rates on synthetic datasets ( N=200 ) under low and high noise regimes.
Effective model selection is critical in symbolic regression (SR) to identify mathematical expressions that balance accuracy and complexity, and have low expected error on unseen data. Many modern implementations of genetic programming (GP) for SR generate a set of Pareto optimal candidate solutions, but reliable automatic selection of solutions that generalize well remains an open issue. Current literature offers various information-theoretic and Bayesian approaches, yet comprehensive comparisons of their performance across different data regimes are limited. This study presents a systematic empirical comparison of widely used selection criteria: the Akaike information criterion (AIC), the corrected AIC (AICc), the Bayesian information criterion (BIC), minimum description length (MDL), as well as Efron's bootstrap estimate for the in-sample prediction error on seven synthetic datasets with Gaussian noise. We rank candidate expressions generated by perturbing ground-truth functions to assess generalization error and selection probability of the ground-truth expression. Our findings reveal that MDL consistently identifies models with the lowest test error and the shortest length across most datasets. While no single criterion dominates all results, MDL and BIC produced the highest probability of selecting the ground-truth expressions.
Ali Soltani, Gabriel Kronberger, Fabricio Olivetti de Franca +2
Aarhus University Department of Mechanical Engineering Aarhus, Denmark · Heuristic and Evolutionary Algorithms Laboratory (HEAL) University of Applied Sciences Upper Austria Hagenberg, Austria · Federal University of ABC Center for Mathematics, Computing, and Cognition Brazil +1
Symbolic regression (SR) seeks concise and interpretable mathematical expressions from data for scientific equation discovery. Existing SR benchmarks face a tradeoff between evaluation cost and benchmark validity. Repeated evaluation of large task pools is expensive, and compact benchmarks lack systematic evidence of preserved task diversity and algorithm discriminability. SymbolicArena provides a unified infrastructure for benchmark distillation and dynamic evaluation. The framework standardizes 664 heterogeneous tasks with executable ground truth expressions and distills the Full Task Set into Core50, a validated benchmark of 50 tasks. The distillation process preserves task coverage and algorithm discrimination under explicit balance constraints. SymbolicArena applies a unified execution protocol to heterogeneous SR algorithms and produces comparable outputs and search trajectories. Multi Axis Evaluation characterizes numerical quality, symbolic quality, and search behavior. Core50 reduces evaluation workload by 92.5% and maintains agreement with Full Task Set evaluations. Experiments show that SymbolicArena achieves 72.6% to 86.7% lower approximation error than alternative selectors, further supporting its fidelity to the Full Task Set. Evaluation reveals a substantial gap between numerical fitting and symbolic recovery across current SR methods, suggesting that reliable equation recovery remains an open challenge.
Ziwen Zhang, Xiju Wu, Yuheng Jing +9
School of Artificial Intelligence, University of Chinese Academy of Sciences · C2DL, Institute of Automation, Chinese Academy of Sciences · Nanjing University of Science and Technology +1
Symbolic regression aims to uncover explicit scientific laws from data. Recent methods use LLMs to guide mutation from background text, which is more directed than random genetic programming. However, exact symbolic recovery requires both semantic guidance and explicit structure, so that domain-informed search are carried out through valid symbolic representation. Current LLM-driven systems remain structure-blind: they select among opaque candidates, lack explicit mechanisms for local mutation, and rely on brittle coefficient fitting that can undervalue correct skeletons. We propose FunctionEvolve, an evolutionary framework using expression trees to organize the whole search: structural summaries promote diverse parent selection, local tree edits preserve useful subexpressions, and structure-aware fitting decomposes, constrains, and simplifies coefficients for more reliable scoring. It uses only elementary function families, without additional domain-specific rules limiting generalization. On the 129-task synthetic subset of LLM-SRBench, FunctionEvolve with \emph{Claude Opus 4.6} recovers 107 exact forms, reaching 82.9% SA@50, 4.5x above same-backbone baselines, and 55.8% SA@1, 3.6x above the strongest previously published top-1 result. Ablations show that structure-visible search is central to reliable recovery, with LLM-guided refinements and structure-aware coefficient optimization serving as essential proposal and scoring mechanisms. We also audit the benchmark and show that collinearity in its materials-science subset creates identifiability issues.
Zeyu Xia, Jun Zhu, Dong Yan
Bosch Center for Artificial Intelligence · Department of Computer Science and Technology, Tsinghua University · Department of Computer Science and Technology +1