SymbolicArena: A Unified Infrastructure for Benchmark Distillation and Dynamic Evaluation in Symbolic Regression
Organizations: School of Artificial Intelligence, University of Chinese Academy of Sciences · C2DL, Institute of Automation, Chinese Academy of Sciences · Nanjing University of Science and Technology · School of Future Technology, University of Chinese Academy of Sciences
Abstract
Symbolic regression (SR) seeks concise and interpretable mathematical expressions from data for scientific equation discovery. Existing SR benchmarks face a tradeoff between evaluation cost and benchmark validity. Repeated evaluation of large task pools is expensive, and compact benchmarks lack systematic evidence of preserved task diversity and algorithm discriminability. SymbolicArena provides a unified infrastructure for benchmark distillation and dynamic evaluation. The framework standardizes 664 heterogeneous tasks with executable ground truth expressions and distills the Full Task Set into Core50, a validated benchmark of 50 tasks. The distillation process preserves task coverage and algorithm discrimination under explicit balance constraints. SymbolicArena applies a unified execution protocol to heterogeneous SR algorithms and produces comparable outputs and search trajectories. Multi Axis Evaluation characterizes numerical quality, symbolic quality, and search behavior. Core50 reduces evaluation workload by 92.5% and maintains agreement with Full Task Set evaluations. Experiments show that SymbolicArena achieves 72.6% to 86.7% lower approximation error than alternative selectors, further supporting its fidelity to the Full Task Set. Evaluation reveals a substantial gap between numerical fitting and symbolic recovery across current SR methods, suggesting that reliable equation recovery remains an open challenge.
Figures & tables
| Benchmark | Tasks | Focus | HP | CP | VC | ST |
| SRBench ( La Cava et al., 2021 ) | 252 | Algorithm comparison | ||||
| SRBench 2025 ( Imai Aldeia et al., 2025 ) | 24 | Accuracy and complexity | ||||
| SRSD ( Matsubara et al., 2024 ) | 240 | Scientific equation rediscovery | ||||
| LLM-SRBench ( Shojaee et al., 2025b ) | 240 | LLM based equation discovery | ||||
| Nguyen ( Uy et al., 2011 ) Keijzer ( 2003 ) Korns ( 2011 ) Vladislavleva ( 2009 ) | 12 15 15 8 | Equation recovery and extrapolation | ||||
| SymbolicArena | 664 50 | Validated benchmark distillation |
| Numerical | Symbolic | Search | |||||
| Paradigm | Algorithm | ID | OOD | SYM | MIN | EFF | STAB |
| Structured Search | iMCTS | 77.91 | 73.35 | 42.15 | 80.37 | 91.66 | 23.31 |
| QLattice | 34.79 | 25.37 | 29.82 | 55.12 | 87.81 | 31.12 | |
| JAXSR | 37.99 | 30.21 | 32.00 | 75.03 | 99.99 | 97.43 | |
| Evolutionary Search | PySR | 70.34 | 66.28 | 44.60 | 82.62 | 91.94 | 31.03 |
| PyOperon | 38.89 | 30.64 | 25.06 | 44.88 | 97.28 | 15.34 | |
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
| Field group | Stored information |
| Identity and provenance | dataset_id , source family, and subgroup. Original basename and source benchmark |
| Variables and target | Ordered feature names and target name. Number of features and dummy variable indicator |
| Ground truth expression | Raw source formula and executable expression. Normalized symbolic forms and expression tree |
| Data splits | Fixed train, ID test, and OOD test samples. Split sizes and parameters required for deterministic regeneration |
| Structural metadata | Formula complexity and operator types. Variable count bin and OOD type |
| Duplicate metadata | Semantic duplicate group identifier and provenance links used for duplicate auditing |
| Dimension | Categories and task counts |
| Task sources | LLM-SRBench (240), SRSD (232), SRBench 1.0 (133), Keijzer (15), SRBench 2025 first-principles subset (12), Nguyen (12), Korns (12), Vladislavleva (8) |
| Formula types | Rational (282), Trigonometric (137), Polynomial (124), Exponential / logarithmic (86), Mixed (34), Unknown (1) |
| Formula complexity | Simple (284), Moderate (188), Complex (192) |
| Rank | Candidate combination | Mean usability |
| 1 | DSO, PyOperon, iMCTS, uDSR | 0.971250 |
| 2 | DSO, PyOperon, QLattice, uDSR | 0.969375 |
| 3 | PyOperon, iMCTS, QLattice, uDSR | 0.968750 |
| Selector | MAE | Empirical 95% range |
| Uniform random | ||
| Smoothed family stratified | ||
| Metadata diverse | — | |
| Response K-medoids | — | |
| Top information | — | |
| Difficulty balanced | — |
| Idx | Task | Variables | Ground Truth |
|---|---|---|---|
| SRSD 15 tasks | |||
| 1 | Feynman II.27.18 | ||
| 2 | Feynman I.43.31 | ||
| 3 | Feynman II.34.11 | ||
| 4 | Feynman I.18.16 | ||
| 5 | Feynman II.34.2a | ||
| Algorithm | Probe4 | Inclusion (%) | Any optimum (%) | |
| uDSR | Yes | 0.9875 | 100.0 | 100.0 |
| PyOperon | Yes | 0.9600 | 98.8 | 99.3 |
| DSO | Yes | 0.9700 | 88.1 | 91.5 |
| iMCTS | Yes | 0.9675 | 81.4 | 85.4 |
| QLattice | 0.9600 | 30.5 | 35.0 | |
| gplearn | 0.9300 | 1.2 | 1.9 |
| Removed | Replacement | Best | Paradigms | |
| iMCTS | QLattice | 0.969375 | 0.001875 | 3 |
| DSO | QLattice | 0.968750 | 0.002500 | 3 |
| uDSR | QLattice | 0.964375 | 0.006875 | 3 |
| PyOperon | gplearn | 0.963750 | 0.007500 | 3 |
| Perturbed weight | Weight range | Median Jaccard vs. default | Overlapping tasks | Median within- group Jaccard | Full-ranking rate |
| Coverage | 0.724 | 42 | 0.667 | ||
| MeanInfo | 0.710 | 41.5 | 0.667 | ||
| Balance | 0.710 | 41.5 | 0.667 |
| Optimal set | Selections | Share | Jaccard vs. default |
| S0 ( Core50 ) | 2012 | 1.000 | |
| S1 | 1803 | 0.724 | |
| S2 | 1173 | 0.786 | |
| S3 | 12 | 0.754 |
| Statistic | Result |
| Mean Jaccard vs. default | 0.850 |
| Minimum overlapping tasks | 42 / 50 |
| Full-ranking rate | |
| Tasks selected in all runs | 37 |
| Tasks in of runs | 38 |
| Tasks in of runs | 42 |
| ID | OOD | ||||||||
| Paradigm | Algorithm | Clean | 1% | 5% | Clean | 1% | 5% | ||
| Structured Search | iMCTS | 77.91 | 55.08 | 46.69 | -31.22 | 73.35 | 49.40 | 40.73 | |
| QLattice | 34.79 | 32.36 | 29.18 | 25.37 | 23.09 | 20.27 | |||
| JAXSR | 37.99 | 32.62 | 31.00 | 30.21 | 25.91 | 24.56 | |||
| Evolutionary Search | PySR | 70.34 | 48.42 | 41.05 | 66.28 | 41.50 | 35.53 | ||
| PyOperon | 38.89 | 36.29 | 32.97 | 30.64 | 27.85 | 24.42 | |||
| Numerical | Symbolic | Search | |||||
| Paradigm | Algorithm | ID | OOD | SYM | MIN | EFF | STAB |
| Structured Search | iMCTS | 55.08 | 49.40 | 37.22 | 72.88 | 89.34 | 12.82 |
| QLattice | 32.36 | 23.09 | 29.59 | 54.60 | 90.87 | 38.32 | |
| JAXSR | 32.62 | 25.91 | 30.88 | 76.16 | 99.95 | 91.33 | |
| Evolutionary Search | PySR | 48.42 | 41.50 | 30.53 | 60.26 | 86.08 | 6.02 |
| PyOperon | 36.29 | 27.85 | 25.14 | 43.76 | 97.30 | 15.92 | |
| Numerical | Symbolic | Search | |||||
| Paradigm | Algorithm | ID | OOD | SYM | MIN | EFF | STAB |
| Structured Search | iMCTS | 46.69 | 40.73 | 36.43 | 66.26 | 86.11 | 5.24 |
| QLattice | 29.18 | 20.27 | 29.19 | 54.99 | 85.95 | 29.66 | |
| JAXSR | 31.00 | 24.56 | 31.29 | 77.47 | 99.95 | 91.95 | |
| Evolutionary Search | PySR | 41.05 | 35.53 | 30.66 | 59.16 | 85.28 | 6.73 |
| PyOperon | 32.97 | 24.42 | 23.60 | 39.83 | 97.27 | 15.33 | |
| Paradigm | Method | Description |
| Structured Search | iMCTS ( Huang et al., 2025 ) | Monte Carlo tree search for guided expression exploration. |
| QLattice ( Broløs et al., 2021 ) | Probabilistic graph search for compact symbolic models. | |
| JAXSR ( Kitchin, 2024 ) | Structured symbolic model construction with optimization in JAX. | |
| Neural Policy Search | uDSR ( Landajuela et al., 2022 ) | Neural symbolic search with evolutionary refinement. |
| DSO ( Petersen et al., 2021 ) | Policy search using risk seeking optimization. | |
| Evolutionary Search | PySR ( Cranmer, 2023 ) | Population search with complexity aware model selection. |
| Algorithm | Search setting | Operators and basis | Fixed settings |
|---|---|---|---|
| gplearn | Population 1000. | add, sub, mul, div, sqrt, log, sin, cos | 1 job, tournament 20, initial depth 2 to 6, MAE metric, parsimony 0.001, crossover 0.9, subtree, hoist, and point mutation 0.01. |
| PyOperon | Population 500, pool 500. | add, sub, mul, div, aq, exp, log, sin, cos, tanh, sqrt, square, constants, variables | 4 threads, max length 50, max depth 10, tournament 5, keep best reinsertion, LM optimizer, local search probability 1.0. |
| PySR | 8 populations of size 64, 500 cycles per iteration. | Binary . Unary square, cube, exp, log, sin, cos. | Serial execution with 1 process, max size 30, max depth 10, parsimony 0.001, deterministic mode, precision 32, best model selection. |
| QLattice | BIC criterion. | QLattice/Feyn internal symbolic basis | 4 threads, regression mode, BIC model selection, 4 significant digits. |
| DSO | Batch size 1000. | add, sub, mul, div, sin, cos, exp, log | 4 batch cores, reward inv NRMSE, , learning rate , entropy weight 0.03, expression length 4 to 64. |
| uDSR | Batch size 1000, GP meld population 100 for 20 generations. | add, sub, mul, div, sin, cos, exp, log, sqrt, 1.0, constants, poly token | 1 batch core, reward inv NRMSE, , learning rate , entropy weight 0.03, expression length 4 to 100, GP meld and linear poly enabled. |
| Model | Stage | Primary role | Affects search |
| Opus 5 | Postprocessing | Symbolic processing and judgment | No |
| Llama | LLM-SR search | Equation structure generation | Yes |
| Llama | DrSR search | Generation and diagnostic feedback | Yes |
| Axis | Role of Opus 5 |
| ID | None |
| OOD | None |
| SYM | Simplification and equivalence judgment |
| MIN | Simplification before deterministic tree size computation |
| EFF | None |
| STAB | Pairwise structural judgment of final expressions |