Organizations: Department of Electronic Engineering, BNRist, Tsinghua University, Beijing, China · College of AI, Tsinghua University, Beijing, China · Department of Earth System Science, Tsinghua University, Beijing, China
Recent agentic symbolic regression approaches increasingly rely on large language models to analyze data, select scientific operations, and refine hypotheses over long search trajectories. In such systems, performance depends not only on the underlying model and search strategy, but also on the runtime infrastructure that supports scientific search. We introduce SRHarness, a domain-specific harness for agentic symbolic regression built around three mechanisms: composable scientific actions that provide a common interface over raw, transformed, and candidate-derived quantities; persistent scientific state that retains evaluated hypotheses and exposes compact model-facing views; and trajectory lifecycle management that coordinates continuation, branching, restart, and termination. On LLM-SRBench, SRHarness consistently improves both numerical generalization and symbolic recovery under matched LLM backbones. With DeepSeek-v4-flash-0731, it achieves 93.69% symbolic accuracy on LSR-Transform, compared with 62.16% for SR-Scientist, and retains 72.97% accuracy on an anonymized variant that removes scientific descriptions and variable semantics, versus 39.64% for SR-Scientist. Under the same DeepSeek-v4-flash-0731 backbone, SRHarness also substantially outperforms Codex (72.97% vs. 20.72%) and reaches performance comparable to Codex with GPT-5.5, while simply providing Codex with the same scientific tools does not reproduce this advantage. These results show that effective agentic symbolic regression depends not only on models or tools, but also on structured runtime support for organizing scientific actions, accumulated hypotheses, and long-horizon search.
Figures & tables
Figure 1: Overview of the three core mechanisms in SRHarness.
Base Model
Method
Physics (Acc0.1↑)
Chemistry (Acc0.1↑)
Biology (Acc0.1↑)
Material (Acc0.1↑)
Overall
ID
OOD
ID
OOD
ID
OOD
ID
OOD
SA (%) ↑
Time (min)
(no LLM)
PySR
86.83%
82.87%
92.49%
66.68%
75.52%
59.26%
99.90%
100.00%
6.20%
25.46
Deepseek-v4-flash-0731
LLM-SR
71.41%
60.53%
72.02%
47.52%
68.78%
55.30%
95.00%
92.10%
0.00%
57.67
IGSR
72.37%
61.92%
91.71%
76.14%
86.95%
71.84%
98.89%
100.00%
0.78%
24.02
SR-Scientist
77.48%
69.83%
78.08%
60.63%
70.17%
60.43%
95.76%
97.19%
0.78%
16.62
Ours
84.46%
78.00%
95.24%
78.87%
95.14%
87.09%
95.06%
96.00%
6.20%
11.11
Table 1: Performance on LSR-Synth across domains, reporting ID/OOD numerical accuracy, symbolic accuracy (SA), and average search time.
Base Model
Method
LSR-Transform
Test Acc0.1 (%)↑
Test NMSE↓
SA (%) ↑
Complexity↓
Time (min)
(no LLM)
PySR
77.33%
1.44E-04
35.14%
17.57
25.49
Deepseek-v4-flash-0731
LLM-SR
79.48%
2.80E-14
45.05%
64.46
72.74
IGSR
86.44%
1.05E-14
48.65%
42.14
23.67
SR-Scientist
79.62%
2.50E-14
62.16%
17.50
6.72
Ours
88.37%
1.24E-14
93.69%
14.64
12.47
Table 2: Performance on LSR-Transform, reporting numerical performance, symbolic accuracy (SA), expression complexity, and average search time.
Base Model
Method
LSR-Transform (Anonymized)
Test Acc0.1 (%)↑
Test NMSE↓
SA (%) ↑
SA Retention (%)↑
Complexity↓
Time (min)↓
Time Ratio (×)↓
(no LLM)
PySR
77.33%
1.44E-04
35.14%
100%
17.57
25.49
1
Deepseek-v4-flash-0731
LLM-SR
67.59%
1.21E-03
12.61%
28%
87.85
286.13
3.93
IGSR
47.36%
8.36E-02
13.51%
28%
34.05
55.60
2.35
SR-Scientist
73.74%
6.00E-05
39.64%
64%
30.73
22.30
3.32
Ours
91.69%
1.24E-14
72.97%
78%
33.36
12.34
0.99
Table 3: Performance on LSR-Transform-Anon after removing problem-specific semantic cues, reporting numerical performance, symbolic accuracy (SA), SA retention, expression complexity, and search efficiency.
Base Model
Harness
LSR-Transform (Anonymized)
Test Acc0.1 (%)↑
Test NMSE↓
SA (%) ↑
Complexity↓
Deepseek-v4-flash-0731
Ours
91.69%
1.24E-14
72.97%
33.36
Codex
66.74%
1.59E-02
20.72%
91.72
GPT-5.5
Codex
93.20%
1.63E-14
68.47%
2547.55
Codex (+tools)
91.45%
2.49E-14
64.86%
902.47
Table 4: Comparison of SRHarness and Codex on LSR-Transform-Anon under different model and tool configurations.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Base Model
Method
Physics (NMSE↓)
Chemistry (NMSE↓)
Biology (NMSE↓)
Material (NMSE↓)
ID
OOD
ID
OOD
ID
OOD
ID
OOD
(no LLM)
PySR
8.01E-06
7.57E-05
5.07E-06
4.59E-01
2.50E-06
1.17E-02
1.31E-09
2.56E-07
Deepseek-v4-flash-0731
LLM-SR
8.71E-05
1.99E-03
1.08E-04
9.51E-01
7.79E-06
3.90E-02
8.31E-08
4.41E-06
IGSR
1.34E-04
1.43E-03
9.28E-08
2.84E-03
1.87E-08
6.84E-06
1.94E-11
1.54E-08
SR-Scientist
1.85E-05
3.83E-04
3.05E-05
8.26E-02
2.81E-05
1.02E-02
2.39E-09
5.13E-07
Ours
1.48E-05
8.21E-05
1.80E-08
8.70E-04
1.75E-07
2.25E-05
3.21E-14
1.91E-10
Appendix
Table 5: Domain-level ID and OOD NMSE on LSR-Synth.
Base Model
Method
Time (min)
Token (M)
Cost ($)
(no LLM)
PySR
25.46
/
/
Deepseek-v4-flash-0731
LLM-SR
57.67
1.00
0.05
IGSR
24.02
0.125
0.0198
SR-Scientist
16.62
1.63
0.05
Ours
11.11
1.53
0.03
GLM-5.3-flash
LLM-SR
23.06
0.28
0.09
Appendix
Table 6: Runtime, token usage, and API cost on LSR-Synth.
Method
Base Model
LSR-Synth Performance
Parameters
Release Date
Benchmark Score
ID Acc 0.1 ↑
ID Median NMSE ↓
SA ↑
TerminalBench 2.1
DeepSWE
GPQA Diamond
Ours
DeepSeek-v4-flash-0731
91.51%
1.68×10−7
6.20%
284B / 13B
2026-07-31
82.7
54.4
89.9
DeepSeek-v4-pro-0813
95.00%
5.00×10−9
11.63%
1.6T / 49B
2026-08-13
87.9
62.7
92.4
DeepSeek-v4.1-flash
95.52%
1.96×10−12
15.50%
552B / 8–16B
2026-09-10
90.6
74.2
90.9
GLM-5.3-flash
92.56%
1.84×10−9
10.85%
320B / 18B
2026-08-26
84.3
63.4
91.2 †
GLM-5.3
93.47%
7.58×10−10
10.85%
744B / 40B
2026-08-14
88.2
66.9
88.1
Appendix
Table 7: SRHarness performance with five fully evaluated LSR-Synth backbones and selected model attributes. Parameter counts are total/activated parameters per token; DeepSeek-v4.1-flash activates 8B parameters during prefill and 16B during decoding. Benchmark scores use the publicly reported evaluation settings of each model and are therefore descriptive rather than a strictly controlled cross-model evaluation.
Base Model
Method
Time (min)
Token (M)
Cost ($)
(no LLM)
PySR
25.49
/
/
Deepseek-v4-flash-0731
LLM-SR
72.74
1.015
0.0535
IGSR
23.67
0.140
0.0224
SR-Scientist
6.72
0.9151
0.028
Ours
12.47
1.143
0.0204
GLM-5.3-flash
LLM-SR
27.56
0.3454
0.1176
Appendix
Table 8: Runtime, token usage, and API cost on LSR-Transform.
Base Model
Method
Time (min)↓
Token (M)
Cost ($)
(no LLM)
PySR
25.49
/
/
Deepseek-v4-flash-0731
LLM-SR
286.13
2.253
0.1187
IGSR
55.60
0.163
0.0269
SR-Scientist
22.30
1.587
0.0449
Ours
12.34
1.300
0.0252
Deepseek-v4-flash-0731
Ours
12.34
1.300
0.0252
Appendix
Table 9: Runtime, token usage, and API cost on LSR-Transform-Anon.
Base Model
Harness
LSR-Transform (Anonymized)
Time (min)
Token (M)
Cost ($)
Deepseek-v4-flash-0731
Ours
12.34
1.300
0.0252
Codex
12.81
1.146
0.0333
GPT-5.5
Codex
5.54
0.618
/
Codex (+tools)
7.47
1.231
/
Appendix
Table 10: Resource usage for the SRHarness–Codex comparison on LSR-Transform-Anon.
Base Model
Harness
LSR-Transform (Anonymized)
Test Acc0.1 (%)↑
Test NMSE↓
SA (%) ↑
Complexity↓
Time (min)
Token (M)
Cost ($)
Deepseek-v4-flash-0731
Ours (run 1)
91.69%
1.24E-14
72.97%
33.36
12.34
1.300
0.0252
Ours (run 2)
92.30%
1.09E-14
74.77%
18.67
15.92
1.379
0.0403
Ours (run 3)
92.77%
1.09E-14
75.68%
17.55
16.55
1.382
0.0404
Ours (3 runs average)
92.25%
1.24E-14
74.47%
23.19
14.93
1.354
0.0353
Appendix
Table 11: Variation across the three runs used for the full-configuration ablation reference.
Base Model
Harness
LSR-Transform (Anonymized)
Test Acc0.1 (%)↑
Test NMSE↓
SA (%) ↑
Complexity↓
Time (min)
Token (M)
Cost ($)
Deepseek-v4-flash-0731
Ours (Full, 3 runs average)
92.25%
1.24E-14
74.47%
23.19
14.93
1.354
0.0353
Ours (Code executor only)
88.15%
1.66E-14
68.47%
26.39
14.34
1.124
0.0349
Ours (Additional fitting actions)
92.39%
1.07E-14
79.28%
33.16
17.39
1.449
0.0419
Ours (w/o Expression-based views)
90.29%
1.54E-14
67.57%
25.07
11.59
1.238
0.0311
Appendix
Table 12: Ablation of the scientific action space on LSR-Transform-Anon.
Base Model
Harness
LSR-Transform (Anonymized)
Test Acc0.1 (%)↑
Test NMSE↓
SA (%) ↑
Complexity↓
Time (min)
Token (M)
Cost ($)
Deepseek-v4-flash-0731
Ours (Pareto view, 3 runs average)
92.25%
1.24E-14
74.47%
23.19
14.93
1.354
0.0353
Ours (Best-formula view)
90.84%
1.61E-14
66.67%
25.75
17.49
1.341
0.0400
Ours (w/o Best-formula view)
93.12%
1.24E-14
72.97%
19.56
15.04
1.297
0.0385
Appendix
Table 13: Ablation of the model-facing scientific-state view on LSR-Transform-Anon.
Base Model
Harness
LSR-Transform (Anonymized)
Test Acc0.1 (%)↑
Test NMSE↓
SA (%) ↑
Complexity↓
Time (min)
Token (M)
Cost ($)
Deepseek-v4-flash-0731
Ours (Default, R1-C1-L30-K1, 3 runs average)
92.25%
1.24E-14
74.47%
23.19
14.93
1.354
0.0353
Ours (More restarts, R2-C1-L15-K1)
91.34%
1.14E-14
66.67%
20.35
27.60
1.050
0.0379
Ours (More independent branches, R1-C2-L15-K1)
91.80%
1.24E-14
72.07%
22.09
22.98
0.950
0.0358
Ours (Local response sampling, R1-C1-L15-K2)
95.38%
9.76E-15
73.87%
16.30
23.73
1.296
0.0411
Appendix
Table 14: Ablation of trajectory-budget allocation on LSR-Transform-Anon.
Symbolic regression aims to discover closed-form equations from data, but existing LLM-guided methods often rely on a unified proposal loop that compresses heterogeneous search failures into a scalar score and a single prompt. We propose A-SR, a self-evolving agentic framework that shifts the control unit from expression edits to role-conditioned evidence views. A-SR coordinates formula discovery through routing among coordination protocols, an online evaluator-reward role policy, and state-routed process memory. During search, evaluator feedback characterizes reliability and productivity, updates role-level utilities, and routes elite motifs, failure traces, and validity diagnostics to different agents. The framework self-evolves at two timescales: within a run, it adapts the search process without updating LLM parameters; across runs, recorded trajectories can be distilled into open-source LLMs as role-conditioned proposal priors. Averaged over the four LSR-Synth scientific domains in LLM-SRBench, A-SR improves Acc@0.01 over baselines from 25.79% to 48.30% with Llama3.1-8B, while A-SR-LoRA improves the corresponding Qwen3-4B result from 24.58% to 38.29%. On four real-world scientific discovery tasks, A-SR obtains the best in-distribution or out-of-distribution normalized mean squared error on 7 of 8 reported metrics.
Wenxiao Zhao, Dong Liu, Kaiyi Xu +9
Shanghai AI Laboratory · University of California at Los Angeles · Tongji University
Symbolic regression (SR) discovers compact mathematical expressions from data, yet recent LLM-based evolutionary methods remain sample-inefficient because they rely mainly on scalar feedback such as MSE. We identify a core limitation: existing methods conflate candidate proposal with search guidance, requiring the LLM to infer how to evolve an expression, diagnose its errors, and reuse past experience from a single score. To address this, we propose Deliberate Evolution (DE), an agentic framework that decouples symbolic generation from search control. DE guides LLM proposals with adaptive operators for search direction, analytical tools for structural diagnosis, and reflective memory for trajectory-level experience. Experiments on LLM-SRBench show that DE consistently outperforms representative LLM-based SR baselines across diverse scientific domains while using only 40% of the standard sample budget.
Xinyu Pang, Zhanke Zhou, Xuan Li +5
Beijing National Research Center for Information Science and Technology (BNRist), Department of Automation, Tsinghua University, Beijing, P.R. China · TMLR Group, Department of Computer Science, Hong Kong Baptist University · Lenovo Research
Scientific equation discovery must combine broad domain priors with strict numerical testing. Symbolic regression supplies numerical grounding but faces a combinatorial search space, whereas many language-model systems ask the model to propose or select formulas directly. We test a different division of labour. We compare role specifications in which the language model acts as equation author, candidate decider or search controller, alongside end-to-end language-model and purely numerical baselines. In the controller setting we propose here, implemented as LLM-PySR, language models specify variables, operators, transformations and search depth; symbolic regression enumerates and fits expressions; and deterministic metrics govern retention. Across 74 AI-Feynman equations and seven complex formula-recovery tasks, search control achieved the strongest observed balance of accuracy, complexity, stability and cost. On an independent battery dataset, LLM-PySR identified a compact piecewise-linear relation between early voltage-curve displacement and cycle life. The results suggest that language models should shape hypothesis exploration rather than decide which equations survive.
Zikai Xie, Wenmei Li, Man Luo +2
State Key Laboratory of Precision and Intelligent Chemistry, Hefei National Research Center for Physical Sciences at the Microscale, School of Chemistry and Materials Science, University of Science and Technology of China, Hefei, China. · Center for Scientific Intelligence Innovation, Hefei, China · School of Chemistry, School of Computer Science, University of Birmingham, Birmingham, UK