SciWalker: Synthesizing Scientific Coding Problems with Operator Graphs and Execution Feedback
Organizations: The Hong Kong University of Science and Technology · Shanghai AI Laboratory · Tsinghua University · Nanyang Technological University
Abstract
Improving the scientific coding capabilities of large language models (LLMs) requires high-quality training data. However, such data remain scarce because manually authoring realistic problems is costly and time-consuming, while systematically covering diverse scientific domains and algorithmic combinations remains challenging. To address this, we introduce SciWalker, a framework for synthesizing scientific coding problems through operator-chain sampling and execution feedback. The framework combines scientific library interfaces with operation modes to instantiate operators, organizes them into operator graphs, and samples operator chains as computational workflow cues. Guided by these cues, we adopt LLMs to generate scientifically grounded problem statements, reference solutions, and tests, with failed generations iteratively repaired using execution feedback. By combining structured workflow composition with verification and quality review, SciWalker enables scalable task generation while promoting scientific grounding, computational diversity, and executability. Using this framework, we construct 8,178 high-quality problems spanning 5 scientific domains and 32 subdomains. To evaluate their training utility, we conduct reinforcement learning on Qwen3.5-9B using the GSPO algorithm. This training improves SciCode subproblem accuracy by 9.9 percentage points, from 29.3% to 39.2%, with gains across scientific code generation, code repair, and reasoning benchmarks. The code for SciWalker is available at https://github.com/lichenx1/SciWalker.
Figures & tables
| Item | Shared setting |
|---|---|
| Learning rate | |
| Optimizer parameters | Adam, , , weight decay |
| GSPO clipping | , |
| Gradient clipping | 1.0 |
| KL / entropy terms | KL reward, KL loss, and entropy regularization disabled |
| Training budget | 150 steps, 32 problems per step, 8 trajectories per problem |
| Scientific computing and code generation | |||||
|---|---|---|---|---|---|
| Model | SciCode | DS-1000 | HumanEval | LCB Code Gen. v6 | APPS Introductory |
| Baseline | 29.3 | 45.6 | 94.7 | 61.9 | 69.1 |
| PPO | 32.6 (+3.3) | 49.4 (+3.8) | 93.5 (-1.2) | 57.6 (-4.3) | 72.3 (+3.2) |
| GSPO | 39.2 (+9.9) | 62.9 (+17.3) | 97.8 (+3.1) | 74.9 (+13.0) | 84.6 (+15.5) |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Substep | Computational task |
|---|---|
| 1 | Compute the residual vector and Jacobian matrix |
| 2 | Compute one Gauss–Newton parameter update |
| 3 | Solve for the parameters using Gauss–Newton iterations with backtracking |
| 4 | Compute the condition number of the approximate Hessian |
| 5 | Filter the data using an absolute residual threshold |
| Version | Main changes and validation results during generation |
|---|---|
| Initial generation | The problem statement directly hints at interfaces; parameter recovery, differential testing, and other checks fail |
| First repair | Backtracking is added to student and oracle, and one differential test case is modified; parameter recovery errors remain after outlier handling |
| Second repair | A lower-bound constraint on the decay rate is added to both implementations; full validation still fails |
| Third repair | The normal equations are solved using a least-squares method, and the iteration stopping criterion is adjusted; full validation passes |
| Stage | Retained | Reduction | Cumulative retention |
|---|---|---|---|
| Candidate operator chains | 66,889 | — | 100.0% |
| Problem-generation potential assessment | 21,543 | 45,346 | 32.2% |
| Scientific context selection and review | 21,527 | 16 | 32.2% |
| Training problem generation and execution validation (including repair) | 17,166 | 4,361 | 25.7% |
| Problem statement polishing and quality grading (retaining only high-quality problems) | 8,178 | 8,988 | 12.2% |
| Evaluation | Tokens | Result |
|---|---|---|
| First | 1,496 | Generates code but omits the velocity-state update; fails the tests |
| Second | 1,326 | Generates code but omits the velocity-state update; fails the tests |
| Third | 32,768 | Repeatedly discusses dependency imports, reaches the length limit, and produces no final code |
| Mean / Total | 11,863.33 | 0/3 pass |
| Domain | Libraries |
|---|---|
| Mathematics | arch, ArviZ, CVXPY, geomdl, JAXopt, NLopt, NumPy, PyGMO, PyLops, PyMC, pymoo, Pyomo, PyVista, Riskfolio-Lib, scikit-learn, SciPy, SfePy, statsmodels, trimesh |
| Physics | alchemlyb, ASE, Astropy, Awkward Array, boost-histogram, Cirq, coffea, DecayLanguage, Diffractio, discretize, FiPy, Gala, galpy, GWpy, HCIPy, healpy, hist, Lightkurve, mplhep, NumPy, particle, POPPY, py_pol, PyDMD, PySINDy, Qiskit, QuTiP, qutip-qip, REBOUND, SciPy, sisl, Stim, SymPy, zfit |
| Chemistry | Biopython, Cantera, datamol, DeepChem, edlib, molmass, MolVS, OpenMM, parasail, periodictable, pysam, PySCF, RDKit, scikit-bio, SciPy |
| Biology | agentpy, Biopython, COBRApy, DendroPy, edlib, GillesPy2, Mesa, msprime, NumPy, parasail, PyMC, pysam, scikit-bio, SciPy, tskit |
| Materials science | CHGNet, Dans_Diffraction, datamol, jarvis-tools, jobflow, matminer, molmass, periodictable, pyFAI, pymatgen, RDKit |