Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
Authors: Piyush Jha, Aishik Ghosh, Vijay Ganesh
Organizations: School of Computer Science, Georgia Institute of Technology, USA · School of Physics, Georgia Institute of Technology, USA · Lawrence Berkeley National Laboratory, USA
The upgrade and rewriting of large scientific codebases has traditionally been a major challenge. While evolutionary search with large language models (LLMs) can port and accelerate legacy code, repair feedback in prompts alone does not prevent subsequent candidates from repeating the same errors. We introduce Certificate-Driven Evolutionary Search (CDES), which extends evolutionary search with enforceable restrictions derived from failed candidates, recorded as certificates of assumptions, checker evidence, and justified restrictions. Its control logic enforces these restrictions through rejection, backtracking, and targeted repair while preserving compatible edits. We apply CDES to CPU-to-GPU translation of two particle-simulation functions from the Geant4 toolkit, evaluated with a harness that goes beyond unit tests to combine formal checks, numerical comparisons, physics checks, and GPU safety tests. Generated implementations achieve 13.78x and 23.54x function-level speedups over CPU code, including data conversion and transfers; for one function, GPU throughput exceeds an expert implementation by 14.9%, reaching 16.1% when complementary components are combined. In an ablation over execution settings, certificate feedback increases the fraction of candidates passing required correctness checks from 55% to 90%.
Figures & tables
Function
Geant4 CPU
CDES-generated CUDA
Speedup
Compton
30.162
2.189
13.78 ×
Fermi
340.090
14.447
23.54 ×
Table 1: Generated CUDA achieves 13.78 × and 23.54 × speedups over standalone CPU functions under the stated timing protocol. Times are in milliseconds. GPU times sum separately measured conversion, transfers, and computation, not full-simulation runtime.
Version
Execution time
GPU time
GPU throughput gain
vs. Celeritas
vs. CDES
Celeritas
2.4753
0.5093
CDES-generated
2.4000
0.4433
14.9%
Hybrid
2.3905
0.4387
16.1%
1.2%
Table 2: Combining Celeritas components with generated code gives the highest GPU throughput. Times are medians in milliseconds over six paired timing rounds. Throughput gains are medians of per-round time ratios, not ratios of displayed median times. Execution time directly measures transfers, computation, and completion waits in the shared benchmark. Settings are selected separately for each timing metric. Celeritas and the hybrid use a shared direction correction.
Without certificates
With certificates
Proposals passing correctness checks
55%
90%
Passing with throughput gain >10%
50%
80%
Passing with throughput gain >20%
30%
35%
Best throughput gain on withheld inputs
66.5%
68.6%
Model-call cost (USD)
$0.0535
$0.0782
Table 3: Certificate feedback produces more proposals that pass correctness checks and improve performance. Costs cover this ablation’s Qwen3-Coder-Plus calls.