Organizations: State Key Laboratory of Novel Software Technology, Nanjing University, China · School of Artificial Intelligence, Nanjing University, China · Huawei Noah’s Ark Lab, China
Large language models are increasingly participating in complex real-world tasks in the form of algorithm-design agents, designing and refining algorithms. Many successful algorithm-design agents adopt pure in-context evolutionary frameworks, but they may quickly plateau in domains that require specialized knowledge. Parametric adaptation offers a way to internalize specialized knowledge, but conventional training requires abundant domain-specific corpora while high-quality algorithms are scarce in complex algorithm-design scenarios. In this paper, we propose sample-efficient parametric self-evolution where agents can explore and learn from self-generated algorithms. First, we characterize in-context evolutionary stagnation and analytically propose the Improvement Chain proposition, showing how learning successive self-generated algorithms can locally increase the likelihood of neighboring algorithms. Motivated by this local-transfer perspective, we further propose Population-Curated Policy Optimization (PCPO) to utilize a global population and a hybrid policy update scheme for retaining and reusing high-quality, diverse self-generated algorithms, shifting the policy towards stronger algorithms. In the task of learning rate schedule design for global placement in electronic design automation, trained only on 4 chip cases, PCPO outperforms the state-of-the-art in-context evolutionary methods (e.g., OpenEvolve and ShinkaEvolve) on average across 16 chip cases. With an 8B-size base model, PCPO achieves competitive performance compared to frontier closed-source models such as GPT-5.5. PCPO also reduces inference-time token cost by internalizing grounded domain knowledge and prompt distillation. Moreover, PCPO achieves significant speedups on four GPU kernel designs, with an average of 8.27× speedup against the PyTorch Eager baseline.
Figures & tables
Figure 1: Average online improvement on 4 chips against total training and inference cost. PCPO achieves the largest improvement within acceptable computational cost, and the trained model can generalize to other chip cases with minimal inference-time token cost as shown in Section 6.1 . Frozen-model methods underperform the base model due to the stagnation described in Section 4.1 and incur substantial computational cost due to repeated generation refinement and long thinking.
Figure 2: Diagnostic experiments based on Qwen3-8B.
Figure 3: An illustration of the ideal Improvement Chain where each successor in the chain gradually approaches A∗ . The optimal algorithm is approached by progressively increasing the reachability of successive chain nodes, generating and learning on chain nodes one after another.
Figure 4: The workflow of PCPO. The left part illustrates the Global Population, filtering promising candidates while avoiding incompatible reward scales and difficulties. The right part illustrates the Hybrid Policy Update scheme, supporting sustained improvement, exploration, and rapid correction.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Symbol / name
Value
Reward and legality
Illegal / crash / missing API
—
−109
Not achieving target overflow
—
−1000
Valid placement
r
DREAMPlace get_fitness (higher better)
Illegal threshold (crash band)
rinv
−108
PCPO
Appendix
Table 4: Hyperparameters of PCPO with Qwen3-8B on Global Placement. Every T on-policy steps, we additionally run an off-policy update on the global-population top- k programs ranked by composite score S=αr~+βd , where r~ is min–max normalized reward and d is semantic diversity. Near-duplicates with diversity ≤δmax are merged unless the fitness gain is at least γmin .
Figure 5: A representative OpenEvolve thinking trace on bigblue1 versus a strong direct generation based on Qwen3-8B. Panel (a) is lightly abridged from the model’s <think> block; highlighted claims H1–H3 are factually false or internally inconsistent. Panel (b) is the program emitted from that trace: extra multiplicative decay on learning_rate_prev while overflow exceeds 0.07 compounds into ηt=η0(0.995⋅0.95)t , the optimizer stalls, and overflow never reaches the 0.07 target. Panel (c) reconstructs the expert schedule η0⋅0.995t and applies non-compounding state-dependent modulators; HPWL improving ( log_hpwl < log_hpwl_prev ) shrinks the step, the opposite of (b). Ellipses in (a) mark omitted sentences; wording is otherwise verbatim.
ID
Type
Typical claim in <think> vs. the working policy
H1
Causal inversion
Overflow remains ≈0.65⇒ “LR is not small enough”; extra ×0.99/0.95/0.9 on learning_rate_prev . Direct: non-compounding soft scale on η0⋅0.995t .
H2
Semantic inversion
“If log_hpwl is decreasing (meaning hpwl is getting worse)”. log is monotone: a drop is an improvement .
H3
Phantom input
Invents log_gradient_norm_prev , which is not in the API.
Treats −88.90 / −89.26 as “better than baseline” −88.56 (higher is better).
Appendix
Table 5: Recurring errors in OpenEvolve chain-of-thought on bigblue1 . “Direct” denotes the best sample produced by the base model directly from the same task (Fig. 5 c).
Figure 6: A second OpenEvolve trace on the legal island. H5 inverts fitness polarity ( −89.26 is worse than baseline −88.56 ). H4 compares log∥∇∥ to logHPWL , two unrelated quantities, then treats the predicate as “gradient is decreasing.” The model later notes that the previous gradient is unavailable, yet still emits the comparison.
Figure 7: Joint log-probability assigned by Qwen3-8B to five global placement learning-rate policies under a shared prompt, before and after one training update step on the anchor program. The orange bar is the log-probability of the anchor produced by PCPO; blue bars are the log-probabilities of short-edit neighbors with respect to the anchor; bars with lighter colors show the log-probability after one training update step. Token-level edit distance from the anchor is annotated on each bar.
Figure 8: Detailed modifications and improvement of the global placement learning rate schedule algorithm evolution made by PCPO on adaptec1.
Method
adaptec1
bigblue1
superblue1
superblue7
Avg rank
OpenEvolve
70.99
87.53
389.18
554.48
4.25
ShinkaEvolve
71.10
88.29
389.18
551.10
4.63
JitRL
70.99
88.36
387.61
556.13
4.38
ThetaEvolve
70.68
87.83
388.79
562.56
4.25
GRPO
70.49
87.51
388.09
551.60
2.50
PCPO
70.04
87.23
385.64
547.11
1.00
Appendix
Table 6: Average inference-time 64-rollout best HPWL ( ×106 ) of 4 independent runs on four global placement cases. Lower is better.
Method
adaptec1
bigblue1
superblue1
superblue7
Average
PCPO
70.04
87.23
385.64
547.11
–
Qwen3-8B + Skill
70.94
87.58
388.76
553.16
–
Relative gap (%)
+1.29
+0.40
+0.81
+1.11
+0.90
Appendix
Table 7: Average inference-time 64-rollout best HPWL ( ×106 ) of 4 independent runs produced by PCPO versus Qwen3-8B equipped with a generalizable skill distilled by GPT-6-Astra from PCPO’s training logs. Lower is better.
Figure 9: Ablation study on the training performance of the Regularized Off-policy Update and the Refreshing On-policy Update in the Hybrid Policy Update module of PCPO.
Figure 10: Best-so-far normalized reward along the training trajectories of PCPO and GRPO on 4 training cases.
Figure 11: Inference-time token cost statistics of PCPO and two evolutionary methods.
Figure 12: Diagnostic experiments based on Gemma4-E4B.