From Search to Signal: Online Post-Training in Automatic Heuristic Design
Organizations: Wenzhou University
Abstract
Large language model (LLM)-based automatic heuristic design (AHD) iteratively proposes and refines heuristics, pairing design rationales with executable code. Task-specific evaluators assess programs; execution outcomes and performance scores guide search. Many AHD systems keep the generator frozen; EvoTune and Co-Evolution of Algorithms and Language Model (CALM) instead update it from evaluated candidates. When such outcomes drive reinforcement learning with verifiable rewards (RLVR), they create a search-coupled loop: the evaluated candidate stream supplies both search-state updates and training signals for the model that generates future candidates. Yet validity and performance do not uniquely determine useful model updates; converting them into learning signals must account for the prompt and evolving search state that produced each candidate. We formulate online post-training of small open-weight LLMs in AHD as context-dependent signal construction and develop alternative mappings from program validity, task performance, and generation context to update signals. Using shared evaluated rollouts and matched update budgets, controlled experiments across AHD tasks and model families compare these mappings with online post-training baselines, testing their effects on validity, performance among valid proposals, and the yield of valid proposals that improve under contextual comparisons. Complementary checkpoint, frozen-search, and live-system evaluations assess whether proposal-level gains appear in updated checkpoint behavior and subsequent search, rather than arising solely from accumulated search state. A resource-matched comparison under pre-specified cost accounting tests whether online updating adds value beyond additional search with a frozen generator. Together, this design avoids treating end-to-end search gains alone as evidence of stronger heuristic-design capabilities.
Figures & tables
| Construction | Evidence and reference | Reward supplied to GRPO | Role |
|---|---|---|---|
| Native CALM | Typed failures; nonlinear comparison with | Released piecewise reward | Online baseline |
| Factorized Validity–Quality | Binary validity; signed, RMS-scaled for valid candidates | Separate feasibility from quality | |
| Pre-Generation Tail-Weighted | Typed outcomes; pre-generation score novelty, repetition, and gap above | Emphasize the group upper tail | |
| Search-Exposure Residual | Typed validity; ; counterfactual one-step parent exposure | Search-derived mechanism probe | |
| Frozen CALM | Same search and evaluator; updates disabled | — | No-update control |
| Search outcomes | Proposal stream | ||||
|---|---|---|---|---|---|
| Condition | Final best | Trajectory AUC | Valid (%) | Valid-only perf. | Improve yield (%) |
| Frozen | -6.2329 0.0105 | -6.2417 0.0046 | 60.85 0.55 | -9.8800 0.8639 | 2.83 0.50 |
| Native | -6.2090 0.0289 | -6.2203 0.0196 | 93.92 1.68 | -6.4966 0.0986 | 2.98 1.16 |
| Factorized | -6.2041 0.0116 | -6.2179 0.0087 | 95.60 0.64 | -6.7488 0.1168 | 3.60 1.25 |
| Tail-weighted | -6.1963 0.0248 | -6.2123 0.0148 | 85.12 2.50 | -7.0468 0.2214 | 4.63 0.33 |
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
| Layer | Tasks | Models | Seeds | Shared evidence | Claim supported |
|---|---|---|---|---|---|
| Normalizer replay | TSP, CVRP, OBP, OP | recorded generators | – | evaluated records | whether reward differences survive group normalization |
| Matched update | TSP, CVRP | Qwen2.5-7B, DeepSeek-Coder-7B | 1/cell | tokens, masks, reference log probabilities, update budget | whether an advantage difference reaches parameter updates |
| Formal live cohort | TSP | Qwen2.5-7B | 3 | starts, seeds, 500 groups | proposal-stream and equal-horizon live-search effects |
| Frozen restart | TSP | Qwen2.5-7B | 3 | final adapter, reset population, zero updates | checkpoint effect after removing accumulated live population |
| Directional breadth | CVRP; TSP (DeepSeek only) | Qwen2.5-7B (CVRP); DeepSeek-Coder-7B | 1/cell | task/model protocol within cell | directional breadth across three task–model cells, not multi-seed efficacy |
| Resource anchor | TSP | Qwen2.5-7B | 3 | preregistered timing-derived group budgets | endpoint sensitivity to allocating compute to extra Frozen search |
| Python | 3.10.20 | PyTorch | 2.5.1 |
| Transformers | 4.49.0 | TRL | 0.15.1 |
| PEFT | 0.14.0 | Unsloth | 2025.3.18 |
| vLLM | 0.7.3 | Ray | 2.40.0 |
| NumPy | 1.26.4 | bitsandbytes | 0.45.2 |
| Host class | Accelerator | CPU / memory / OS | Evidence role |
|---|---|---|---|
| Local workstation | 3 NVIDIA RTX 4090, 24,564 MiB each; driver 550.144.03 | Intel Xeon w5-2455X, 12 cores/24 threads; 125 GiB RAM; Linux 5.15 | most formal live runs, including at least one run from every condition; Qwen breadth; Gate-4 prefixes; API-served frozen runs; common-context probe |
| Shared GPU host | 4 NVIDIA RTX 3090, 24,576 MiB each | CPU model, RAM, and OS not archived in formal manifests | selected Tail-weighted live/restart runs; one Frozen formal seed; DeepSeek-Coder breadth; sequential Frozen references |
| A100 host | 1 NVIDIA A100, 40,960 MiB | CPU model, RAM, and OS not archived in formal manifests | selected mechanism and breadth cells; one Tail-weighted formal seed |
| Search / generation | Value | Optimization | Value |
|---|---|---|---|
| Task / live model | TSP / Qwen2.5-7B-Instruct | Adapter | LoRA rank 32, alpha 64 |
| Seeds | 42, 3407, 1926000 | Target modules | q, k, v, o, gate, up, down projections |
| Groups / completions | 500 / 2,000 | Quantization | 4-bit base model |
| Prompts per step / group size | 1 / 4 | Learning rate | , constant |
| Population size | 10 | Optimizer | 8-bit AdamW |
| Operator weights | simplification 1; injection 1; replacement 2; crossover 4 | Adam betas / weight decay | 0.9, 0.99 / 0.1 |
| Condition | Final best | Trajectory AUC | Valid (%) | Valid-only perf. | Improve (%) |
|---|---|---|---|---|---|
| Frozen | -6.2329 0.0105 | -6.2417 0.0046 | 60.85 0.55 | -9.8800 0.8639 | 2.83 0.50 |
| Native | -6.2090 0.0289 | -6.2203 0.0196 | 93.92 1.68 | -6.4966 0.0986 | 2.98 1.16 |
| Factorized | -6.2041 0.0116 | -6.2179 0.0087 | 95.60 0.64 | -6.7488 0.1168 | 3.60 1.25 |
| Tail-weighted | -6.1963 0.0248 | -6.2123 0.0148 | 85.12 2.50 | -7.0468 0.2214 | 4.63 0.33 |
| Task | Model | Variant | Final best | Trajectory AUC | Valid (%) | Valid-only perf. | Improve (%) |
|---|---|---|---|---|---|---|---|
| CVRP | DeepSeek Coder | Frozen | -9.0676 | -9.3136 | 49.20 | -12.2861 | 3.35 |
| CVRP | DeepSeek Coder | Native | -8.7337 | -8.8805 | 81.00 | -10.6994 | 3.45 |
| CVRP | DeepSeek Coder | Factorized | -8.9977 | -9.0951 | 81.50 | -11.2008 | 2.85 |
| CVRP | Qwen2.5 | Frozen | -9.1271 | -9.4523 | 63.90 | -11.0669 | 5.60 |
| CVRP | Qwen2.5 | Native | -9.1517 | -9.4046 | 88.85 | -10.3968 | 5.00 |
| CVRP | Qwen2.5 | Factorized | -9.1097 | -9.4950 | 94.80 | -10.5522 | 5.60 |
| Seed | Online condition | Groups (online/Frozen) | Online final | Frozen final | final | valid (pp) | valid-only |
|---|---|---|---|---|---|---|---|
| 42 | Native | 500/883 | -6.2176 | -6.2167 | -0.0010 | +27.01 | +0.742 |
| 42 | Factorized | 500/762 | -6.2221 | -6.2167 | -0.0055 | +25.42 | +0.545 |
| 3407 | Native | 500/572 | -6.2217 | -6.2303 | +0.0086 | +34.45 | +2.360 |
| 3407 | Factorized | 500/536 | -6.1939 | -6.2303 | +0.0365 | +34.43 | +2.004 |
| 1926000 | Native | 500/581 | -6.2419 | -6.2003 | -0.0415 | +39.88 | +1.045 |
| 1926000 | Factorized | 500/500 | -6.2208 | -6.2003 | -0.0205 | +40.50 | +1.617 |
| Seed | Checkpoint | Valid / 256 | Valid (%) | Valid-only perf. | Improve valid (%) | Best-of-4 improve (%) |
|---|---|---|---|---|---|---|
| 3407 | Frozen | 7 | 2.73 | -8.9086 | 0.00 | 0.00 |
| 3407 | Native | 12 | 4.69 | -8.8106 | 8.33 | 1.56 |
| 3407 | Factorized | 33 | 12.89 | -9.5113 | 12.12 | 3.12 |
| 1926000 | Frozen | 7 | 2.73 | -8.9086 | 0.00 | 0.00 |
| 1926000 | Native | 8 | 3.12 | -8.8336 | 0.00 | 0.00 |
| 1926000 | Factorized | 3 | 1.17 | -8.6397 | 0.00 | 0.00 |
| Seed | Valid (%) | Valid-only | Improve (%) | Frontier (%) | Final best | AUC | Prompt tok. | Completion tok. |
|---|---|---|---|---|---|---|---|---|
| 42 | 82.05 | -9.0526 | 3.25 | 0.25 | -6.2202 | -6.2435 | 1,745,317 | 672,208 |
| 3407 | 82.15 | -9.8845 | 4.30 | 0.30 | -6.2254 | -6.2395 | 1,890,022 | 713,864 |
| 1926000 | 83.15 | -9.6897 | 4.65 | 0.75 | -6.2109 | -6.2518 | 1,842,358 | 711,707 |
| Mean | 82.45 | -9.5423 | 4.07 | 0.43 | -6.2188 | -6.2449 | 1,825,899 | 699,260 |