Automatic Harness Evolution for Hardware Design Verification: Can LLMs Consolidate Gains Across Discovered Harnesses?
Organizations: NVIDIA
Abstract
Agent behavior depends on the harness surrounding a language model, but it remains unclear whether language models can reliably improve such harnesses for hardware-design tasks. We study automatic harness evolution around a fixed subject model on 12 proprietary design-verification root-cause localization tasks. Across five trials per task, automatically evolved harnesses increased completed attempts by 71-76% and any-hit task coverage by 80-100%, while total correct attempts improved by only 18-24%. The strongest success reproducible at least twice result improved by one task, and later candidates exchanged gains across tasks rather than preserving them. An auxiliary candidate improved on a four-task validation set excluded from search but tied its baseline on a subsequent 12-task replay containing both search and validation tasks, so the selected gain did not persist across the full pool. Across the tested lineage, useful search, evidence, and finalization behaviors appeared in different candidates but did not consistently consolidate into a single harness that dominated across tasks and metrics. In a separate CVDP cross-benchmark case study, an automatically evolved defined-width repair harness produced 35.6% more functional passes than its 142-task reference baseline; the final functional verifier scored completed outputs but was not shown to the subject agent during repair. These results support archive-aware selection when evolution yields complementary specializations without consistent consolidation.
Figures & tables
| Evaluation track | Task type | Scoring criterion | Role in the study |
|---|---|---|---|
| Proprietary design-verification benchmark | 12 root-cause localization tasks | Format-normalized function exact match | Primary within-domain diagnostic evaluation |
| Selected CVDP subset | 12 RTL generation and repair tasks | Executable functional verifier | Selected-subset phase of the cross-benchmark case study |
| 142-task CVDP follow-up | 142 RTL generation and repair tasks | Executable functional verifier | Baseline-to-evolved comparison in the cross-benchmark case study |
| Harness | First trial | Success 2/5 | Pass@5 | Correct attempts | Completed |
|---|---|---|---|---|---|
| Kimi K3 baseline | 2/12 | 5/12 | 5/12 | 17/60 | 34/60 |
| Automatically evolved evidence-grounded | 6/12 | 6/12 | 9/12 | 20/60 | 58/60 |
| Automatically evolved batched-causal | 5/12 | 5/12 | 10/12 | 21/60 | 60/60 |
| Stage | Principal addition | Intended effect |
|---|---|---|
| Evidence-first | Bounded orientation, literal search, anchored reads, forced finalization | Reduce speculative search and timeouts |
| Coverage ledger | Observed-text and inspected-file ledger with search/read budgets | Balance premature fixation and uncontrolled exploration |
| Evidence-grounded | Queries restricted to observed text; bounded sequential reads | Prevent invented queries and expensive paging |
| Causal commit | Multiple evidence types; source-to-propagation-to-detection ordering | Avoid naming the checker as the cause |
| Batched causal | Independent calls per turn plus standardized read windows | Gather more evidence within the turn budget |
| Domain-context treatment | Mean candidate pass rate | Best result | Interpretation |
|---|---|---|---|
| Static curated context | 16.7% | 4/12 | Matched standalone control |
| Live knowledge retrieval | 17.5% | 4/12 | Matched standalone control |
| Low-density injection | 10.4% | 3/12 | No reliable improvement |
| Specificity instruction | 10.3% | 3/12 | No demonstrated improvement |
| Heavy domain context | 10.6% | 2/12 | Content changed; score stayed flat |
| CVDP harness condition | Subject LLM | Functional passes |
|---|---|---|
| Reference baseline: verification loop | GLM-5.2 | 45/142 (31.7%) |
| Automatically evolved: defined-width repair | GLM-5.2 | 61/142 (43.0%) |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Design-verification task subfamily | Tasks | Typical failure or reasoning pattern |
|---|---|---|
| Local testbench or resource accounting | T1, T2, T3, T4, T5, T7 | Guards, counters, resource accounting, and tracing a local symptom to its producer |
| RTL or interface-state update | T8, T12 | Assignment semantics, state transitions, and calling-logic behavior |
| Simulation, launch, or cross-layer behavior | T6, T9, T10, T11 | Stimulus, checking, memory modes, restore behavior, and multi-layer failures |
| Treatment | Knowledge treatment | Mean candidate pass rate | Best result |
|---|---|---|---|
| Static curated DV context | Documentation in proposer context | 16.7% | 4/12 |
| Live DV retrieval | Proposer could retrieve specialized knowledge | 17.5% | 4/12 |
| Low-density injection | About 11 terms per 1,000 words | 10.4% | 3/12 |
| Specificity instruction | About 18 terms per 1,000 words | 10.3% | 3/12 |
| Heavy DV context | About 25 terms per 1,000 words | 10.6% | 2/12 |
| Configuration | Evaluated configurations | Mean pass rate | Best score | Input tokens | Agent time |
|---|---|---|---|---|---|
| Empty-policy baselines | 4 | 14.6% | 3/12 | 5.3–6.7M | about 1,250 s |
| Evolved domain scaffolds | 58 | n.r. | 3/12 | 7.4–8.2M | about 1,600 s |
| Model | Passed attempts / 36 | Timeouts |
|---|---|---|
| GLM 5.2 | 12/36 | 6 |
| Kimi K2.6 | 6/36 | 17 |
| Nemotron Ultra | 4/36 | 4 |
| MiniMax | 2/36 | 23 |