cs.SESep 24, 2026

Automatic Harness Evolution for Hardware Design Verification: Can LLMs Consolidate Gains Across Discovered Harnesses?

Authors: Kidus Seyoum, Ajay Mittur

Organizations: NVIDIA

Abstract

Agent behavior depends on the harness surrounding a language model, but it remains unclear whether language models can reliably improve such harnesses for hardware-design tasks. We study automatic harness evolution around a fixed subject model on 12 proprietary design-verification root-cause localization tasks. Across five trials per task, automatically evolved harnesses increased completed attempts by 71-76% and any-hit task coverage by 80-100%, while total correct attempts improved by only 18-24%. The strongest success reproducible at least twice result improved by one task, and later candidates exchanged gains across tasks rather than preserving them. An auxiliary candidate improved on a four-task validation set excluded from search but tied its baseline on a subsequent 12-task replay containing both search and validation tasks, so the selected gain did not persist across the full pool. Across the tested lineage, useful search, evidence, and finalization behaviors appeared in different candidates but did not consistently consolidate into a single harness that dominated across tasks and metrics. In a separate CVDP cross-benchmark case study, an automatically evolved defined-width repair harness produced 35.6% more functional passes than its 142-task reference baseline; the final functional verifier scored completed outputs but was not shown to the subject agent during repair. These results support archive-aware selection when evolution yields complementary specializations without consistent consolidation.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. HarnessCompass: Guiding Automatic Harness Evolution toward Generalizable and Effective Agent Harnesses

    Aug 3, 2026Luan Zhang, Ruochen Zhou, Dandan Song +9Evolving HarnessAgent Harness

  2. Rethinking the Evaluation of Harness Evolution for Agents

    Jul 14, 2026Yike Wang, Huaisheng Zhu, Zhengyu Hu +7Large Language Model AgentsTest-Time Scaling

  3. HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?

    Sep 1, 2026Yuhao Wu, Jingyuan Zhang, Jiajun Shi +16Agent Harness