BEHAVE: Functional Behavior Modeling Enables Self-Improving Agents for Hardware Design and Verification
Organizations: Stanford University · Workato, Inc.
Abstract
Developing agents for hardware design and verification requires reliable correctness feedback. As a hardware specification may permit correct implementations with different latencies, matching design and reference outputs cycle by cycle can reject valid designs. To address this, we introduce BEHAVE, an agentic framework for multi-turn joint hardware design and verification through functional behavior modeling. We define Behavior IR to express task functionality as executable behavior models without prescribing implementation timing beyond the specification. The agent iteratively develops a register-transfer-level (RTL) design and a behavior model as the design's verification reference. Our evaluator, BEHAVE-Sim, checks both artifacts separately against a hidden golden behavior model using input stimuli generated by random sampling and solver-guided search. BEHAVE thus supports power, performance, and area (PPA) exploration across task-permitted latencies and microarchitectures. During training, the same evaluator provides verifiable reinforcement learning (RL) rewards from specification-behavior pairs without reference RTL. For self-improvement, the agent continually searches for high-level implementations relevant to its capability gaps, constructs and checks specification-behavior pairs, and trains on the expanded task pool. We release BEHAVE-Train and BEHAVE-Eval with 600 human-reviewed specification-behavior pairs for realistic hardware workloads. Starting from 60 seed tasks and acquiring 100 new tasks, self-improvement raises Qwen3.8-27B's RTL pass@1 on BEHAVE-Eval from 55.0% to 75.0%, reaching performance comparable to RL using a 540-task pool.
Figures & tables
| VerilogEval-v2 | RTLLM-v2 | CVDP-cid003 | BEHAVE-Eval | |||||||||
| Model | RTL | Behav. | Cov. | RTL | Behav. | Cov. | RTL | Behav. | Cov. | RTL | Behav. | Cov. |
| Qwen3.5-4B | 34.6 | 28.8 | 6.9 | 12.0 | 20.0 | 0.0 | 4.3 | 4.3 | 0.9 | 0.0 | 6.7 | 0.0 |
| Qwen3.5-9B | 46.8 | 40.4 | 13.4 | 18.0 | 22.0 | 2.1 | 7.0 | 9.6 | 2.4 | 0.0 | 15.0 | 1.8 |
| Claude Haiku 4.5 | 92.9 | 89.1 | 62.7 | 88.0 | 84.0 | 87.2 | 72.2 | 54.8 | 83.8 | 36.7 | 70.0 | 48.1 |
| GPT-5.6 Luna | 94.2 | 94.9 | 84.1 | 98.0 | 98.0 | 85.3 | 93.0 | 93.9 | 91.6 | 68.3 | 73.3 | 89.7 |
| Claude Sonnet 5 | 92.3 | 92.3 | 58.1 | 84.0 | 86.0 | 75.7 | 64.3 | 60.9 | 66.9 | 68.3 | 65.0 | 60.7 |
| Inference | Training | ||||||
| Method | Task | Design | Verification | Alignment | Timing Flex. | Training Source | Self- Improv. |
| MAGE [ 20 ] | Design | RTL | Verilog TB G/P | Cycle | No | - | No |
| ChatModel [ 8 ] | Verification | - | SystemC RM G | Cycle | No | - | No |
| PRO-V-R1 [ 11 ] | Verification | - | Python RM G | Cycle | No | Spec-RTL | No |
| FAVer [ 9 ] | Design | RTL | Python RM G | Cycle | No | Spec-RTL | No |
| AutoVeriFix+ [ 10 ] | Both | RTL | Python RM G | Cycle | No | - | No |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Capability | Interaction requirements | Typical interfaces |
| Sampling and flow control | Sampling events, transfer detection and backpressure | Plain, valid-only and ready/valid interfaces |
| Framing and byte validity | Frame boundaries, valid bytes and sideband fields | AXI4-Stream [ 115 ] |
| Phased transfers | Setup, access, wait states and completion | APB [ 116 ] |
| Multi-channel assembly | Independently arriving address and data | AXI4-Lite [ 117 ] |
| Outstanding requests | Pending-request limits and response ordering | Memory request/response interfaces |
| Invocation and completion | Start, busy, completion and result consumption | Start/done interfaces |
| Feature | Evaluation and replay | Goal-guided generation | Exact bounded analysis |
| Behavior models | Behavior models use a restricted Python subset. I/O values may be Boolean, bounded integers, fixed-size arrays, records or enumerations. | The process() implementation must have statically proved finite bounds. | Every analyzed operation and loop must have a finite exact encoding. |
| RTL designs | Fixed-width SystemVerilog designs that pass structural checks execute in isolated Verilator. Native source-event observation does not require a statically proved loop trip count. | Each target goal requires a finite symbolic query relating selected input fields to whether that goal is reached. | A faithful encoding must cover the complete RTL transition relation within the analysis bound. |
| Protocols | Behavior uses logical transactions. RTL supports plain, valid-only, ready/valid and start/done interfaces, task-configured APB, AXI4-Lite and framed AXI4-Stream subsets, and bounded request/response. | Requires a finite symbolic query covering the selected protocol and transaction mapping. Generation varies logical input fields under fixed protocol rules, transaction order and environment schedule. | Requires an exact encoding of the selected protocol, transaction mapping and monitors. Schedules remain fixed except for single-clock ready/valid with one input and one output. |
| Endpoints and clock domains | Replay supports multiple endpoints and clock domains for Behavior and RTL. A task’s RTL interfaces may use different protocols. | Behavior generation supports multiple endpoints within one clock domain. RTL generation with multiple endpoints or clocks requires a fixed global schedule. | Behavior analysis supports multiple endpoints. Multi-clock Behavior and multiple-endpoint or multi-clock RTL require all analyzed stimuli to share one complete global clock schedule. |
| Reset | Replay follows the reset events in each stimulus. Reset during traffic requires the task to define how pending transactions are handled. | Goal-guided generation does not add or move reset events. | Reset timing remains fixed unless it is explicitly represented as a symbolic schedule choice. |
| State and internal memory | Replay supports persistent Behavior state, RTL registers and writable internal memories. | Behavior state must be finite and initialized. RTL goal-guided generation does not support writable internal memories. | All analyzed state must have a finite exact encoding. RTL exact analysis does not support writable internal memories. |
| Family | Goal schema | Eligible instances | Goal is hit when |
| Program structure | location: execution-site entry | each admitted executable source site of the Behavior or RTL artifact | execution enters any occurrence of that site |
| Program structure | control choice: binary branch or multiway selection outcome | the true and false branches of each admitted binary decision, and each alternative of an admitted multiway selection | execution takes that branch or selects that alternative |
| Program structure | expression: Boolean-expression evaluation pattern | each policy-generated partial Boolean assignment to the terms of an admitted expression | the expression is evaluated, and every term named by the pattern is evaluated with the specified value |
| Program structure | iteration: loop zero- or non-zero activation | admitted run-time loops, including procedural RTL loops but excluding generate loops | on a dynamic entry, control exits before body entry for the zero goal, or takes the first body-entry edge for the non-zero goal |
| Value/state | value observation: state/value bin | a sampling point and a total Boolean predicate over one or more values sampled there | the sampling point occurs and the predicate is true |
| Value/state | operation event: semantic predicate | an admitted operation site and a total predicate over its operands, result and semantic flags | the operation occurs and the predicate is true |
| Model | Model ID | Access | Reasoning |
| Qwen3.5-4B | Qwen/Qwen3.5-4B | SGLang | Thinking enabled |
| Qwen3.5-9B | Qwen/Qwen3.5-9B | SGLang | Thinking enabled |
| Qwen3.8-27B | Qwen/Qwen3.8-27B | SGLang | Extra-high |
| GPT-5.6 Luna | gpt-5.6-luna | OpenAI | Medium (default) |
| GPT-5.6 Terra | gpt-5.6-terra | OpenAI | Medium (default) |
| GPT-6 Astra | gpt-6-astra | OpenAI | Extra-high / High |
| Configuration | VerilogEval-v2 | RTLLM-v2 | CVDP-cid003 | BEHAVE (Train + Eval) | Avg. |
| Baseline | 91.75 | 96.61 | 87.83 | 94.76 | 93.48 |
| + Reference-directed | 92.23 | 97.61 | 88.88 | 95.24 | 94.06 |
| + Artifact-directed | 92.24 | 97.10 | 89.10 | 95.39 | 94.16 |
| + Two-sided | 92.24 | 97.43 | 88.39 | 95.59 | 94.21 |
| Setting | Value |
| Model | Qwen3.8-27B |
| Parameter updates and precision | Full-parameter, BF16 |
| GPUs | 8 NVIDIA H200 |
| Training parallelism (tensor/context/pipeline) | 4 / 2 / 1 |
| Rollout serving | 2 SGLang engines, tensor parallelism 4 |
| Maximum concurrent conversations | 32 |
| Experiment | Scope | Resource use / API cost |
| Fixed-dataset RL | 40 updates | 883.5 H200 GPU-hours |
| Self-improvement | 18 updates with task acquisition | 537.7 H200 GPU-hours |
| Seed-only | 15 updates after shared update 3 | 340.7 H200 GPU-hours |
| Evaluation | 5 API models 381 tasks | $959.12 |
| Pipeline comparison | 5 workflows 60 tasks | $10.54 |
| PPA exploration | 12 MXFP4 + 10 NVFP4 searches | $236.92 |
| Domain | Tasks | Representative families | Representative sources |
| AI/ML | 190 | Attention, normalization, quantization | Transformers, TorchAO |
| Numerics | 160 | Linear algebra, reductions, interpolation | CMSIS-DSP, LAPACK |
| Signal processing | 80 | Filtering, transforms, resampling, modulation | liquid-dsp, librosa, Opus |
| Cryptography | 80 | Symmetric primitives, field arithmetic, post-quantum kernels | PyCryptodome, liboqs, HEXL |
| Control | 30 | Coordinate transforms, controllers, estimation | SimpleFOC, FilterPy |
| Compression | 30 | Entropy coding, stream packing, match operations | FLAC, Brotli, Zstd |
| Feature | Representative behavior | Train | Eval | Total |
| Single-stream 1:1 input/output | Block quantization | 492 | 56 | 548 |
| Single-stream non-1:1 input/output | Filtering, accumulation, interpolation | 47 | 4 | 51 |
| Multiple independent streams | Three-input, two-output crossbar | 1 | 0 | 1 |
| External memory access | Memory-backed vectors, lookup tables | 26 | 2 | 28 |
| External memory writes | In-place updates, writeback | 8 | 1 | 9 |
| Collection | Tasks | Spec words | Behavior LOC | RTL LOC |
| VerilogEval-v2 | 156 | 270.5 [211.5, 327.5] | 9 [7, 13] | 11.5 [7, 21] |
| RTLLM-v2 | 50 | 312 [231.75, 427] | 14 [9, 24] | 20.5 [15, 33.75] |
| CVDP-cid003 | 115 | 473 [358.5, 615.5] | 28 [16.5, 47] | 34 [21, 53] |
| BEHAVE (ours) | 600 | 633 [496.75, 861.5] | 40 [24, 69] | 97 [56, 183] |
| Train | 540 | 623.5 [498.75, 869.75] | 39 [24, 67.25] | 97 [56, 177.25] |
| Eval | 60 | 645.5 [481.25, 808] | 43.5 [22, 70.25] | 102.5 [60.75, 209.5] |
| Review item | Researcher checks | Tasks revised, |
| Specification consistency | Requirements and examples agree, with valid inputs, parameter constraints and boundary behavior defined. | 6 |
| Numerical semantics | Widths, signedness, overflow, rounding, special values and operation order are explicit where relevant. | 2 |
| Interface and state | Packing, transaction counts and ordering, reset, backpressure and timing are defined where applicable. | 14 |
| Implementation fidelity | Behavior models and auxiliary RTL match the specified function, state transitions and memory effects. | 7 |
| Evidence integrity | Source and execution records match reviewed versions and runtime settings, including retests. | 1 |
| Total | 30 |
| Stage | Revised, (%) | 0 rounds | 1 round | rounds |
| comparison | 6 (0.65%) | 915 | 6 | 0 |
| Agent review | 14 (1.52%) | 907 | 14 | 0 |
| Human review | 30 (3.26%) | 891 | 29 | 1 |