SpecRead: A Benchmark for Measuring Whether Language Models Understand Hardware Specifications
Organizations: Independent Researcher
Abstract
Existing benchmarks for large language models (LLMs) in hardware design evaluate downstream artifacts such as generated RTL, assertions, or testbenches. When a model fails such a benchmark, the failure is ambiguous: it may have misread the specification, or it may have understood the specification and failed to write the code. We present SpecRead, a benchmark that isolates specification comprehension from generation ability. SpecRead v2.1 contains 385 questions over 10 open-source OpenTitan IP blocks: exact retrieval, cross-section reasoning, contradiction detection in mutated specifications, and spec-RTL consistency checking, plus 82 controls (41 distractor, 41 consistent-RTL). Type-4 items are built from real RTL mutations; we retain only mutations that Icarus Verilog simulation shows to change observable behavior. A with-spec vs. without-spec ablation suggests the questions require the excerpt, not training recall alone (without-spec accuracy 3/20 on the t1/t2 subset), though memorization of the source text may still help spot mutations. As an initial characterization with a small model, Ministral-3B scores 33.2% overall (128/385; macro average 39.0%): 55.2% on retrieval, 51.7% on cross-section reasoning. On the two contradiction-focused types, the verdict-plus-location measure gives 48.0% (t3) and 63.3% (t4), with a 51.2% false-positive rate on distractors and 100% on consistent-RTL controls. Layered scoring shows the model locates contradictions well (78.9-81.6% location accuracy) but scores lower on their category (43.9-49.7%). A structured "rule-table" prompting intervention lowers accuracy on every question type except t2 (tied). SpecRead is automatically scorable by deterministic checks, with gray-zone cases counted wrong under the conservative main scoring. The benchmark is regenerable for type-3 items via mutation injection, and built exclusively from public sources.
Figures & tables
| IP block | questions |
|---|---|
| i2c | 55 |
| spi_host | 55 |
| aes | 50 |
| spi_device | 50 |
| kmac | 45 |
| hmac | 35 |
| Type | verdict+loc. correct | verdict+loc. acc. | full-credit acc. | 95% CI (v+l) | |
| t1 (exact retrieval) | 58 | 32 | 55.2% | 55.2% | [0.425, 0.673] |
| t2 (cross-section reasoning) | 58 | 30 | 51.7% | 51.7% | [0.392, 0.641] |
| t3 (spec contradiction) | 171 | 82/171 | 48.0% | 24.6% | [0.406, 0.554] |
| t4 (spec–RTL consistency) | 96 | 63/96 | 65.6% | 26.0% | [0.557, 0.744] |
| Overall | 383 | — | — | 33.7% (129/383) | [0.291, 0.386] |
| Macro avg. (mean of 4 types) a | 39.4% |
| Type | with spec | without spec | |
|---|---|---|---|
| t1 ( ) | 6/10 = 60.0% | 0/10 = 0.0% | +60.0 pp |
| t2 ( ) | 4/10 = 40.0% | 3/10 = 30.0% | +10.0 pp |
| t3 ( ) | 9/20 = 45.0% | 0/20 = 0.0% | +45.0 pp |
| t4 ( ) | 3/20 = 15.0% | 0/20 = 0.0% | +15.0 pp |
| Overall | 22/60 = 36.7% | 3/60 = 5.0% | +31.7 pp |
| gold pred | CROSSREF | MISSING | NUMERIC | RULE_INV |
|---|---|---|---|---|
| CROSSREF_CONFLICT | 24 | 2 | 9 | 4 |
| MISSING_CONDITION | 18 | 2 | 3 | 14 |
| NUMERIC_MISMATCH | 8 | 0 | 39 | 3 |
| RULE_INVERSION | 19 | 2 | 4 | 20 |
| gold pred | CROSSREF | MISSING | NUMERIC | RULE_INV |
|---|---|---|---|---|
| CROSSREF_CONFLICT | 1 | 0 | 3 | 1 |
| MISSING_CONDITION | 2 | 9 | 0 | 5 |
| NUMERIC_MISMATCH | 4 | 7 | 12 | 26 |
| RULE_INVERSION | 2 | 2 | 0 | 21 |
| Operator | correct | acc. | |
|---|---|---|---|
| changed_number | 21 | 53 | 39.6 |
| crossref_conflict | 9 | 30 | 30.0 |
| rule_inversion | 12 | 51 | 23.5 |
| deleted_condition | 0 | 37 | 0.0 |
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
| Type | correct | acc. (lenient) | 95% CI | |
| t1 (exact retrieval) | 58 | 32 | 55.2% | [0.425, 0.673] |
| t2 (cross-section reasoning) | 58 | 30 | 51.7% | [0.392, 0.641] |
| t3 (spec contradiction) | 171 | 69/171 | 40.4% | [0.333, 0.478] |
| t4 (spec–RTL consistency) | 96 | 35/96 | 36.5% | [0.275, 0.464] |
| Overall | 383 | 166/383 | 43.3% | [0.385, 0.483] |
| Macro avg. (mean of 4 types) | 45.9% |