Long-Horizon Analog Design Bench: Benchmarking Agents on Hours-Long Analog and Mixed-Signal Circuit Design Tasks
Abstract
Coding agents now sustain hours-long, tool-driven loops, yet their ability to carry long-horizon analog and mixed-signal circuits to electrical specification remains unmeasured. We introduce Analog Design Bench, a long-horizon agentic benchmark of 50 transistor-level design tasks contributed by 17 chip designers. Agents work with an open-source simulator, while an isolated verifier evaluates the submitted circuit using specification-based electrical tests. We evaluate 15 agent configurations across 2,250 two-hour attempts and observe full-specification pass rates from 8.0% to 78.0%. Coding-benchmark performance correlates with analog results but leaves much of the performance spread unexplained. Our failure analysis shows that most unsuccessful submissions have no recorded legality rejection but fail electrical acceptance, identifying electrical closure as the dominant endpoint challenge. We test time, reasoning effort, agent harness, and supplied design knowledge as interventions. Longer budgets and higher reasoning effort improve performance, while general skill documents provide little benefit and sometimes reduce performance. Supplying a task-matched reference topology, an idealized form of circuit-IP retrieval, raises DeepSeek V4 Pro by 18.7 percentage points and mainly accelerates GPT-5.6 Sol.
Figures & tables
| Work | Task | Scoring | # Tasks | Original tasks | Long horizon * | Tool-use autonomy | Isolated verification |
|---|---|---|---|---|---|---|---|
| General Coding | |||||||
| Terminal-Bench ( Merrill et al., 2026 ) | Terminal tasks | AT | 89 | ||||
| DeepSWE ( Huang et al., 2026 ) | Repository changes | AT | 113 | ||||
| Digital Circuit | |||||||
| VerilogEval ( Liu et al., 2023b ) | RTL design | AT | 156 | ||||
| RTLLM ( Lu et al., 2024 ) | RTL design | AT | 50 | ||||
| pass@1 | R1 / R2 / R3 | pass@3 | SpecScore | Cost | Out tok | Turns | |
|---|---|---|---|---|---|---|---|
| Configuration | (%) | (%) | (%) | (%) | (USD) | (k) | |
| MiMo 2.5 Pro [thinking] | 8.00 | 6 / 12 / 6 | 16.00 | 22.45 | 2.66 | 332 | 78 |
| GLM-5.2 [max] | 8.00 | 6 / 8 / 10 | 18.00 | 22.87 | 1.57 | 101 | 39 |
| Doubao Seed 2.1 Pro [high] | 14.00 | 16 / 16 / 10 | 26.00 | 28.99 | 3.34 | 63 | 56 |
| DeepSeek V4 Flash [max] | 16.00 | 16 / 20 / 12 | 24.00 | 30.68 | 0.12 | 191 | 84 |
| GLM-5.3 Flash [max] | 18.67 | 20 / 24 / 12 | 38.00 | 33.09 | 0.49 | 382 | 70 |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Slot | Family | Solved by | Checks | Design objective |
|---|---|---|---|---|
| 1 | Power management and references | 15 | 30 | Implement a half-bridge class-D power amplifier targeting > 95% peak efficiency and > 30 mW output power. |
| 2 | Power management and references | 15 | 360 | Implement a programmable ICC/IPTAT NMOS current mirror targeting 1.00 mA nominal output current and 8% output-current error. |
| 3 | Signal-chain amplifiers and active filters | 12 | 75 | Implement a source-degenerated GM-R differential amplifier targeting 1.98–2.02 V/V gain and > 60 MHz bandwidth. |
| 4 | Signal-chain amplifiers and active filters | 14 | 25 | Implement a source-degenerated GM-R differential amplifier targeting 1.98–2.02 V/V gain and > 70 MHz bandwidth. |
| 5 | Interfaces and drivers | 5 | 243 | Implement an all-NMOS half-bridge bootstrap gate driver with a designable off-chip bootstrap capacitor, targeting 3–7 ns dead time and < 3.5 mW driver power. |
| 6 | Signal-chain amplifiers and active filters | 9 | 13 | Implement a three-stage switched-capacitor ring amplifier targeting 7.2–8.8 V/V closed-loop gain and 50 ns settling time. |
| Functional family | Tasks | Full pass (%) | SpecScore (%) |
|---|---|---|---|
| Data conversion and sampling | 9 | 30.4 | 45.5 |
| General-purpose op amps and OTAs | 9 | 28.9 | 41.1 |
| Interfaces and drivers | 3 | 33.3 | 50.4 |
| Power management and references | 11 | 50.7 | 67.9 |
| RF, timing, and high-speed | 9 | 48.9 | 62.2 |
| Signal-chain amplifiers and active filters | 9 | 31.6 | 43.8 |
| Harness | Version |
|---|---|
| Codex | 0.144.1 |
| Claude Code | 2.1.220 |
| Kimi Code | 0.32.0 |
| MCode | 0.2.20 |
| OpenCode | 1.18.18 |
| Pi | 0.80.7 |
| Skill | Package | Size |
| Held-out guidance : distilled by GPT-5.6 Sol [max] from the 1,260 main-experiment trajectories of tasks 1 to 30; evaluated on the held-out tasks 31 to 50 | ||
| 1 | General handbook | 1,295 words |
| 2 | Compact workflow | 668 words |
| 3 | Failure-diagnosis decision tree | 1,258 words |
| Full-suite design documents : distilled by GPT-5.6 Sol [max] from every successful and failed main-experiment trajectory then available, versions 2 and 3 revising version 1; evaluated on all 50 tasks | ||
| 4 | Design document, version 1 | 5,716 words |