Agents increasingly build on code written by other agents, and they reimplement rather than reuse, growing the codebases later agents must work in. To measure how well agents design libraries for other agents, we introduce LibraryDesignBench, a two-phase benchmark in which an agent implements a full-featured library from a specification that defines required capabilities and potential use cases without prescribing the design. We evaluate the library through the correctness and simplicity of programs written by three user agents from different model families. The benchmark spans 242 expert-validated programming problems across 15 library-design tasks in four languages. On eleven of the fifteen tasks, agent designers reproduce the abstractions of the human-written production library. Downstream agents adopt agent- and human-written libraries alike but underuse them, reimplementing capabilities the library already provides. Our failure analysis finds that downstream agents write extra code mainly because agent-written libraries are rigid or hard to use, not because capabilities are missing. We also experiment with giving designers more prescriptive, agent-first guidance and having them test their library with subagents; this improves downstream scores and yields simpler programs. LibraryDesignBench provides both a testbed for evaluating library-design practices for agent users and an initial design baseline that improves downstream reuse.
Figures & tables
Figure 1: LibraryDesignBench’s two-phase setup evaluates the library through real usage. The agent under evaluation designs the library from a non-prescriptive specification. Three downstream agents then solve problems with it. The score reflects how correct and simple their programs are.
Model (Harness)
Score ( ↑ )
% Pass ( ↑ )
Simplicity ( ↑ )
Library (\downarrow$ )
Problem (\downarrow$ )
DeepSeek V4 Pro
31.2 [30.4, 32.1]
84.9 ± 1.7
42.0 ± 1.1
0.31\pm$ 0.02
0.175\pm$ 0.008
Fable 5.1
47.5 [46.4, 48.5]
86.1 ± 1.6
62.7 ± 1.8
14.81\pm$ 1.37
0.198\pm$ 0.013
Fable 5.1 (CC)
39.9 [37.4, 42.4]
84.9 ± 1.6
58.3 ± 2.5
14.39\pm$ 2.23
0.211\pm$ 0.014
GLM 5.3
41.8 [40.8, 42.9]
84.1 ± 1.7
57.1 ± 1.7
21.06\pm$ 1.89
0.234\pm$ 0.013
GPT-5.6 Sol (Codex)
39.5 [38.4, 40.5]
84.1 ± 2.1
52.9 ± 1.2
2.14\pm$ 0.20
0.199\pm$ 0.011
GPT-6 Astra
45.1 [44.3, 45.8]
85.7 ± 1.7
58.7 ± 1.5
3.63\pm$ 0.24
0.155\pm$ 0.008
Table 1: Overall results for each library setup by designer. All results average the same implementer set. Designers use mini-SWE-agent ( Yang et al., 2025 ) unless another harness is named in parentheses, where CC is Claude Code. Production and No library are settings in which the implementer is given a human-written library or no library at all ( ), respectively. “Library ”istheaveragecost,inUSD,togenerateasinglelibrary,while“Problem” is the average implementer cost per evaluation problem. “% Pass” is the mean share of tests passed, not of fully solved problems. Scores show 95% CIs. Other values show clustered standard errors ( Section 2.3 ). Bold marks the best per column.
Figure 2: Agent-written libraries outperform production libraries on a subset of tasks. Difference in score per task compared with that of the production library. The rightmost column is the no-library setup. Outlined cells score below no library.
Setup
Score (95% CI, ↑ )
% Pass ( ↑ )
Simplicity ( ↑ )
Problem tokens ( ↓ )
Implementer: DeepSeek V4.1 Flash
Astra
46.9 [45.8, 47.9]
87.4 ± 1.7
58.5 ± 1.4
5.80 ± 0.34M
GPT-6 Sol
41.3 [40.2, 42.5]
87.0 ± 1.8
51.6 ± 1.2
4.86 ± 0.28M
Fable
50.2 [48.9, 51.5]
87.8 ± 1.7
63.2 ± 1.8
6.35 ± 0.35M
Opus 5.5
51.6 [50.4, 52.8]
88.4 ± 1.5
64.3 ± 1.8
6.47 ± 0.36M
GLM 5.3
45.1 [43.8, 46.4]
86.2 ± 1.8
57.3 ± 1.6
8.43 ± 0.46M
Table 2: Results per implementer across library setups. Only mini-SWE-agent designers are shown. “Problem tokens” is the mean number of tokens per problem.
Figure 3: Agents converge on the same designs. README quick-starts from Astra (top) and Fable (bottom) across six clirs libraries each. Teal : all six; orange : Astra only; violet : Fable only.
Figure 4: Primary failure category per designer. Failed Tests classifies why at least one test failed. Excess Code classifies why the solution was longer than the reference. Definitions are in Table 8 .
Figure 7
Figure 6: Explicit guidance results . Left: score gain per language, split into simplicity and correctness contributions ( Appendix I ). Right: clirs examples under each prompt.
Pinned requirements.txt in /workspace/.venv ; uv cache warmed
TypeScript (1)
tsc , tsx , prettier , Chromium
Pinned package.json
Rust (5)
cargo , rustfmt , clippy ; Rust 1.85–1.91
Prefetched Cargo.lock ; CARGO_NET_OFFLINE
Haskell (3)
GHC 9.8.4, cabal , fourmolu
Prefetched frozen Cabal closure
Appendix
Table 4: Per-task Docker images, shared by both phases and all library conditions. Counts are tasks.
Design Phase (design)
Evaluation Phase (one problem)
Agent wall-clock
4 hours
60 minutes
Verifier wall-clock
15 minutes
5–60 minutes, set per problem
Sandbox
4 CPU, 8 GB
2 CPU, 4 GB
Spend cap
None
$2.50
Attempts
1 ( K=3 independent runs per task)
1
Agent network
Model-provider API allowlist only
Model-provider API allowlist only
Appendix
Table 5: Harness-enforced limits per phase. The agent budget is wall-clock inside the container. Design runs against wall-clock alone; a problem ends at whichever of its two budgets binds first.
Language
Formatter (pinned)
Config
tree-sitter grammar
Python
ruff format 0.16.6
ruff.toml
python
TypeScript
prettier 3.9.6
prettier.json
typescript
Rust
rustfmt 1.9.0-stable
rustfmt.toml
rust
Haskell
fourmolu 0.20.1.0
fourmolu.yaml
haskell
Appendix
Table 6: Pinned normalizers and grammars used for static measurement. The same versions are applied to the optimized references and to every generated program. Measurement uses its own pinned Rust 1.98.1 toolchain for rustfmt , separate from the per-task build toolchains in Table 4 ; TypeScript problems additionally load the JavaScript grammar for embedded sources.
Library
Language
Production library
Domain
Problems
canon
Python
pydantic
Schema validation
14
pyda
Python
pandas
Dataframes
21
roadkill
Python
uxsim
Traffic simulation
16
sapi
Python
fastapi
HTTP API framework
11
simu
Python
simpy
Discrete-event simulation
13
uglypie
Python
beautifulsoup4
HTML parsing
11
Appendix
Table 7: The fifteen LibraryDesignBench library-design problems.
Category
Leaf
Definition
Limited by the library: no path through the library as shipped does better.
Coverage
Absent operation
No operation or documented composition performs this reusable domain computation, and none does nearly this. The fix is a new operation.
Correctness
Contract violation
The library returns wrong output on a legitimate input.
Misleading diagnostic
An error pointed away from the actual cause, and the implementer followed it.
Performance defect
The library path is too slow for the problem’s limits.
Rigidity
Fixed policy
An operation hard-codes how it works, such as ordering, rounding, error handling, or output format, with no parameter that selects what the problem needs.
Appendix
Table 8: Failure-taxonomy categories and leaves; Appendix F gives the classification procedure.
Failed tests
Excess code
Category
Leaf
Astra
GPT-6 Sol
Fable
Opus 5.5
GLM 5.3
Grok
All
Astra
GPT-6 Sol
Fable
Opus 5.5
GLM 5.3
Grok
All
Limited by the library
Coverage
Absent operation
20
32
21
23
19
16
131
28
19
21
17
16
15
116
Correctness
Contract violation
9
17
22
17
25
33
123
1
2
3
2
9
9
26
Misleading diagnostic
0
0
0
0
0
0
0
0
0
0
0
0
0
0
Performance defect
5
0
1
0
1
1
8
0
0
0
0
1
0
1
Appendix
Table 9: Failure-taxonomy leaf counts per designer, over the standardized subset of Failed tests and Excess code cells classified in Section 3.2 . Each designer contributes 135 cells per symptom stratum. Leaf definitions are in Table 8 .
Minimal
Low
Medium
High (default)
Never touch library (%)
23.7
0.0
0.0
0.1
Docs/examples read before first write (%)
13
88
87
100
Library grep before first write (%)
44
88
98
100
Distinct library files read
3.5
7.5
10.0
14.8
Distinct grep terms
12
20
33
46
Grep share of library commands
.32
.27
.33
.29
Appendix
Table 10: GPT-5.6 Luna library interaction by prescription level (high effort).
Low
Medium
High
Distinct library files read
2.0
5.7
14.8
Library commands per trial
3.1
5.3
12.2
Any library grep (%)
45
60
91
Any library source read (%)
23
50
81
Only listings or docs (%)
42
18
0.4
Return to library after first write (%)
22
30
47
Appendix
Table 11: GPT-5.6 Luna library interaction by reasoning effort (default prompt).