Agents increasingly build on code written by other agents, and they reimplement rather than reuse, growing the codebases later agents must work in. To measure how well agents design libraries for other agents, we introduce LibraryDesignBench, a two-phase benchmark in which an agent implements a full-featured library from a specification that defines required capabilities and potential use cases without prescribing the design. We evaluate the library through the correctness and simplicity of programs written by three user agents from different model families. The benchmark spans 242 expert-validated programming problems across 15 library-design tasks in four languages. On eleven of the fifteen tasks, agent designers reproduce the abstractions of the human-written production library. Downstream agents adopt agent- and human-written libraries alike but underuse them, reimplementing capabilities the library already provides. Our failure analysis finds that downstream agents write extra code mainly because agent-written libraries are rigid or hard to use, not because capabilities are missing. We also experiment with giving designers more prescriptive, agent-first guidance and having them test their library with subagents; this improves downstream scores and yields simpler programs. LibraryDesignBench provides both a testbed for evaluating library-design practices for agent users and an initial design baseline that improves downstream reuse.
Figures & tables
Figure 1: LibraryDesignBench’s two-phase setup evaluates the library through real usage. The agent under evaluation designs the library from a non-prescriptive specification. Three downstream agents then solve problems with it. The score reflects how correct and simple their programs are.
Model (Harness)
Score ( ↑ )
% Pass ( ↑ )
Simplicity ( ↑ )
Library (\downarrow$ )
Problem (\downarrow$ )
DeepSeek V4 Pro
31.2 [30.4, 32.1]
84.9 ± 1.7
42.0 ± 1.1
0.31\pm$ 0.02
0.175\pm$ 0.008
Fable 5.1
47.5 [46.4, 48.5]
86.1 ± 1.6
62.7 ± 1.8
14.81\pm$ 1.37
0.198\pm$ 0.013
Fable 5.1 (CC)
39.9 [37.4, 42.4]
84.9 ± 1.6
58.3 ± 2.5
14.39\pm$ 2.23
0.211\pm$ 0.014
GLM 5.3
41.8 [40.8, 42.9]
84.1 ± 1.7
57.1 ± 1.7
21.06\pm$ 1.89
0.234\pm$ 0.013
GPT-5.6 Sol (Codex)
39.5 [38.4, 40.5]
84.1 ± 2.1
52.9 ± 1.2
2.14\pm$ 0.20
0.199\pm$ 0.011
GPT-6 Astra
45.1 [44.3, 45.8]
85.7 ± 1.7
58.7 ± 1.5
3.63\pm$ 0.24
0.155\pm$ 0.008
Table 1: Overall results for each library setup by designer. All results average the same implementer set. Designers use mini-SWE-agent ( Yang et al., 2025 ) unless another harness is named in parentheses, where CC is Claude Code. Production and No library are settings in which the implementer is given a human-written library or no library at all ( ), respectively. “Library ”istheaveragecost,inUSD,togenerateasinglelibrary,while“Problem” is the average implementer cost per evaluation problem. “% Pass” is the mean share of tests passed, not of fully solved problems. Scores show 95% CIs. Other values show clustered standard errors ( Section 2.3 ). Bold marks the best per column.
Figure 2: Agent-written libraries outperform production libraries on a subset of tasks. Difference in score per task compared with that of the production library. The rightmost column is the no-library setup. Outlined cells score below no library.
Setup
Score (95% CI, ↑ )
% Pass ( ↑ )
Simplicity ( ↑ )
Problem tokens ( ↓ )
Implementer: DeepSeek V4.1 Flash
Astra
46.9 [45.8, 47.9]
87.4 ± 1.7
58.5 ± 1.4
5.80 ± 0.34M
GPT-6 Sol
41.3 [40.2, 42.5]
87.0 ± 1.8
51.6 ± 1.2
4.86 ± 0.28M
Fable
50.2 [48.9, 51.5]
87.8 ± 1.7
63.2 ± 1.8
6.35 ± 0.35M
Opus 5.5
51.6 [50.4, 52.8]
88.4 ± 1.5
64.3 ± 1.8
6.47 ± 0.36M
GLM 5.3
45.1 [43.8, 46.4]
86.2 ± 1.8
57.3 ± 1.6
8.43 ± 0.46M
Table 2: Results per implementer across library setups. Only mini-SWE-agent designers are shown. “Problem tokens” is the mean number of tokens per problem.
Figure 3: Agents converge on the same designs. README quick-starts from Astra (top) and Fable (bottom) across six clirs libraries each. Teal : all six; orange : Astra only; violet : Fable only.
Figure 4: Primary failure category per designer. Failed Tests classifies why at least one test failed. Excess Code classifies why the solution was longer than the reference. Definitions are in Table 8 .
Figure 7
Figure 6: Explicit guidance results . Left: score gain per language, split into simplicity and correctness contributions ( Appendix I ). Right: clirs examples under each prompt.
Pinned requirements.txt in /workspace/.venv ; uv cache warmed
TypeScript (1)
tsc , tsx , prettier , Chromium
Pinned package.json
Rust (5)
cargo , rustfmt , clippy ; Rust 1.85–1.91
Prefetched Cargo.lock ; CARGO_NET_OFFLINE
Haskell (3)
GHC 9.8.4, cabal , fourmolu
Prefetched frozen Cabal closure
Appendix
Table 4: Per-task Docker images, shared by both phases and all library conditions. Counts are tasks.
Design Phase (design)
Evaluation Phase (one problem)
Agent wall-clock
4 hours
60 minutes
Verifier wall-clock
15 minutes
5–60 minutes, set per problem
Sandbox
4 CPU, 8 GB
2 CPU, 4 GB
Spend cap
None
$2.50
Attempts
1 ( K=3 independent runs per task)
1
Agent network
Model-provider API allowlist only
Model-provider API allowlist only
Appendix
Table 5: Harness-enforced limits per phase. The agent budget is wall-clock inside the container. Design runs against wall-clock alone; a problem ends at whichever of its two budgets binds first.
Language
Formatter (pinned)
Config
tree-sitter grammar
Python
ruff format 0.16.6
ruff.toml
python
TypeScript
prettier 3.9.6
prettier.json
typescript
Rust
rustfmt 1.9.0-stable
rustfmt.toml
rust
Haskell
fourmolu 0.20.1.0
fourmolu.yaml
haskell
Appendix
Table 6: Pinned normalizers and grammars used for static measurement. The same versions are applied to the optimized references and to every generated program. Measurement uses its own pinned Rust 1.98.1 toolchain for rustfmt , separate from the per-task build toolchains in Table 4 ; TypeScript problems additionally load the JavaScript grammar for embedded sources.
Library
Language
Production library
Domain
Problems
canon
Python
pydantic
Schema validation
14
pyda
Python
pandas
Dataframes
21
roadkill
Python
uxsim
Traffic simulation
16
sapi
Python
fastapi
HTTP API framework
11
simu
Python
simpy
Discrete-event simulation
13
uglypie
Python
beautifulsoup4
HTML parsing
11
Appendix
Table 7: The fifteen LibraryDesignBench library-design problems.
Category
Leaf
Definition
Limited by the library: no path through the library as shipped does better.
Coverage
Absent operation
No operation or documented composition performs this reusable domain computation, and none does nearly this. The fix is a new operation.
Correctness
Contract violation
The library returns wrong output on a legitimate input.
Misleading diagnostic
An error pointed away from the actual cause, and the implementer followed it.
Performance defect
The library path is too slow for the problem’s limits.
Rigidity
Fixed policy
An operation hard-codes how it works, such as ordering, rounding, error handling, or output format, with no parameter that selects what the problem needs.
Appendix
Table 8: Failure-taxonomy categories and leaves; Appendix F gives the classification procedure.
Failed tests
Excess code
Category
Leaf
Astra
GPT-6 Sol
Fable
Opus 5.5
GLM 5.3
Grok
All
Astra
GPT-6 Sol
Fable
Opus 5.5
GLM 5.3
Grok
All
Limited by the library
Coverage
Absent operation
20
32
21
23
19
16
131
28
19
21
17
16
15
116
Correctness
Contract violation
9
17
22
17
25
33
123
1
2
3
2
9
9
26
Misleading diagnostic
0
0
0
0
0
0
0
0
0
0
0
0
0
0
Performance defect
5
0
1
0
1
1
8
0
0
0
0
1
0
1
Appendix
Table 9: Failure-taxonomy leaf counts per designer, over the standardized subset of Failed tests and Excess code cells classified in Section 3.2 . Each designer contributes 135 cells per symptom stratum. Leaf definitions are in Table 8 .
Minimal
Low
Medium
High (default)
Never touch library (%)
23.7
0.0
0.0
0.1
Docs/examples read before first write (%)
13
88
87
100
Library grep before first write (%)
44
88
98
100
Distinct library files read
3.5
7.5
10.0
14.8
Distinct grep terms
12
20
33
46
Grep share of library commands
.32
.27
.33
.29
Appendix
Table 10: GPT-5.6 Luna library interaction by prescription level (high effort).
Low
Medium
High
Distinct library files read
2.0
5.7
14.8
Library commands per trial
3.1
5.3
12.2
Any library grep (%)
45
60
91
Any library source read (%)
23
50
81
Only listings or docs (%)
42
18
0.4
Return to library after first write (%)
22
30
47
Appendix
Table 11: GPT-5.6 Luna library interaction by reasoning effort (default prompt).
Today's agents are highly effective at implementing well-scoped software design plans, but user intent is often vague and admits multiple equally valid solutions. In this paper, we introduce SpecBench, a new benchmark for evaluating an agent's ability to translate user intent into a structured, executable specification that aligns with user preferences. The agent is given access to past user conversations and may interact with the user for a fixed number of rounds to ask clarifying questions. We find that existing agents exhibit two extreme behaviors: they either (i) struggle to collaborate proactively with users, entering implementation mode too quickly while overestimating their understanding of user preferences, or (ii) exhaust their question budget by asking about every ambiguous design choice. To address this limitation, we introduce a user-assistant agent: Buddy. It follows a workflow inspired by classical morphological analysis, decomposing user intent into a structured space of design dimensions and candidate choices. It then creates simulated users to evaluate these choices, before engaging the real user to resolve remaining ambiguities and finalize the specification. By shifting the focus from execution to specification, SpecBench and Buddy emphasize agent-user collaboration (not just code generation) as a key frontier in future agent design.
Hao Wang, Ligong Han, Kai Xu +1
Red Hat AI Innovation · MIT-IBM Watson AI Lab · IBM Core AI
We present AgentLens, a production-assessed benchmark for interactive code agents. Most code-agent benchmarks reduce a run to a single bit -- did the task pass? -- but the people who actually use these agents experience the entire trajectory: how the agent follows instructions, uses its tools, verifies its own work, recovers from mistakes, and talks to them along the way. AgentLens evaluates that whole trajectory. It pairs formal verification, where an objective check exists, with LLM-written trajectory reviews and side-by-side comparisons, so that each run yields a readable explanation of why the score is what it is. This makes AgentLens useful for more than ranking models: we use it to diagnose model behavior, compare successive versions of our own agent, and catch product regressions in a nightly evaluation pipeline. We release the benchmark as open source at https://github.com/agent-lens/agent-lens-bench.
Benchmark scores tell you what an agent got right; they do not tell you how it got there. In this work, we introduce methods for comparing agents procedurally in different contexts, where the model, tasks, and approaches vary. We compare ten agents and find that they are identifiable by their behavioral habits, which we define as fingerprints: a probe over these procedural signatures attributes an unseen trajectory to the correct agent at 85.7% accuracy, controlling for leakage across tasks. We develop procedural representations for agent problem-solving procedures with an emergent vocabulary induction technique that is meant to be maximally compressive to avoid surface-level variation while being expressive enough to unveil the quirks of the models' patterns. We apply our framework to the software engineering evaluation dataset SWE-Bench to study the structural distinctness of agent trajectories and find that behavior is most similar between models from similar release periods and those that are distilled from one another (e.g., a distilled student model and its teacher have a Jensen-Shannon divergence of 0.25, about half the distance between other model pairs). As more models saturate evaluations, we believe that it will be important to probe model behavior along more holistic dimensions than success rates alone. We introduce ProcGrep, a library for auditing and evaluating agents for how they approach tasks at a procedural level given their traces in a top-down fashion. We believe this work has a range of applications to help developers work with and program coding agents, such as task-aware model routing, agent monitoring, and finer-grained cost analysis.