DoGBench: Can Agents Meet Expert Standards for User-Facing Documentation?
Organizations: PROMPTLESS · PROMPTLESS / DOC DETECTIVE
Abstract
We introduce DoGBENCH (Documentation Generation Benchmark), to our knowledge, the first benchmark for generating and maintaining real user-facing software documentation. It asks whether an agent can produce documentation that experienced technical writers would accept in review. The benchmark contains 292 items from open source projects, including Helm, PostHog, and Mautic. Each item gives the agent a pre-change repository and a trigger, such as a code pull request or a reported documentation gap. The agent must first decide whether the documentation needs an update. For items that need one, the agent must produce an acceptable patch in one attempt. For items that do not need updates, the agent must abstain. Task-specific rubrics, validated with project maintainers, score each patch on accuracy, completeness, reader guidance, placement, and repository conventions. The composite score combines patch quality with correct abstention, and a score of 100 means an agent meets every requirement for the task. Scores should not be interpreted as a percentage of an expert's capability. We evaluated seven agents. The highest-scoring agent reached 47.3 out of 100 on the 117-item held-out split. In a separate audit of 1,267 patches, the most common failure modes were task-completion gaps (45.5%), technical inaccuracies (36.6%), and incomplete conceptual or reference coverage (32.5%). Analysis of the corresponding trajectories identified three key patterns associated with these failures: (1) describing interfaces without examining how readers use them (36.0%), (2) missing decisive evidence and filling the gaps with plausible assumptions (33.1%), and (3) stopping after finding the first plausible documentation surface and leaving other affected pages stale (30.1%).
Figures & tables
| Agent | Score | Accuracy | Patch recall | Abstention recall | P0-clean delivery | Delivered quality | Conditional quality |
|---|---|---|---|---|---|---|---|
| Qwen3.8 Max +OpenCode | 47.3 | 79.5 | 80.5 | 77.1 | 26.8 | 34.1 | 42.3 |
| GPT-5.6 Sol +Codex | 46.2 | 81.2 | 95.1 | 48.6 | 39.0 | 44.1 | 46.3 |
| GLM 5.2 +OpenCode | 44.2 | 83.8 | 84.1 | 82.9 | 23.2 | 30.1 | 35.8 |
| Kimi K2.7 Code +OpenCode | 43.9 | 82.1 | 91.5 | 60.0 | 24.4 | 34.6 | 37.8 |
| Claude Opus 4.8 +Claude Code | 41.2 | 81.2 | 79.3 | 85.7 | 24.4 | 27.1 | 34.2 |
| Claude Sonnet 4.6 +Claude Code | 40.0 | 82.9 | 93.9 | 57.1 | 20.7 | 30.8 | 32.8 |
| Agent | Accuracy | Patch recall | Abstention recall |
|---|---|---|---|
| GLM 5.2+OpenCode | 83.8 (98/117) | 84.1 (69/82) | 82.9 (29/35) |
| Claude Sonnet 4.6+Claude Code | 82.9 (97/117) | 93.9 (77/82) | 57.1 (20/35) |
| Kimi K2.7 Code+OpenCode | 82.1 (96/117) | 91.5 (75/82) | 60.0 (21/35) |
| GPT-5.6 Sol+Codex | 81.2 (95/117) | 95.1 (78/82) | 48.6 (17/35) |
| Claude Opus 4.8+Claude Code | 81.2 (95/117) | 79.3 (65/82) | 85.7 (30/35) |
| Qwen3.8 Max+OpenCode | 79.5 (93/117) | 80.5 (66/82) | 77.1 (27/35) |
| Agent | Mean | 95% CI | Median | Correct patches ( ) |
|---|---|---|---|---|
| GPT-5.6 Sol+Codex | 46.3 | [40.8, 51.3] | 50.0 | 78 |
| GPT-5.5+Codex | 43.6 | [37.8, 48.7] | 42.9 | 79 |
| Qwen3.8 Max+OpenCode | 42.3 | [35.0, 48.7] | 42.8 | 66 |
| Kimi K2.7 Code+OpenCode | 37.8 | [31.9, 43.3] | 35.7 | 75 |
| GLM 5.2+OpenCode | 35.8 | [29.7, 41.8] | 35.7 | 69 |
| Claude Opus 4.8+Claude Code | 34.2 | [28.5, 40.1] | 36.4 | 65 |
| Agent | Pre | Pre score | Post | Post score | Repo 95% CI | |
|---|---|---|---|---|---|---|
| Claude Opus 4.8+Claude Code | 86 | 29.4 | 119 | 27.1 | [ , 4.8] | |
| Claude Sonnet 4.6+Claude Code | 50 | 34.6 | 155 | 30.9 | [ , 5.3] | |
| GPT-5.5+Codex | 69 | 42.9 | 136 | 40.9 | [ , 7.1] | |
| GPT-5.6 Sol+Codex | 94 | 44.3 | 111 | 44.0 | [ , 6.7] |
| Technical-writing failure mode | Effect on the documentation | Rate |
|---|---|---|
| Task-completion gap | Omits a decision point, procedure, verification step, or recovery path needed to complete the reader’s task. | 45.5% |
| Technical inaccuracy | Misstates an interface or behavior, including its scope, default, lifecycle, or compatibility. | 36.6% |
| Incomplete conceptual or reference coverage | Omits the core concept, capability, or contract the reader needs to understand and use the change. | 32.5% |
| Missing supporting information | Covers the main task but omits material rationale, boundaries, examples, operational detail, or secondary cases. | 26.5% |
| Missing prerequisites | Omits permissions, dependencies, credentials, versions, resources, enablement, or other setup conditions. | 21.0% |
| Information-architecture or findability failure | Places content outside the reader’s likely path or omits navigation, cross-references, and findable terminology. | 20.7% |
| Trajectory-level root cause | Observable reasoning failure | Submissions | Rate |
|---|---|---|---|
| Stopped at explaining the interface without examining how it is used in practice | The agent described changed fields, settings, callbacks, or lifecycle mechanics, but did not test the explanation against the reader’s setup, decision, execution, verification, or recovery path. | 456 | 36.0% |
| Did not inspect decisive evidence and filled the gap with a plausible assumption | The agent found related material but stopped before the controlling implementation, schema, test, or public contract, then completed the explanation with a convention that sounded reasonable. | 420 | 33.1% |
| Stopped searching after finding the first plausible documentation surface | The agent found a reasonable page to edit and did not continue checking other maintained, generated, mirrored, migration, or workflow surfaces affected by the same change. | 382 | 30.1% |
| Inspected relevant evidence but did not convert it into a complete coverage checklist | The agent reached evidence bearing on the requirement but began drafting without tracking the claims, setup, boundaries, examples, and reader actions that needed to survive into the final patch. | 349 | 27.5% |
| Committed too early to a narrow interpretation of the task | Before completing the investigation, the agent declared the task to be a rename, reference update, single-page edit, or similarly narrow deliverable and ignored evidence outside that frame. | 346 | 27.3% |
| Overgeneralized or misinterpreted partial evidence | The agent inspected relevant evidence but converted one branch, example, implementation detail, or deployment pattern into a broader or different public rule. | 333 | 26.3% |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting |
| Item: pantsbuild-pants-pr23219 . Agent: GLM 5.2+OpenCode. Source: Pants PR #23219 . This is a documentation-only item, so the agent receives no implementation code diff. An update is required. |
| Task request supplied to the agent |
| ⬇ Document combining Python coverage from sharded Pants test runs and applying the coverage threshold to the combined result . |
| Human merged patch (excerpts) |
| From docs/docs/python/goals/test.mdx . [...] marks omissions. |
| Incomplete coverage, raw output, and threshold placement |
| Setting |
| Item: jj-vcs-jj-pr9347 . Agent: GPT-5.6 Sol+Codex. Source: Jujutsu PR #9347 . In this code-triggered item, the change introduces a byte-string template type. The agent receives the code diff and the pre-change documentation and must update docs/templates.md . |
| Agent-visible code diff (excerpts) |
| Annotation-line content: cli/src/commit_templater.rs |
| ⬇ let out_property = self_property . map (| line | line . content ); - // TODO : Add Bytes or BString template type ? - Ok ( P :: wrap_template ( out_property . into_template ())) + Ok ( out_property . into_dyn_wrapped ()) |
| Fallible byte-to-string conversion: cli/src/template_builder.rs |
| ⬇ + let from_bytes = + | s : BString | Ok ( String :: from_utf8 ( s . into ()). map_err (| err | err . utf8_error ())?); + let property = match self . property . try_into_string () { + Ok ( string_property ) => return Some ( string_property ), + Err ( property ) => property , + }; + let property = match property . try_into_byte_string () { + Ok ( bytes_property ) => return Some ( bytes_property . and_then ( from_bytes ). into_dyn ()), + Err ( property ) => property , + }; |
| Subset | Acc. | Prec. | Recall | F1 | TP | FP | FN | TN | |
|---|---|---|---|---|---|---|---|---|---|
| Documentation-need gate | 42 | 0.976 | 0.958 | 1.000 | 0.979 | 23 | 1 | 0 | 18 |
| Human-rubric recall | Count | Percentage |
|---|---|---|
| Covered | 220 | 90.9% |
| Missing | 17 | 7.0% |
| Conflicting | 5 | 2.1% |
| Total | 242 | 100.0% |
| Generated-rubrics precision | Count | Percentage |
|---|---|---|
| Covered | 216 | 59.8% |
| Missing | 137 | 38.0% |
| Conflicting | 8 | 2.2% |
| Total | 361 | 100.0% |
| Grader | Coverage | Matches | Agreement |
|---|---|---|---|
| GPT-5.6 Terra | 20/20 | 310/330 | 93.9% |
| Claude Sonnet 5 | 20/20 | 301/330 | 91.2% |
| Case / criterion / score | Faulty human-patch excerpt | Why the failure remains P0 |
|---|---|---|
| Pants #22034 C8; 50.0 | pants experimental-deploy src/k8s/:webpages | The exact-version address parser rejects the empty path component before the colon, so the deployment command cannot resolve its target. Removing the slash fixes it: src/k8s:webpages . The defect is small but blocks the advertised operation. |
| OpenCost #102 C6; 44.4 | Secret creation: kubectl create secret generic azure-service-key -n kubecost (excerpt). Workload update: helm upgrade opencost . --namespace opencost -f values.yaml | The Secret is created in kubecost , but the workload that mounts it is updated in opencost . A workload cannot mount a Secret from another namespace. The supplied credential-injection procedure therefore fails unless the namespaces are made consistent. |
| Strawberry #4342 C6; 31.2 | Resolver: def create_user(self, email: str) -> str: xx return email Schema: strawberry.Schema( mutation=Mutation, extensions=[ PydanticErrorExtension() ],) | The usage example omits the required query root and never invokes Pydantic validation. It cannot produce the advertised validation errors. The same PR’s working test supplies both a query root and Pydantic model construction, which gives a direct implementation contrast. |
| dlt #2292 C8; 55.6 | iceberg_tables[ "my_iceberg_table"] .optimize.compact() | The helper returns native PyIceberg Table objects. The retained source check for supported version 0.8.1 finds no optimize API or dynamic fallback. The copied Delta-style operation cannot run on that object, so a supported Iceberg operation must replace it. |
| Agent | Score | Accuracy | Patch recall | Abstention recall | P0-clean delivery | Delivered quality | Conditional quality |
|---|---|---|---|---|---|---|---|
| Qwen3.8 Max+OpenCode | 48.1 | 81.2 | 84.4 | 73.6 | 30.7 | 35.8 | 42.4 |
| GLM 5.2+OpenCode | 46.9 | 84.6 | 84.9 | 83.9 | 26.3 | 32.6 | 38.4 |
| GPT-5.6 Sol+Codex | 45.0 | 81.2 | 96.1 | 46.0 | 38.5 | 44.1 | 45.9 |
| Claude Opus 4.8+Claude Code | 42.2 | 80.8 | 79.0 | 85.1 | 21.0 | 28.1 | 35.5 |
| Kimi K2.7 Code+OpenCode | 42.0 | 79.1 | 88.8 | 56.3 | 25.9 | 33.5 | 37.7 |
| Claude Sonnet 4.6+Claude Code | 39.1 | 81.5 | 94.6 | 50.6 | 25.9 | 31.8 | 33.6 |
| Agent | FN | FP |
|---|---|---|
| GLM 5.2+OpenCode | 31 | 14 |
| Claude Sonnet 4.6+Claude Code | 11 | 43 |
| Qwen3.8 Max+OpenCode | 32 | 23 |
| GPT-5.6 Sol+Codex | 8 | 47 |
| Claude Opus 4.8+Claude Code | 43 | 13 |
| Kimi K2.7 Code+OpenCode | 23 | 38 |
| Agent | Low-footprint | Popular | low popular | Repo 95% CI |
|---|---|---|---|---|
| Claude Opus 4.8+Claude Code | 26.3 | 28.5 | [ , 5.9] | |
| Claude Sonnet 4.6+Claude Code | 31.7 | 29.9 | [ , 7.2] | |
| GLM 5.2+OpenCode | 33.6 | 29.9 | [ , 12.9] | |
| GPT-5.5+Codex | 45.7 | 36.0 | [ , 16.1] | |
| GPT-5.6 Sol+Codex | 47.2 | 42.4 | [ , 10.9] | |
| Kimi K2.7 Code+OpenCode | 34.2 | 32.5 | [ , 9.3] |
| Technical-writing failure mode | Effect on the documentation | Rate |
|---|---|---|
| Low information density or poor scannability | Repeats information, adds unnecessary structure, or uses disproportionately long prose for the information conveyed. | 7.6% |
| Fabricated content | Invents an interface, command, control, behavior, version requirement, or guarantee; distinct from a false description of a real interface. | 6.1% |
| Nonfunctional example | Supplies a code block, command, configuration, or API example that would fail or teach the wrong call shape. | 2.5% |
| Documentation-system defect | Breaks links, markup, rendering, terminology, or documentation-system conventions. | 0.9% |
| Ambiguous or contradictory guidance | Gives incompatible instructions or leaves a material rule ambiguous within the edited documentation. | 0.4% |
| Scope creep | Edits unrelated files or topics beyond the documentation need; excludes companion edits required for consistency. | 0.2% |
| Trajectory-level root cause | Observable reasoning failure | Submissions | Rate |
|---|---|---|---|
| Found the correct fact, then dropped or contradicted it during drafting | The correct distinction appeared in the evidence or reasoning but disappeared, weakened, or reversed in the patch. | 81 | 6.4% |
| Selected the wrong or conflicting source as authoritative | The agent trusted stale documentation, generated output, an internal representation, or permissive runtime behavior over the maintained public contract. | 72 | 5.7% |
| Validated presentation or file mechanics, but not the substantive claim | The agent checked syntax, links, formatting, or file existence without validating the underlying command, example, route, or factual statement. | 50 | 3.9% |
| Did not perform a reader-priority and compression pass | The agent stopped after inserting relevant content without removing repetition, artificial structure, or low-density prose. | 22 | 1.7% |
| Did not run documentation-specific validation | The agent omitted the relevant link, markup, navigation, generation, spelling, or build check. | 9 | 0.7% |
| Lost requirements or edit state during a long trajectory | Earlier requirements vanished after extended investigation, rewriting, or scope expansion. | 7 | 0.6% |