Between the Commits: Process, Error, and Claim Reliability in a Wholly AI-Authored Codebase
Organizations: Trinity College Dublin Dublin, Ireland
Abstract
We present: (i) a new dataset consisting of the full development history of a 21,000-line Python tool built entirely by Claude AI, with no human-authored code or tests, (ii) two code-provenance tracing tools, (iii) three taxonomies for instruction intent, commit provenance, and response reliability, (iv) application of these to analyse the dataset. We find that: (i) user coding agent CLI instructions differ in kind from IDE-chat instructions, with a greater focus on comprehension, planning and consultation, (ii) code development is mainly proactive, (iii) 14.3% of AI code-generation events contain a real error later caught by the AI-authored test suite, (iv) roughly 1 in 4-5 of the AI's interactive responses contains one or more factual errors.
Figures & tables
| Size (src/test LOC) | Tests | Commits | Sessions | User instructions |
| 21,000 / 21,829 | 1,467 | 210 | 25 | 678 |
| Category | Examples |
|---|---|
| 1.1 New Implementation | “look into adding a check for use of google protobufs e.g. for GPBMessage, so we can give the tool user an informative message”; “then look into improving the test coverage so we can properly test later changes for regressions” |
| 1.2 Iterative Modification | “then swap into production”; “update tool to reflect this change” |
| 1.3 Alignment Correction | “the reference protos are not ground truth, they’re known to contain errors” |
| 2.2 Symptom Description | “fix heuristic_freq.py”; “with locationd the regex gate fails with garbage.” |
| 3.1 Planning & Consultation | “plan out what step 1 would involve”; “suggest develop a proof of concept first that can be tested” |
| 3.2 Project Comprehension | “no field types with swift? is that unavoidable or a tool limitation?”; “we already talked at length about enum-name-prefix and decided the enum type name was cosmetic, look back to see if you can find that conversation and confirm” |
| Category | Definition |
|---|---|
| Proactive-Spec | Content directly translated from a documented format spec/ABI (no investigation needed to write it). |
| Proactive-Analysis | Content found by directly inspecting a raw artifact, because no documented spec covers it (e.g. disassembling compiled code). |
| Proactive-Design | Content is a general software-design/architecture decision. |
| Proactive-Audit | Found by systematically re-reviewing already-working code or tests for quality/robustness, via one of two sub-routes: general review (a self-review pass identifies a list of follow-up items) or necessity verification (a candidate mechanism is disabled/relaxed and the output re-diffed; zero diff proves it dead/redundant before retirement). |
| Reactive-Spec | Found reactively, the fix was determined from a documented format spec/ABI. |
| Reactive-Analysis | Found reactively, the fix required directly inspecting the raw artifact under study to determine. |
| Category | Commits (Codex) | Commits (Claude) | Code blocks (Codex) | Code blocks (Claude) |
|---|---|---|---|---|
| Proactive-Audit | 71 (33.8%) | 76 (36.2%) | 437 (23.6%) | 673 (36.4%) |
| Proactive-Design | 35 (16.7%) | 31 (14.8%) | 513 (27.8%) | 476 (25.8%) |
| Proactive-Analysis | 27 (12.9%) | 31 (14.8%) | 706 (38.2%) | 496 (26.8%) |
| Proactive-Spec | 0 (0.0%) | 0 (0.0%) | 6 (0.3%) | 6 (0.3%) |
| Housekeeping | 56 (26.7%) | 56 (26.7%) | – | – |
| Reactive-Analysis | 10 (4.8%) | 7 (3.3%) | 111 (6.0%) | 117 (6.3%) |
| Category | Definition |
|---|---|
| Clean (no failure) | A pytest invocation reporting zero failures. The code under test worked as checked, whether on the first attempt or a later confirmation run. |
| Test/fixture defect | The check failed because the newly-written test itself (its assertion or its synthetic fixture data) was wrong. |
| Production-logic defect | The check failed because the reverse-engineering/decoding logic under test was genuinely wrong. |
| Failed to propagate change to existing tests | The check failed because a correct change made elsewhere in the same session (a new method, a changed diagnostic message, a build-script fix) had not been carried through to a dependent mock, assertion, or script. |
| Category | Runs | % |
|---|---|---|
| Clean (no failure) | 90 | 85.7% |
| Test/fixture defect | 8 | 7.6% |
| Production-logic defect | 4 | 3.8% |
| Failed to propagate change to existing tests | 3 | 2.9% |
| Total | 105 | 100% |
| Category | Definition | Verdict categories |
|---|---|---|
| Checkable-fact | A single fact about one artifact (a count, a presence/absence, a structural property) e.g. “Full suite green: 1174 passed, 0 failed.” | Matches / Contradicts / Unreproducible |
| Mechanism-explanation | Relates multiple facts or code locations into an explanation of how existing logic works (data flow, control flow, a multi-step process) e.g. “The revert at extractor.py:4025-4028 fires when a field was tagged ENUM.” | Confirmed / Wrong / Unreproducible |
| Root-cause-diagnosis | Diagnoses the cause of a stated or implied symptom/problem (something not working, unexpected, wrong, lost, or failing) e.g. “The root cause: my scan uses stop_at_ret=False with a 256-insn window, which runs past the wrapper’s end into neighboring functions, picking up other classes’ cores.” | Confirmed / Partial / Right-Conclusion-Wrong-Mechanism / Wrong / Unreproducible |
| Design-fix-proposal | Describes a proposed or implemented code change/fix and what happened to it e.g. “The fix: attribute by control-flow block instead.” | Adopted-Worked / Adopted-Revised / Adopted-Worked-But-Superseded / Not-Adopted / Failed / Mixed / Unreproducible |
| Other | A statement of intent, an interrogative, an opinion, a hedge, or pure process narration, e.g. “The exact cause hasn’t been pinned down yet.” | Not scored |
| Category | Scorer | Accurate | 95% CI |
|---|---|---|---|
| Checkable-fact | Claude | 2321/2461 (94.3%) | 93.3–95.2% |
| Codex | 2305/2443 (94.4%) | 93.4–95.2% | |
| Mechanism-explanation | Claude | 381/418 (91.1%) | 88.0–93.5% |
| Codex | 376/420 (89.5%) | 86.2–92.1% | |
| Root-cause-diagnosis | Claude | 275/308 (89.3%) | 85.3–92.3% |
| Codex | 275/304 (90.5%) | 86.6–93.3% |