Cross-Provider Review as a Runtime Contract for Coding Agents: A Controlled Pilot and Fault-Injection Study
Organizations: Department of Biology, Stanford University, Stanford, CA, USA · Institute of Health Informatics, University College London, London, UK
Abstract
Coding agents increasingly share a workstation while drawing on separate providers and subscription allowances. A second agent can inspect a completed answer, but the call spends another pool and may provide no substantive finding. We describe an advisory cross-provider review contract: distinct resource pools, bounded execution, restricted reviewer capabilities, complete input delivery, usable semantic output, explicit failure states and durable per-attempt evidence. In a controlled, agent-authored pilot of 20 paired development turns, eight had a material reviewer finding (95% exact interval 19.1-63.9%). A boundary-condition scan across both reviewer backends reproduced a previously discovered false success on partial input: four truncation levels passed historically and failed after repair. The scan also found and repaired cancellation during process reaping. In real CLI probes, Claude had no writing tools; Codex attempted writes in five of five read-only trials, each write tool failed, and no disposable repository changed. These tests cover specified paths and versions, not field reliability. A preregistered shadow study of metadata-only review allocation accrued 25 formal observations before an exact-runtime regression found a third defect: a reviewer exiting nonzero with a well-formed verdict was counted as complete. Exit status was not recorded per attempt, so exposure cannot be resolved retrospectively. The 25 formal and two pending records remain an audit cohort; the measurement-valid cohort restarted at zero and collection has begun. No gate result is reported.
Figures & tables
| Property | Failure exercised | Protected measurement |
|---|---|---|
| Distinct resource pool | Same-pool configuration refused before spawn | Independent reviewer-call cost |
| Bounded lifetime | Child does not drain input; cancellation during review | Attempt denominator and timeout/cancellation attrition |
| Restricted capabilities | Write-enabled launch arguments rejected; real CLI asked to write in disposable repos | Observed working-tree integrity under tested backends |
| Complete input delivery | Plausible clean verdict after four truncation levels, from 1% delivered to one character missing | Valid review denominator |
| Semantic result and typed failure | Empty, malformed, exit-zero empty output, and nonzero exit with well-formed verdict | No failure miscounted as "no material finding" |
| Durable per-attempt evidence | Exit status not recorded when defect found after collection | Retrospective scoping of affected observations |
| Direction | Pairs | Mat. | Nonmat. | P med. | R med. |
|---|---|---|---|---|---|
| Codex Claude | 10 | 2 | 8 | 77.7 | 117.5 |
| Claude Codex | 10 | 6 | 4 | 282.9 | 24.7 |
| Condition | Observed classification or filesystem outcome |
|---|---|
| Clean-looking verdict after 1%, 50%, 99%, or all but one character delivered | Both backends: historical completed 4/4; current input_delivery_failed 4/4 |
| Cancellation during input write, output read, or process reap | Both: pre-reap-repair accepted the reap case; current cancelled in all three phases |
| Nonzero exit with complete, well-formed verdict | Both: every collection-era revision completed; post-repair candidate process_crash ; all six exit/output scenarios classified as specified per backend (12/12) |
| Output at 0.8, 1.0, or 1.2 s with 1.0 s budget | Both: completed, timeout , timeout , respectively, in this timed scan |
| Malformed JSON; exit-zero empty result | Claude: malformed_output , empty_response ; Codex: empty_response for either; no clean verdict |
| Same-pool or write-enabled launch configuration | same-pool refusal or unsafe_argv , before launch |