cs.SEOct 7, 2026

Cross-Provider Review as a Runtime Contract for Coding Agents: A Controlled Pilot and Fault-Injection Study

Authors: Bowen Xu, Boyu Chen

Organizations: Department of Biology, Stanford University, Stanford, CA, USA · Institute of Health Informatics, University College London, London, UK

Abstract

Coding agents increasingly share a workstation while drawing on separate providers and subscription allowances. A second agent can inspect a completed answer, but the call spends another pool and may provide no substantive finding. We describe an advisory cross-provider review contract: distinct resource pools, bounded execution, restricted reviewer capabilities, complete input delivery, usable semantic output, explicit failure states and durable per-attempt evidence. In a controlled, agent-authored pilot of 20 paired development turns, eight had a material reviewer finding (95% exact interval 19.1-63.9%). A boundary-condition scan across both reviewer backends reproduced a previously discovered false success on partial input: four truncation levels passed historically and failed after repair. The scan also found and repaired cancellation during process reaping. In real CLI probes, Claude had no writing tools; Codex attempted writes in five of five read-only trials, each write tool failed, and no disposable repository changed. These tests cover specified paths and versions, not field reliability. A preregistered shadow study of metadata-only review allocation accrued 25 formal observations before an exact-runtime regression found a third defect: a reviewer exiting nonzero with a well-formed verdict was counted as complete. Exit status was not recorded per attempt, so exposure cannot be resolved retrospectively. The 25 formal and two pending records remain an audit cohort; the measurement-valid cohort restarted at zero and collection has begun. No gate result is reported.

Figures & tables

Explore similar work

CardsList
  1. Groundability, Not Scale Alone: When Weak Reviewers Can Audit Strong Coding Agents

    Oct 1, 2026Junyu Guo, Shangding Gu, Ming Jin +1AI Agent AuditingCode Generation Evaluation

  2. Cross-Model LLM Code Review: Should you use Claude to review Codex or vice versa?

    Jul 22, 2026Zuodong Xiang, Yike Zhang, YueMing Zhang +1Code Generation EvaluationAutomated Code Review