Some defects in an AI-generated study can be identified from its artifacts; others require knowledge of what was approved before execution. We propose study contracts that bind declared experimental choices, run obligations and claim scope to recorded execution evidence, and distinguish this contract-relative verification from scientific truth. A diagnostic using eight self-authored clean/mutated pairs illustrates the information boundary. A deterministic checker applying a registered, fault-specific rule to approved and executed objects detected all eight registered mutations. Across eighteen recorded judge aliases given individual metadata-filtered packages without pair context or the registry-selected fault label, 104 of 144 mutated evaluation cases received defect flags; the remaining cases comprised 32 abstentions and eight terminal failures, with no explicit clean decisions on mutated cases. Some packages retained approval and execution fields, including digests. The prompt instructed judges to abstain when evidence was insufficient. These results characterize a deliberately information-asymmetric development setting; they do not isolate the effect of authoritative information from differences in task specification and rule selection, and they are not comparative verifier quality or agent reward hacking. We identify full-information comparisons, legitimate-adaptation controls and closed-loop agent evaluations as necessary tests of whether contract checks improve useful compliant completion under optimization.
Figures & tables
Gold case class
Cases
Defect flag
Explicit clean
Abstention
Terminal failure
Registered mutation
144
104
0
32
8
Paired clean control
144
1
74
63
6
Total
288
105
74
95
14
Table 1: Judge outcomes over 288 evaluation cases (18 aliases × 16 packages). Deterministic checker: 8/8 mutations blocked, 0/8 false blocks. No mutated case received an explicit clean decision.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Fixture
Mutation type
Flagged by
compilation_drift_dev_001
approval/compiler digest drift
18/18
comparison_budget_001
mismatched seed/budget/stopping
17/18
external_validation_001
overbroad external generalization
17/18
reconstruction_dev_001
reconstruction manifest incomplete
17/18
runtime_contract_001
material workload mutation
16/18
integrity_digest_dev_001
artifact digest mismatch
16/18
Appendix
Table 2: Mutated packages flagged by how many of 18 aliases (development-corpus diagnostic only).
Alias
Mut. flags
Clean flags
Mut. non-dec.
Clean non-dec.
Decisions
/8
/8
/8
/8
/16
OpenAI gpt-5.6-sol
6
0
2
2
12
OpenAI gpt-5.6-terra
6
0
2
6
8
OpenAI gpt-5.6-luna
5
0
3
5
8
OpenAI gpt-5.5
6
0
2
2
12
OpenAI gpt-5.4
6
0
2
6
8
Appendix
Table 3: Per-alias counts. Non-decisions pool abstentions and terminal failures per alias; the pooled totals are 32+8 mutated and 63+6 clean. Aliases are not eighteen independent datasets nor necessarily immutable model snapshots.
Prop.
Mechanism
Status
Not evidence of
P1
Versioned study specification with per-field provenance ( explicit , inferred , missing )
that denominators improve conclusions; cells are fixtures
P3
Typed claim scope over machine-checkable dimensions with recorded human remainder
X4 ≡ Y4 at git 648338e : promotion blocked, seven failed dimensions, largest admissible scope returned, audit_finding_pass=0
passing overall finding; solved semantic containment; completed human review; any detector/clinical result
P3
Obligation-local and claim-level contradiction represented separately
Implemented
pilot-validated disposition behaviour
Appendix
Table 4: Engineering observations, not comparative results. X3 ≡ Y3 and X4 ≡ Y4 denote shared artifacts. Records report human approval; this paper does not independently authenticate every recorded actor.
AI research agents combine public information and experimental feedback to produce measurable results. The Discovery Certification Protocol (DCP) turns an outcome claim into an executable audit under a registered model, information boundary, and budget. Gate 1 validates useful improvement. Gate 2 tests recovery by matched agents given the starting information and observed Web content, with run history and new measurements withheld. Core requires adequate registered controls, zero recoveries, and a finite-sample recovery bound. Optional Gate 3 compares truthful and neutral feedback from a shared checkpoint; Evidence adds a supported effect and a null-policy equivalence check. Controlled SQLite and virtual catalyst audits pass both decision kernels. On real-data response surfaces, Yacht and Ionosphere pass the Core kernel after zero recoveries in 96 attempts, with an upper bound of 0.0468. Each target combines ten observed utilities and six predictions into a 16-entry data product. Yacht scores 0.7677 on reconstruction of all 32 switch effects, with utility-prediction MAE 0.0315 on its six unmeasured configurations. Fresh truthful continuations recover the target level in 9/30 and 16/30 trials, respectively, separating achieved utility from process repeatability. A deterministic verifier reproduces these local decisions from frozen records.
Jingjie Ning, Shanshan Zhong, Xiaochuan Li +1
School of Computer Science, Carnegie Mellon University
As LLM agents move from answering questions to carrying out procedures, failures can be unwarranted rather than visibly wrong: the final response looks acceptable even though the system skipped the check, branch, dependency, or invariant that made the answer justified. Output-only evaluation sees the answer, and trace-aware judging sees activity, but neither identifies which obligations were active for the query. We introduce CONTRACTEVAL, a diagnostic framework for making those active obligations explicit. It represents procedural instructions as query-active obligations and matches them against response or trace evidence, turning omissions, wrong branches, ordering errors, extra actions, invariant breaches, and output-contract violations into distinct conformance failures. On a controlled suite of audited procedural contracts, output-only and trace-aware LLM judges miss many injected structural failures; under gold expected and observed graphs, ContractEval detects and localizes all of them. LLM-backed extraction preserves much of this signal but remains calibration-sensitive. ContractEval is therefore not a compliance guarantee; it makes procedural conformance auditable rather than implicit in final-answer quality.
AI coding agents increasingly accept assigned software tasks, modify repositories under bounded authority, and return work packages for review. Prior work proposed the software delegation contract, covering the task, authority, returned work package, and acceptance context, as the unit of analysis for delegated coding work, but did not measure its effects. This paper reports a controlled pilot study of explicit delegation contracts for coding agents. We built a dependency-free TypeScript API task environment with seeded defects and documentation gaps, authored ten tasks across five families, and ran 64 agent executions across two model tiers under three conditions: a realistic issue-style prompt, an explicit delegation contract, and a contract with a required evidence bundle. Each run was scored with hidden acceptance tests, mutation checks, and scope analysis, then reviewed by three independent condition-blinded model-based reviewers using a fixed rubric, for 192 reviews. Explicit contracts did not improve objective task outcomes: all 64 runs passed hidden acceptance checks, with zero scope violations. They did improve reviewability. Evidence sufficiency improved in 22 of 30 paired comparisons and worsened in none (+0.83 on a 5-point scale, p < 0.0001, Cliff's delta = 0.66); reviewer ambiguity decreased (p = 0.035); changed-file lists, known-limitations sections, residual-risk sections, and reviewer checklists appeared mostly or only when demanded by the contract. Contracts cost +13% agent tokens and +38% wall-clock time, with larger effects for the weaker model tier. On these small tasks, delegation contracts bought reviewability rather than correctness.