Some defects in an AI-generated study can be identified from its artifacts; others require knowledge of what was approved before execution. We propose study contracts that bind declared experimental choices, run obligations and claim scope to recorded execution evidence, and distinguish this contract-relative verification from scientific truth. A diagnostic using eight self-authored clean/mutated pairs illustrates the information boundary. A deterministic checker applying a registered, fault-specific rule to approved and executed objects detected all eight registered mutations. Across eighteen recorded judge aliases given individual metadata-filtered packages without pair context or the registry-selected fault label, 104 of 144 mutated evaluation cases received defect flags; the remaining cases comprised 32 abstentions and eight terminal failures, with no explicit clean decisions on mutated cases. Some packages retained approval and execution fields, including digests. The prompt instructed judges to abstain when evidence was insufficient. These results characterize a deliberately information-asymmetric development setting; they do not isolate the effect of authoritative information from differences in task specification and rule selection, and they are not comparative verifier quality or agent reward hacking. We identify full-information comparisons, legitimate-adaptation controls and closed-loop agent evaluations as necessary tests of whether contract checks improve useful compliant completion under optimization.
Figures & tables
Gold case class
Cases
Defect flag
Explicit clean
Abstention
Terminal failure
Registered mutation
144
104
0
32
8
Paired clean control
144
1
74
63
6
Total
288
105
74
95
14
Table 1: Judge outcomes over 288 evaluation cases (18 aliases × 16 packages). Deterministic checker: 8/8 mutations blocked, 0/8 false blocks. No mutated case received an explicit clean decision.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Fixture
Mutation type
Flagged by
compilation_drift_dev_001
approval/compiler digest drift
18/18
comparison_budget_001
mismatched seed/budget/stopping
17/18
external_validation_001
overbroad external generalization
17/18
reconstruction_dev_001
reconstruction manifest incomplete
17/18
runtime_contract_001
material workload mutation
16/18
integrity_digest_dev_001
artifact digest mismatch
16/18
Appendix
Table 2: Mutated packages flagged by how many of 18 aliases (development-corpus diagnostic only).
Alias
Mut. flags
Clean flags
Mut. non-dec.
Clean non-dec.
Decisions
/8
/8
/8
/8
/16
OpenAI gpt-5.6-sol
6
0
2
2
12
OpenAI gpt-5.6-terra
6
0
2
6
8
OpenAI gpt-5.6-luna
5
0
3
5
8
OpenAI gpt-5.5
6
0
2
2
12
OpenAI gpt-5.4
6
0
2
6
8
Appendix
Table 3: Per-alias counts. Non-decisions pool abstentions and terminal failures per alias; the pooled totals are 32+8 mutated and 63+6 clean. Aliases are not eighteen independent datasets nor necessarily immutable model snapshots.
Prop.
Mechanism
Status
Not evidence of
P1
Versioned study specification with per-field provenance ( explicit , inferred , missing )
that denominators improve conclusions; cells are fixtures
P3
Typed claim scope over machine-checkable dimensions with recorded human remainder
X4 ≡ Y4 at git 648338e : promotion blocked, seven failed dimensions, largest admissible scope returned, audit_finding_pass=0
passing overall finding; solved semantic containment; completed human review; any detector/clinical result
P3
Obligation-local and claim-level contradiction represented separately
Implemented
pilot-validated disposition behaviour
Appendix
Table 4: Engineering observations, not comparative results. X3 ≡ Y3 and X4 ≡ Y4 denote shared artifacts. Records report human approval; this paper does not independently authenticate every recorded actor.