Coding agents can return plausible patches that omit required behavior. These failures are hard to review because long traces and confident summaries often hide what was missed. We ask when a nominally weaker reviewer can reliably decide whether a patch solves its issue. We study 411 execution-labeled traces from three agents and 101 controlled cases. On 154 GPT-5.4 traces, structured but unchecked evidence raises both defect catch and over-rejection. We then provide official execution evidence as an upper-bound diagnostic. After choosing and freezing one of two formats per reviewer, five of six reviewers improve both rates on 122 held-out traces; two classify every trace correctly. Reviewer size is not a consistent predictor of quality. Because official tests are unavailable in deployment, we also evaluate a frozen cascade with patch-caused static errors and generated tests that first fail on the unpatched repository. On 121 scored held-out GPT-5.4 traces and 59 Gemini traces, its coverage is 0.89 and 0.86, risk is 0.33 and 0.26, catch is 0.76 and 0.80, and over-rejection is 0.66 and 0.67. Most false rejections occur when unresolved cases reach the reviewer. Official execution evidence shows the potential of weak review when decisive checks are available. Producing equally reliable checks without official tests remains the main bottleneck.
Figures & tables
Figure 1: Overview. A plausible patch can hide an omission (A). Official execution evidence shows that a fixed weak reviewer can separate defects from false alarms in an upper-bound diagnostic (B), motivating independent checks before review (C). The starred pass branch is the frozen evaluation rule, not a recommended acceptance rule.
Set
n
Composition
Real, execution-grounded
GPT-5.4 core
154
53 acc. / 76 omis. / 9 regr. / 16 no change
design (Django)
32
21 acc. / 9 omis. / 1 regr. / 1 no change
Gemini transfer
59
15 acc. / 33 omis. / 1 regr. / 10 no change
Real, execution-grounded (accept/no-change)
GPT-5.4 multilingual
78
38 acc. / 40 no change
Table 1: Benchmark composition. The core and transfer sets include real defects. The accept/no-change sets contain no defective submitted patches.
Figure 2: Held-out upper-bound results. Each arrow moves a reviewer from structured evidence to the official-evidence format chosen on the design set. The dashed arrow marks the one exception.
structured
official evidence
Reviewer (size)
risk notes
catch
over-rej.
catch
over-rej.
Llama-3.1 (8B)
removed
0.70
0.75
0.98
0.00
Qwen-2.5 (32B)
kept
0.69
0.66
0.96
0.09
Qwen-2.5 (72B)
kept
0.43
0.44
0.91
0.47
GPT-OSS (120B)
removed
0.51
0.41
1.00
0.00
Qwen3 (235B)
removed
0.28
0.19
0.90
0.00
Table 2: Held-out upper-bound comparison on 122 traces (90 defects, 32 accepts). The official-evidence format is selected on the 32-trace design set and then frozen. “Removed” omits unchecked risk notes; “kept” retains them.
Figure 3: Risk–coverage on held-out GPT-5.4 (top) and Gemini (bottom). Stars mark the cascade; triangles and circles show patch-only and structured review; squares show rules alone. All methods use the same traces. A full-set generated-test-only baseline is unavailable.
Set
cov.
risk
catch
over-rej.
held-out (121)
0.89
0.33
0.76
0.66
[0.84,0.94]
[0.25,0.43]
[0.67,0.85]
[0.49,0.81]
Gemini (59)
0.86
0.26
0.80
0.67
Table 3: Run-once cascade results. Catch and over-rejection include abstentions in their class denominators. Brackets are 95% bootstrap confidence intervals for held-out; Gemini values are point estimates.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure S4: The grounded review cascade on one patch; pill counts are the run-once held-out evaluation ( 121 traces).
Reviewer
Structured
+ missing item
gain
GPT-4.1
0.53
1.00
+0.47
Qwen-2.5-32B
0.40
1.00
+0.60
GPT-4-Turbo
0.13
1.00
+0.87
Llama-3.1-8B
0.07
1.00
+0.93
Appendix
Table S4: Catch on 15 synthetic partial-logic omissions. Naming the missing item helps every reviewer, with the largest gain for Llama-8B.
Figure S5: Descriptive full-set rendering comparison on all 154 core traces. Unlike the held-out main result, this plot chooses the better official-evidence format on the full set. (a) Target-omission catch from structured evidence (open) to official evidence (filled); triangles show Gemini transfer. (b) Crosses mark the other official-evidence format.
Figure S6: Decision composition on the pooled GPT-5.4 multilingual and Claude-Sonnet accept/no-change sets. Grounded evidence makes all three reviewers decisive and nearly always right.
Figure S7: Additional diagnostics. Top: correct-decision rate on the 154-trace core set. A decision is correct when the majority label matches execution ground truth; abstentions count as not correct. (a) Parameter count is not monotonic under structured review. (b) Fresh majority voting over K=1,…,11 for three reviewers; shading gives 95% bootstrap intervals. Averaged over the tested cells, moving from five to eleven votes changes the rate by one to three points, although individual cells vary more. (c) The evidence intervention produces the largest aggregate shift. Panel (b) uses a separate reduced-output run; frozen results elsewhere are unchanged. Bottom: Stages 0–2 do not call the weak reviewer; Stage 3 sends unresolved cases to Llama-8B. Test-generation cost is not shown.