cs.SEOct 1, 2026

Groundability, Not Scale Alone: When Weak Reviewers Can Audit Strong Coding Agents

Authors: Junyu Guo, Shangding Gu, Ming Jin, Javad Lavaei

Organizations: University of California, Berkeley · Virginia Tech

Abstract

Coding agents can return plausible patches that omit required behavior. These failures are hard to review because long traces and confident summaries often hide what was missed. We ask when a nominally weaker reviewer can reliably decide whether a patch solves its issue. We study 411 execution-labeled traces from three agents and 101 controlled cases. On 154 GPT-5.4 traces, structured but unchecked evidence raises both defect catch and over-rejection. We then provide official execution evidence as an upper-bound diagnostic. After choosing and freezing one of two formats per reviewer, five of six reviewers improve both rates on 122 held-out traces; two classify every trace correctly. Reviewer size is not a consistent predictor of quality. Because official tests are unavailable in deployment, we also evaluate a frozen cascade with patch-caused static errors and generated tests that first fail on the unpatched repository. On 121 scored held-out GPT-5.4 traces and 59 Gemini traces, its coverage is 0.89 and 0.86, risk is 0.33 and 0.26, catch is 0.76 and 0.80, and over-rejection is 0.66 and 0.67. Most false rejections occur when unresolved cases reach the reviewer. Official execution evidence shows the potential of weak review when decisive checks are available. Producing equally reliable checks without official tests remains the main bottleneck.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ExecCritic: Learn to Test, Test to Improve for Coding Agents

    Sep 8, 2026Leitian Tao, Baolin Peng, Haorui Wang +7Coding AgentsTest Generation

  2. Coding Agents as Test-Suite Auditors: Finding What Official Suites Miss While Approaching What They Catch

    Aug 3, 2026Shuyang Xie, Shuxiao Xie, Feng Zhu +2Raw Judge OutputsTest Generation

  3. 3100 Opinions on Code Review in an AI World: Building Causal Theory from Practitioner Discourse

    Jul 8, 2026Shyam Agarwal, Courtney Miller, Christian Kästner +1Pull RequestsCode Quality