cs.CLOct 4, 2026

Verification Trap: Understanding Test-Time Selection Failures under False Premises in Code Generation

Authors: Feng He, Hejia Wang, Linghao Meng, Ming Gao, Qiankun Li

Organizations: University of Science and Technology of China, China · Beijing University of Posts and Telecommunications, China · National University of Singapore, Singapore · Nanyang Technological University, Singapore

Abstract

Test-time compute has become a central way to improve code generation: systems sample multiple candidate programs and use verifier-visible evidence to select the final output. This paradigm implicitly assumes that the verifier provides a corrective signal independent from the generator. We challenge this assumption under misleading task premises. When the generator and verifier share a false premise, they become coupled through a mistaken belief: the generator produces premise-consistent shortcuts, while the verifier supplies evidence that fails to expose them. Consequently, the selector may choose a hidden-test-wrong candidate even when a hidden-test-correct program exists in the pool. We call this failure mode Verification Trap. Across three code-generation benchmarks and five code models, false premises consistently degrade first-sample correctness, reduce selector-chosen correctness after 64-sample test-time selection, and amplify recoverable mis-selection. Mechanistically, verifier-written tests inherit the premise-level blind spot, reshaping verifier-visible candidate space away from hidden-test correctness. These traces make Verification Trap predictable before hidden execution: a lightweight gold-free predictor using verifier-visible features reaches 0.846 AUROC. Our results identify decoupled evidence as a key mitigation axis: coupled scaling provides limited recovery, whereas premise-agnostic robustness auditors recover substantial oracle headroom.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Code Monitor Red Teaming for Public-Test-Passing Code

    Jul 23, 2026Junchi Liao, Jiawen Deng, Fuji RenCode QualityRed-Teaming

  2. ACES: Who Tests the Tests? Leave-One-Out AUC Consistency for Code Generation

    Apr 5, 2026Hui Sun, Yun-Ji Zhang, Zheng Xie +4Test GenerationCode Generation

  3. VeriContest: A Competitive-Programming Benchmark for Verifiable Code Generation

    May 8, 2026Zichen Xie, Mrigank Pawagi, Yuxin Liu +5Code GenerationRaw Judge Outputs