Verification Trap: Understanding Test-Time Selection Failures under False Premises in Code Generation
Authors: Feng He, Hejia Wang, Linghao Meng, Ming Gao, Qiankun Li
Organizations: University of Science and Technology of China, China · Beijing University of Posts and Telecommunications, China · National University of Singapore, Singapore · Nanyang Technological University, Singapore
Test-time compute has become a central way to improve code generation: systems sample multiple candidate programs and use verifier-visible evidence to select the final output. This paradigm implicitly assumes that the verifier provides a corrective signal independent from the generator. We challenge this assumption under misleading task premises. When the generator and verifier share a false premise, they become coupled through a mistaken belief: the generator produces premise-consistent shortcuts, while the verifier supplies evidence that fails to expose them. Consequently, the selector may choose a hidden-test-wrong candidate even when a hidden-test-correct program exists in the pool. We call this failure mode Verification Trap. Across three code-generation benchmarks and five code models, false premises consistently degrade first-sample correctness, reduce selector-chosen correctness after 64-sample test-time selection, and amplify recoverable mis-selection. Mechanistically, verifier-written tests inherit the premise-level blind spot, reshaping verifier-visible candidate space away from hidden-test correctness. These traces make Verification Trap predictable before hidden execution: a lightweight gold-free predictor using verifier-visible features reaches 0.846 AUROC. Our results identify decoupled evidence as a key mitigation axis: coupled scaling provides limited recovery, whereas premise-agnostic robustness auditors recover substantial oracle headroom.
Figures & tables
Figure 1: Overview of Verification Trap under false premises. (a) A false premise can make the selector choose a verifier-compatible shortcut despite an available correct candidate. (b) Generator-verifier coupling biases verifier-visible evidence toward premise-consistent shortcuts. (c) Decoupled evidence changes the final evidence channel and recovers candidates missed by coupled scaling.
Dataset
Model
Pass@1
Selected@64
Oracle@64
Trap@64
Base
False
Δ
Base
False
Δ
Base
False
Δ
Base
False
Δ
HumanEval+
Qwen-7B
77.95
70.38
-7.57
82.07
77.03
-5.04
94.48
89.86
-4.62
12.41
12.84
+0.43
Qwen-14B
84.78
68.81
-15.97
85.23
75.17
-10.06
93.29
84.56
-8.73
8.05
9.40
+1.35
Qwen-32B
83.65
69.44
-14.21
85.81
72.00
-13.81
90.54
84.00
-6.54
4.73
12.00
+7.27
DS-Coder-6.7B
64.52
55.43
-9.09
72.86
63.89
-8.97
94.29
89.58
-4.71
21.43
25.69
+4.26
GPT-4o-mini
83.22
70.38
-12.84
86.67
73.33
-13.34
95.33
86.67
-8.66
8.67
13.33
+4.66
Table 1: Verification Trap under false premises across datasets and models. We compare the baseline condition and the false-premise condition. Δ values are computed as false premise minus baseline in percentage points. Selected@64 and Trap@64 are computed using generated-test selection unless otherwise specified.
Figure 2: False premises reduce verifier counterexample coverage. Across model scales, the false-premise condition produces substantially fewer counterexample tests than both the baseline and irrelevant-control.
Figure 4
Feature Group
# Features
AUROC
Contamination
3
0.539
Score distribution
5
0.754
Signature structure
5
0.773
All features, LR
13
0.846
All features, RF
13
0.833
Table 2: Gold-free trap-risk prediction from verifier-visible features. All inputs exclude hidden-test-derived information.
Figure 5: Gold-free trap-risk prediction before hidden-test execution. Ranking tasks by predicted trap risk concentrates recoverable mis-selection into a small high-risk subset, showing that Verification Trap leaves detectable verifier-visible traces before gold execution.
Group
Method
Pass
Δ Pass
Original
Baseline
77.03
0.00
Coupled
Extra verifier tests
78.00
+0.97
Width scaling
76.67
-0.36
Robust regeneration
77.70
+0.67
Decoupled
Source-aware prior
81.08
+4.05
Robustness auditor
81.76
+4.73
Table 3: Representative mitigation results. Δ Pass is computed relative to the original verifier-selected baseline. Additional robustness checks across model and dataset settings are reported in § S5 .
Figure 6: Compute-performance tradeoff of mitigation methods. We compare coupled and decoupled mitigation on the Qwen-7B / HumanEval+ main cell.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Effect
Direction
Pooled Δ
Pooled p
Sign consistency
Cell-level sig.
Δ Selected@64
11/11 lower
−8.43
<0.0001
11/11, p=0.0010
9/11
Δ Trap@64
11/11 higher
+4.31
<0.0001
11/11, p=0.0010
4/11
Appendix
Table S1: Multi-level significance analysis of false-premise selection failures. All effects are computed as false premise minus baseline in percentage points. Pooled p values are two-sided task–cell paired-bootstrap p values over all eleven dataset–model cells with B=10,000 resamples. Sign consistency reports the number of dataset–model cells with the expected direction; p is from an exact two-sided binomial sign test. Cell-level significance reports the number of dataset–model cells with per-cell p<0.05 .
Dataset
Hard
Medium
Easy
HumanEval+
+0.88
+5.92
+6.87
MBPP
+8.60
+6.28
+8.38
LiveCodeBench
+5.12
+3.50
+2.86
Appendix
Table S2: False-premise-induced Trap@64 changes across task-difficulty groups. Values are percentage-point changes from the baseline condition to the false-premise condition. Difficulty groups are defined using only the original clean task prompt.
Model
Δ Selected
Δ Oracle
Δ Trap
Qwen-0.5B
−13.3
+0.7
+14.0
Qwen-1.5B
−10.0
+0.7
+10.7
Qwen-3B
−7.3
−2.7
+4.7
Qwen-7B
−7.9
+0.4
+8.4
Qwen-14B
−9.2
−3.2
+6.0
Qwen-32B
−6.0
0.0
+6.0
Appendix
Table S3: False-premise-induced changes across Qwen2.5-Coder model scales on MBPP. Values are percentage-point changes from the baseline condition to the false-premise condition for Selected@64, Oracle@64, and Trap@64.
Model
Baseline
Irrelevant
False
Δ
HumanEval+
7B
0.436
0.432
0.339
−9.6 pp
14B
0.428
0.420
0.298
−12.9 pp
32B
0.406
0.418
0.218
−18.9 pp
MBPP
7B
0.382
0.351
0.228
−15.4 pp
Appendix
Table S4: Counterexample coverage of verifier-written tests on HumanEval+ and MBPP. We report the fraction of detectable list-input asserts whose input list violates the sortedness premise. Δ is computed as false premise minus baseline.
Verifier
Δ Trap@64
Qwen-7B
−2.86
Qwen-32B
−7.86
GPT-4o-mini
−4.29
Appendix
Table S5: Effect of verifier prompting on fixed false-premise candidate pools. Candidate pools are generated by Qwen-7B on HumanEval+. Δ Trap@64 is computed as clean minus dirty; negative values favor the clean verifier condition.
Cell
Condition
T=10
T=15
T=20
T=25
T=30
HE+ 7B
False
−0.40
−0.40
−0.40
−0.40
+0.00
HE+ 7B
Irrelevant
+0.95
+0.48
−1.43
−0.95
−0.95
HE+ 7B
Baseline
−1.78
−1.78
−2.23
−1.34
−0.45
MBPP 7B
False
−0.71
−0.24
−1.88
−3.06
−2.59
MBPP 7B
Irrelevant
−0.95
−0.48
−1.19
−0.95
−1.91
MBPP 7B
Baseline
−0.69
−0.91
−0.91
+0.00
−0.23
Appendix
Table S6: Adding more coupled verifier-written tests does not systematically reduce Trap@64. Values report the change in Trap@64 relative to T=5 , where T denotes the number of verifier-written tests used for selection. Positive values indicate a higher trap rate than T=5 .
Test Dataset
Test Model
Training Source
Base
P@20%
Lift
Recall@20%
Cross-model transfer
HumanEval+
Qwen-14B
HumanEval+ / Qwen-7B
5.8
18.4
3.2 ×
63.6
HumanEval+
Qwen-32B
HumanEval+ / Qwen-7B
10.1
36.8
3.6 ×
73.7
HumanEval+
Qwen-7B
HumanEval+ / Qwen-14B+32B
11.0
44.4
4.0 ×
80.0
Cross-dataset transfer
MBPP
Qwen-7B
HumanEval+ / Qwen-7B
16.7
31.8
1.9 ×
37.8
Appendix
Table S7: Cross-setting transfer of gold-free trap-risk prediction. Base is the trap rate in the held-out split. P@20% is the trap precision among the top 20% highest-risk tasks, and Lift is P@20% divided by Base. Recall@20% is the fraction of all traps covered by the top 20% highest-risk tasks. All predictors use only verifier-visible features and exclude hidden-test-derived inputs.
Method
Auditor Input
Pass
Δ Pass
Original verifier
False verifier tests
77.03
0.00
Dirty auditor
False prompt + code
80.41
+3.38
Clean auditor
Clean prompt + code
81.08
+4.05
Appendix
Table S8: Auditor prompt-channel ablation on HumanEval+ with Qwen-7B. The dirty auditor receives the same false-premise-conditioned task prompt as the original verifier, while the clean auditor receives the premise-removed task prompt. Dirty auditing recovers most of the clean-auditor gain, indicating that the mitigation is not explained solely by removing the false premise from the auditor input.
Method
Final Selector
Δ Pass
Original verifier
Original verifier
0.00
Augmented verifier
Original verifier
+0.68
Source-aware prior
Provenance-aware rule
+4.05
Robustness auditor
Audit score + verifier
+4.73
Appendix
Table S9: Controlled ablation of gate, pool, and final selection rule on HumanEval+ with Qwen-7B. All non-original variants use the same gold-free high-risk gate and the same augmented candidate pool. Re-ranking the augmented pool with the original verifier gives only limited recovery, whereas source-aware and auditor-based selection recover substantially more.
Group
Intervention
Cell
Pass
Trap
Recovery
Δ Pass
Original selection
Original
Original V-Top1
7B / HE+
77.03
12.84
0.00
0.00
Coupled methods
Coupled
Extra verifier tests
14B / HE
76.51
8.05
14.29
+1.34
Coupled
Width scaling under original V
7B / HE+
76.67
9.33
–
+8.67 †
Coupled
Robust regen + original V
7B / HE+
77.70
12.84
5.26
+0.67
Appendix
Table S10: Operating points from coupled and decoupled mitigation families. Coupled methods keep final selection inside the original verifier channel, while decoupled methods weaken or replace that channel. Recovery is computed relative to the original selector and Oracle@64 within the corresponding cell.
Dataset
Model
Original
Coupled
Δcoup
Decoupled
Δdec
Cross-model robustness
HumanEval+
Qwen-7B
77.03
77.70
+0.67
82.43
+5.40
HumanEval+
Qwen-14B
75.17
75.84
+0.67
79.87
+4.70
HumanEval+
DS-Coder-6.7B
63.89
63.89
0.00
65.97
+2.08
Cross-dataset robustness
MBPP
Qwen-7B
64.67
64.67
0.00
67.33
+2.66
Appendix
Table S11: Cross-model and cross-dataset robustness of decoupled mitigation. Decoupled variants consistently outperform coupled variants across the tested cells.
Model
Coupled Gain
Decoupled Gain
Qwen-7B
0.00
0.00
Qwen-14B
0.00
+0.67
Qwen-32B
0.00
+1.36
Appendix
Table S12: LiveCodeBench mitigation gains across Qwen model scales. Decoupled mitigation matches or exceeds coupled gains in Selected@64 at every evaluated scale.
Visible tests are a common gate for LLM-generated code, but passing them does not certify specification correctness. We study a deployment-like monitoring problem: after code has passed public tests, can a weaker LLM verifier identify the residual hidden bugs? We introduce Code Monitor Red Teaming, a monitor-red-teaming protocol that fixes a public-check information boundary while varying generator pressure, verifier scaffolding, and weak-to-strong capability. We instantiate it as CodeMonitorBench, spanning function-level, data-science, and workflow code. Across 71,000 generated candidates, 43,677 pass public tests and 23,081 of those fail hidden tests. Weak verifiers improve with scaffolding and model family, but still miss most hidden bugs at 5% false-positive rate. As a robustness stress test, adversarial public-test-overfit pressure lowers verifier AUROC and raises low-FPR miss rates in most cells. A GLM-5.1 verifier recovers part of the gap under the same evidence boundary; an inferability audit shows that remaining misses mix verifier failures with M1 evidence limits.
Junchi Liao, Jiawen Deng, Fuji Ren
University of Electronic Science and Technology of China Chengdu, China
Selecting LLM-generated code candidates using LLM-generated tests is challenging because the tests themselves may be incorrect. Existing methods either treat all tests equally or rely on ad-hoc heuristics to filter unreliable tests. Yet determining test correctness requires knowing which codes are correct, creating a \emph{circular dependency}. Our key insight is that we need not determine test correctness at all: \emph{test votes should rank, not merely count}. What matters is not how many codes pass a test, but whether the test can \emph{distinguish} correct from incorrect code. We break the circular dependency via leave-one-out evaluation: hold out one test, rank codes by their aggregate scores on all remaining tests, and measure whether the held-out test's pass/fail pattern agrees with this ranking. We formalize this agreement as the leave-one-out AUC~(LOO-AUC) and prove that the expected LOO-AUC is proportional to each test's ability to separate correct code from incorrect code. Building on this, we propose \textbf{ACES}~(\textbf{A}UC \textbf{C}onsist\textbf{E}ncy \textbf{S}coring) with two complementary variants: ACES-C provides closed-form weights that provably approximate the oracle in expectation under a mild assumption on average test quality; ACES-O drops this assumption and iteratively optimizes a differentiable LOO-AUC objective. Both operate solely on the binary pass matrix with negligible overhead, and achieve state-of-the-art Pass@k on multiple code generation benchmarks.
Hui Sun, Yun-Ji Zhang, Zheng Xie +4
National Key Laboratory for Novel Software Technology, Nanjing University, China · School of Artificial Intelligence, Nanjing University, China
Large language models can generate useful code from natural language, but their outputs come without correctness guarantees. Verifiable code generation offers a path beyond testing by requiring models to produce not only executable code, but also formal specifications and machine-checkable proofs. Progress in this direction, however, is difficult to measure: existing benchmarks are often small, focus on only one part of the pipeline, lack ground-truth proofs or rigorous specification validation, or target verification settings far from mainstream software development. We present VeriContest, a benchmark of 946 competitive-programming problems from LeetCode and Codeforces for verifiable code generation in Rust with Verus. Each problem pairs a natural language description with expert-validated formal specifications, judge-accepted Rust code, Verus-checked proofs, and positive and negative test suites. VeriContest is constructed through a three-phase pipeline that scales from manually verified seed problems to semi-automated expansion with human-in-the-loop review. To further strengthen benchmark quality, we use testing as an additional quality-assurance layer for validating postcondition completeness. VeriContest supports isolated and compositional evaluation of specification generation, code generation, proof generation, and end-to-end verified program synthesis. Evaluating ten state-of-the-art models reveals a sharp gap between coding ability and verifiable code generation: the strongest model reaches 92.18% on natural-language-to-code generation, but only 48.31% on specification generation, 13.95% on proof generation, and 5.29% end-to-end. These results identify proof and specification generation as the central bottlenecks for models and establish VeriContest as a rigorous platform for measuring and training future systems that generate code with machine-checkable correctness.
Zichen Xie, Mrigank Pawagi, Yuxin Liu +5
University of Virginia · 2Indian Institute of Science · 3Rice University +1