Knowing Is Not Choosing: What Explicit Verification Adds Beyond Generative Preference
Organizations: Department of Computer Science University of Wisconsin-Madison Madison, WI 53703, USA
Abstract
Generating a correct answer does not mean that a language model will select it. We separate factual recall into three steps: generating a correct candidate, ranking the available candidates, and selecting the final answer. Pre-generation readouts predict factual recall and which questions sampling will cover across three model families, but say little about whether an available correct answer will ultimately be selected. Explicit verification with improves within-question ranking over mean log-likelihood in Gemma, Qwen3, and Llama, with AUROC gains of --. In a prospectively defined Gemma cohort, verification raises plurality accuracy by about points, and still gains about points over chat-template likelihood, a stronger generative baseline. The advantage is strongest for relations with common-answer priors and depends on access to the entity; masking the entity removes the ranking advantage in larger Qwen models. Finally, the measured benefit depends on how correctness is defined: recall-oriented reference matching can credit option lists favored by likelihood and substantially understate the improvement seen under human semantic judgments. Prior work shows that models can carry latent factual knowledge and judge candidate answers; we show that these capabilities do not collapse into a single notion of ``knowing,'' and trace where information is gained, lost, or mismeasured between availability, ranking, and final choice.
Figures & tables
| Model | AUROC | Plurality accuracy (%) | Gain (points) | |||
|---|---|---|---|---|---|---|
| Gemma 2B | ||||||
| Gemma 9B | ||||||
| Gemma 27B | ||||||
| Llama 8B | ||||||
| Qwen3.5 2B | ||||||
| Qwen3.5 4B | ||||||
Appendix figures & tables51 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | Pair-weighted AUROC | Argmax selection | Plurality selection | Conversion |
|---|---|---|---|---|
| Mean log-likelihood | ||||
| Summed log-likelihood | ||||
| Plurality | — | — | ||
| Oracle | — | — | ||
| Single-answer decoding | ||||
| Coverage | Accuracy | Conversion | |
|---|---|---|---|
| Selection outcome | Questions |
|---|---|
| Both select a correct answer | |
| Only likelihood selects a correct answer | |
| Only selects a correct answer | |
| Neither selects a correct answer | |
| Oracle over the two selections | |
| Candidate coverage |
| Criterion | Selection | Likelihood | Difference | 95% CI | |
|---|---|---|---|---|---|
| Reference | Argmax | ||||
| Stricter | Argmax | ||||
| Reference | Plurality | ||||
| Stricter | Plurality |
| Method | Argmax selection | Plurality selection | Conversion |
|---|---|---|---|
| Mean log-likelihood, filtered |
| Model | Metric | Improvement | 95% CI | ||
|---|---|---|---|---|---|
| Gemma 2 2B-IT | Log loss | ||||
| Brier | |||||
| Qwen3-8B | Log loss | ||||
| Brier | |||||
| Llama 3.1 8B-Instruct | Log loss | ||||
| Brier |
| Gemma 2 2B-IT | Qwen3-8B | Llama 3.1 8B-Instruct | ||||
|---|---|---|---|---|---|---|
| Type | log loss | Brier | log loss | Brier | log loss | Brier |
| Player | ||||||
| City | ||||||
| Movie | ||||||
| Song | ||||||
| Readout | Log-loss improvement (95% CI) | Brier improvement (95% CI) |
|---|---|---|
| Latent | ||
| latent |
| Comparison | Questions | Query-macro AUROC | 95% CI |
|---|---|---|---|
| Greedy answer matches: matching vs. non-matching samples | |||
| Greedy answer misses: matching vs. non-matching samples |
| Temperature | Accuracy | Mixed questions | Query-macro AUROC | Pooled AUROC |
|---|---|---|---|---|
| Study | Model and population | Evaluation |
|---|---|---|
| Prospective recall | Gemma 2 2B-IT; new entities | Fit/calibration/confirmation roles of entities each |
| Recall replication | Qwen3-8B; new entities | Separate fixed feature, cohort, and protocol |
| Recall replication | Llama 3.1 8B-Instruct; entities | Gemma-matched cohort and roles; separate layer- class-mean direction |
| Earlier same-question sample | Gemma; questions | samples at , top- ; retained mixed questions |
| Development sampling cohort | Gemma; separate sample of questions | Stricter-criterion temperature cohorts: name-valued, answerable |
| -entity chat cohort | Gemma 2 2B-IT and 2B; entities | human-labeled chat responses; paired pretrained and instruction-tuned recall study |
| Analysis | Status |
|---|---|
| Gemma recall confirmation | Protocol fixed before measurement; positive |
| Qwen3 independent-cohort recall | Positive under the amended permutation limit |
| Llama matched-cohort recall | Positive; pre-specified controls met |
| Gemma prospective selection | Protocol fixed before generation |
| Llama selection inputs | Same questions and fixed rules as Gemma |
| Qwen3 selection inputs | Same questions and fixed rules as Gemma |
| Direction | SAE | Latent | MaxMin | Pile frequency |
|---|---|---|---|---|
| Known | layer_12/width_16k/average_l0_82 | |||
| Unknown | layer_14/width_16k/average_l0_84 |
| Condition | Refusals |
|---|---|
| Unsteered | |
| Zero direction | |
| Unknown latent | |
| Known latent | |
| Strongest matched random latent, unknown side |
| Predictor | AUROC | 95% interval |
|---|---|---|
| Latent | ||
| Random ordering | ||
| Baseline: exposure, missingness, type, relation | ||
| Baseline and latent |
| Policy | Retention | Correct | Wrong | Undeterminable | Risk bounds |
|---|---|---|---|---|---|
| Answer all | |||||
| Random | |||||
| Latent | |||||
| Baseline | |||||
| Baseline and latent |
| Bare format | Chat correct | Chat wrong | Chat refusal | Undeterminable |
|---|---|---|---|---|
| Hit | ||||
| Miss |
| Criterion | Likelihood | Difference | 95% CI | ||
|---|---|---|---|---|---|
| Stricter | |||||
| Reference | |||||
| Automatic outcome | Responses | Annotator A | Annotator B |
|---|---|---|---|
| Both criteria accept | |||
| Reference matching only | |||
| Neither criterion |
| Selector | Annotator A | Annotator B | Stricter criterion | Reference matching |
|---|---|---|---|---|
| Likelihood | ||||
| Criterion | Rule | Likelihood | Difference | 95% CI | |
|---|---|---|---|---|---|
| Stricter | Argmax | ||||
| Plurality | |||||
| Reference | Argmax | ||||
| Plurality |
| Model | Candidates | Mixed questions | Likelihood | Query-macro gain [95% CI] | |
|---|---|---|---|---|---|
| Gemma | |||||
| Llama |
| Model | Rule | Likelihood | Gain [95% CI] | |
|---|---|---|---|---|
| Gemma | Argmax | |||
| Plurality | ||||
| Llama | Argmax | |||
| Plurality |
| Model | Relations | Mixed | AUROC | Likelihood | Gain [95% CI] | ||
| Gemma | High-prior | ||||||
| Low-prior | |||||||
| High low | — | — | — | — | — | ||
| Llama | High-prior | ||||||
| Low-prior | |||||||
| High low | — | — | — | — | — |
| Model | Wording | Pair-weighted AUROC | Argmax | Plurality | Plurality gain [95% CI] |
|---|---|---|---|---|---|
| Gemma | Original | ||||
| Short | |||||
| True/False | |||||
| No/Yes | |||||
| Llama | Original | ||||
| Short |
| Coverage | Likelihood accuracy | Verification accuracy | Gain | Likelihood cost | Verification cost | |
|---|---|---|---|---|---|---|
| Model | Budget | Uniform | Random | Risk | R+T | R+T+S | Difference [95% CI] |
|---|---|---|---|---|---|---|---|
| Gemma | |||||||
| Llama | |||||||
| Model | Coverage | Likelihood accuracy | Verification | Likelihood time | Verification time | |
|---|---|---|---|---|---|---|
| Gemma | ||||||
| Llama | ||||||
| Model | Quintile | Entities | Questions | Covered | Coverage [95% CI] |
|---|---|---|---|---|---|
| Gemma | 1 (lowest) | ||||
| 2 | |||||
| 3 | |||||
| 4 | |||||
| 5 (highest) | |||||
| Llama | 1 (lowest) |
| Model | Target | [95% CI] | Brier [95% CI] | AUROC | |
|---|---|---|---|---|---|
| Gemma | Coverage | ||||
| Success, likelihood | |||||
| Success, | |||||
| Llama | Coverage | ||||
| Success, likelihood | |||||
| Success, |
| Measure | Likelihood | Difference [95% CI] | |
|---|---|---|---|
| Pair-weighted AUROC | |||
| Query-macro AUROC | |||
| Argmax selection | |||
| Plurality selection | |||
| Conversion (plurality) | — |
| Measure | Likelihood | Difference [95% CI] | |
|---|---|---|---|
| Pair-weighted AUROC | |||
| Argmax selection | |||
| Plurality selection | |||
| Conversion (plurality) | — | ||
| Option-list selections | — |
| Measure | Likelihood | Difference [95% CI] | |
|---|---|---|---|
| Pair-weighted AUROC | |||
| Plurality selection | |||
| Conversion (plurality) | — |
| Conversion | ||||||
|---|---|---|---|---|---|---|
| Model | Eligible | Coverage | Likelihood argmax | Likelihood plurality | plurality | Retention |
| Qwen3.5-2B | ||||||
| Qwen3.5-4B | ||||||
| Qwen3-8B | ||||||
| Qwen3.6-27B | ||||||
| Qwen3.8-27B | ||||||
| Model | Pair-weighted AUROC | Query-macro | Argmax accuracy (%) | Gain (points) | |
|---|---|---|---|---|---|
| Qwen3.5-2B | |||||
| Qwen3.5-4B | |||||
| Qwen3.6-27B | |||||
| Labels | Coverage | Pair-weighted AUROC | Argmax gain | Plurality gain |
| Reference | ||||
| Stricter | ||||
| Semantic | ||||
| Human | — | — | ||
| Semantic reference | — | — | ||
| Stricter semantic | — | — |
| Selection | Likelihood | Difference | Option lists | |
|---|---|---|---|---|
| Argmax | vs. | |||
| Plurality | vs. |
| Model | Target | AUROC | [95% CI] | Brier [95% CI] | |
| Gemma | Coverage | ||||
| Success, likelihood | — | ||||
| Success, | — | ||||
| Llama | Coverage | ||||
| Success, likelihood | — | ||||
| Success, | — |
| Score | High-prior | Low-prior | All |
|---|---|---|---|
| Mean log-likelihood | ( ) | ( ) | ( ) |
| Chat-template likelihood | ( ) | ( ) | ( ) |
| PMI | ( ) | ( ) | ( ) |
| ( ) | ( ) | ( ) |
| Score | Qwen3.6-27B | Qwen3.8-27B |
|---|---|---|
| Mean log-likelihood | / | / |
| Summed log-likelihood | / | / |
| Chat-template likelihood | / | / |
| PMI | / | / |
| / | / | |
| Masked | / | / |
| Model | Measure | chat | Masked chat | Unmasked masked |
|---|---|---|---|---|
| Qwen3.6-27B | Accuracy | |||
| Pair-weighted | ||||
| Query-macro | ||||
| Source | Measurement | Result |
|---|---|---|
| Residual state | Last prompt token, layers | |
| Subject attention | layers, heads | |
| Output probability | Greedy chat answer, raw AUROC | |
| Binding | Subject-by-relation contrast |
| Model | Yes | No | True | False |
|---|---|---|---|---|
| Gemma 2 2B-IT | ||||
| Llama 3.1 8B-Instruct | ||||
| Qwen3-8B | — | — |