cs.CLSep 27, 2026

Knowing Is Not Choosing: What Explicit Verification Adds Beyond Generative Preference

Authors: Yilong Li, Chengpo Yan, Aayan Arish, Suman Banerjee

Organizations: Department of Computer Science University of Wisconsin-Madison Madison, WI 53703, USA

Abstract

Generating a correct answer does not mean that a language model will select it. We separate factual recall into three steps: generating a correct candidate, ranking the available candidates, and selecting the final answer. Pre-generation readouts predict factual recall and which questions sampling will cover across three model families, but say little about whether an available correct answer will ultimately be selected. Explicit verification with P(True)P(\mathrm{True}) improves within-question ranking over mean log-likelihood in Gemma, Qwen3, and Llama, with AUROC gains of 0.080.08--0.120.12. In a prospectively defined Gemma cohort, verification raises plurality accuracy by about 55 points, and still gains about 22 points over chat-template likelihood, a stronger generative baseline. The advantage is strongest for relations with common-answer priors and depends on access to the entity; masking the entity removes the ranking advantage in larger Qwen models. Finally, the measured benefit depends on how correctness is defined: recall-oriented reference matching can credit option lists favored by likelihood and substantially understate the improvement seen under human semantic judgments. Prior work shows that models can carry latent factual knowledge and judge candidate answers; we show that these capabilities do not collapse into a single notion of ``knowing,'' and trace where information is gained, lost, or mismeasured between availability, ranking, and final choice.

Figures & tables

Appendix figures & tables51 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. The Future of Facts: Tracing the Factual Generation-Verification Gap

    May 26, 2026Tim R. Davidson, Anja Surina, Caglar GulcehreFactsReasoning Gaps

  2. Only Say What You Know: Calibration-Aware Generation for Long-Form Factuality

    May 3, 2026Wen Luo, Guangyue Peng, Liang Wang +7FactsLong-Form Generation