cs.AISep 29, 2026

Locating Answer-Correctness Signals in Frozen Large Language Models

Authors: Yuansen Liu, Yixuan Tang, Anthony Kum Hoe Tung

Organizations: School of Computing, National University of Singapore

Abstract

Language models expose internal signals that predict whether an answer is correct, readable from a single forward pass of a frozen model without additional generations. Yet existing probes often commit to one signal family or layer and can be brittle under distribution shift; in retrieval-augmented settings, many specialized detectors instead target passage faithfulness, which can diverge from correctness when retrieved evidence is unhelpful or conflicting. We therefore ask where answer correctness is readable, which internal signal families carry it, and how they should be combined. We search over hidden states, token probabilities, residual-stream features, attention, and their fusion, treating the selected readouts as a predictive measurement rather than a mechanistic localization. We run this analysis separately in closed-book and with-context settings, since context can change which readouts are informative. A consistent anatomy emerges: correctness concentrates in the answer span, recovered from the answer tokens even under retrieval, and the families carry it complementarily, so fusing them helps most out of distribution, where a single signal is weakest. The protocol is effective across two backbones and gates a retrieval controller as one downstream use.

Figures & tables

Appendix figures & tables20 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. What Am I Missing? Question-Answering as Hidden State Probing

    May 29, 2026Chu Fei Luo, Samuel Dahan, Xiaodan ZhuLLM Reasoning StrategiesHidden States

  2. On the Robustness of LLMs' Internal Representation of Code Correctness

    Aug 8, 2026Francisco Ribeiro, Sohaila Abdulsattar, Renata Gonzalez +2CorrectnessDyads