Locating Answer-Correctness Signals in Frozen Large Language Models
Organizations: School of Computing, National University of Singapore
Abstract
Language models expose internal signals that predict whether an answer is correct, readable from a single forward pass of a frozen model without additional generations. Yet existing probes often commit to one signal family or layer and can be brittle under distribution shift; in retrieval-augmented settings, many specialized detectors instead target passage faithfulness, which can diverge from correctness when retrieved evidence is unhelpful or conflicting. We therefore ask where answer correctness is readable, which internal signal families carry it, and how they should be combined. We search over hidden states, token probabilities, residual-stream features, attention, and their fusion, treating the selected readouts as a predictive measurement rather than a mechanistic localization. We run this analysis separately in closed-book and with-context settings, since context can change which readouts are informative. A consistent anatomy emerges: correctness concentrates in the answer span, recovered from the answer tokens even under retrieval, and the families carry it complementarily, so fusing them helps most out of distribution, where a single signal is weakest. The protocol is effective across two backbones and gates a retrieval controller as one downstream use.
Figures & tables
| Closed-book | With context | |||||||
| ID | OOD | ID | OOD | |||||
| Method | Test | PopQA | NQ | TQA | Test | PopQA | NQ | TQA |
| Sampling-based (10 extra generations per query) | ||||||||
| Semantic Entropy [ 18 ] | .790/.891 | .876/.957 | .758/.751 | .917 / .808 | .773/.527 | .751/.811 | .561/.684 | .836/.567 |
| Bayesian SE [ 34 ] | .801/.893 | .884/.957 | .758/.767 | .921 / .821 | .773/.522 | .753/.813 | .564/.692 | .840/.571 |
| SelfCheckGPT [ 21 ] | .805 / .898 | .881/.956 | .802 / .835 | .901/.779 | .699/.546 | .645/.809 | .511/.723 | .821/.581 |
| Closed-book | With context | |||||||||
| ID | OOD | ID | OOD | |||||||
| Variant | Encoder | Test | PopQA | NQ | TQA | Encoder | Test | PopQA | NQ | TQA |
| Hidden-only | ( ) | .834/.910 | .917 / .973 | .795/.828 | .859/.735 | [ ] | .929/.814 | .941/.957 | .880/.926 | .910/.738 |
| w/o Hidden | ( )( )( )( ) | .832/.910 | .900/.969 | .754/.802 | .825/.685 | [ ] | .910/.780 | .937/.958 | .854/.916 | .891/.663 |
| Early fusion | ( ) | .839 / .915 | .913/.972 | .813 / .841 | .870 / .755 | [ ] | .927/.808 | .944/.962 | .877/.926 | .922 /.739 |
| Late fusion (single) | ( )( )( )( ) | .837 /.910 | .912/.971 | .785/.824 | .853/.721 | [ ][ ][ ][ ] | .929/.819 | .942/.959 | .878/.928 | .917/.736 |
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
| Closed-book | With-context | ||||
| Split | Dataset | Correct (%) | Orig. | Rewr. | |
| Train | HotpotQA | 39.8 | 95.2 | 95.5 | |
| MuSiQue | 17.8 | 60.8 | 62.5 | ||
| 2WikiMultihopQA | 42.8 | 95.4 | 96.0 | ||
| All | 33.5 | 83.8 | 84.7 | ||
| Val | HotpotQA | 40.4 | 92.0 | 91.6 | |
| In-domain | Out-of-distribution | ||||||
| Method | MuS | HoP | 2Wi | Test | PopQA | NQ | TQA |
| Random AUPRC (error rate) | .860 | .622 | .636 | .706 | .776 | .573 | .292 |
| BEATRICE | .804 / .955 | .844/.866 | .792 / .871 | .839 / .913 | .924 / .976 | .814 / .847 | .874/.762 |
| SelfCheckGPT-NLI | .756/.944 | .860 / .909 | .789/.855 | .805/.898 | .881/.956 | .802/.835 | .901/.779 |
| Bayesian SE | .708/.924 | .848/.883 | .776/.862 | .801/.893 | .884/.957 | .758/.767 | .921 / .821 |
| LLMsKnow | .735/.934 | .802/.848 | .757/.853 | .795/.891 | .859/.947 | .749/.787 | .828/.646 |
| In-domain | Out-of-distribution | ||||||
| Method | MuS | HoP | 2Wi | Test | PopQA | NQ | TQA |
| Random AUPRC (error rate) | .486 | .066 | .106 | .219 | .607 | .675 | .219 |
| BEATRICE | .874 / .869 | .829 / .437 | .954 / .800 | .932 / .826 | .948 / .965 | .896 / .941 | .930 / .774 |
| SAPLMA (mean) | .819/.822 | .726/.209 | .925/.738 | .892/.744 | .924/.937 | .847/.899 | .883/.637 |
| LLMsKnow | .815/.781 | .743/.250 | .929/.671 | .884/.703 | .912/.935 | .831/.898 | .851/.620 |
| SAPLMA (last) | .794/.775 | .760/.269 | .924/.626 | .883/.696 | .908/.930 | .805/.875 | .856/.601 |
| Setting | Data | Strongest baseline | AUROC [ CI] | |
| Closed- book | Test | SelfCheck-NLI (.805) | ||
| PopQA | SelfCheck-ngram (.906) | |||
| NQ | SelfCheck-NLI (.802) | |||
| TQA | Bayesian SE (.921) | |||
| With context | Test | SAPLMA (.892) | ||
| PopQA | SAPLMA (.924) |
| Setting | Data | Champion | Hidden only | (95% CI) |
| Closed-book | Test | .839 | .832 | .007 † [.000, .014] |
| PopQA | .924 | .917 | .007 [ .000, .014] | |
| NQ | .814 | .793 | .021 † [.011, .032] | |
| TQA | .876 | .865 | .010 † [.000, .020] | |
| With context | Test | .926 | .921 | .005 [ .001, .011] |
| PopQA | .948 | .942 | .006 [ .000, .012] |
| Closed-book | With context | |||||||
| Dropped span | Test | PopQA | NQ | TQA | Test | PopQA | NQ | TQA |
| None (full) | .839/.913 | .924/.976 | .814/.848 | .874/.762 | .932/.826 | .948/.965 | .896/.941 | .930/.774 |
| Context | – | – | – | – | .927/.809 | .944/.959 | .883/.929 | .925/.758 |
| Question | .837/.914 | .920/.974 | .818/.848 | .882/.757 | .926/.800 | .943/.962 | .880/.931 | .927/.767 |
| Answer | .831/.912 | .911/.971 | .803/.841 | .842/.714 | .919/.801 | .944/.963 | .884/.932 | .903/.711 |
| Closed-book | With context | |||||||
| ID | OOD | ID | OOD | |||||
| Predictor | Test | PopQA | NQ | TQA | Test | PopQA | NQ | TQA |
| Text-only classifiers (question, answer, and retrieved passage as text) | ||||||||
| Majority class | .500 | .500 | .500 | .500 | .500 | .500 | .500 | .500 |
| TF-IDF + logistic regr. | .785 | .774 | .676 | .643 | .901 | .863 | .819 | .807 |
| TF-IDF + linear SVM | .746 | .760 | .658 | .629 | .869 | .836 | .787 | .765 |
| Closed-book | With-context | |||
| flip rate | flip precision | flip rate | flip precision | |
| Qwen3.5-9B | ||||
| Hidden | .385 | .598 [.560, .641] | .668 | .964 [.956, .972] |
| Prob | .039 | .746 [.644, .864] | .028 | .583 [.488, .690] |
| Resid | .120 | .556 [.489, .622] | .074 | .897 [.857, .933] |
| Attn | .034 | .569 [.431, .706] | .065 | .809 [.758, .861] |
| test | NQ | TriviaQA | |
| Generator answer accuracy (%) | |||
| closed-book | |||
| with the retrieved passage | |||
| Wrong answers present in passage | |||
| Attention flip precision | |||
| answer in passage, but wrong |
| Single-hop | Multi-hop | ||||||
| Method | PopQA | NQ | TQA | HotpotQA | 2Wiki | MuSiQue | All |
| Random AUPRC (error rate) | .504 | .396 | .158 | .492 | .706 | .882 | .523 |
| BEATRICE | .866 / .868 | .836 / .781 | .899/ .696 | .864 / .866 | .832 / .909 | .798/.964 | .891 / .894 |
| SAPLMA (mean) | .813/.809 | .770/.644 | .910 /.643 | .826/.780 | .757/.865 | .811 / .969 | .852/.836 |
| SAPLMA (last) | .811/.799 | .778/.653 | .844/.566 | .790/.774 | .775/.892 | .766/.958 | .836/.827 |
| LLMsKnow | .800/.789 | .756/.656 | .815/.562 | .795/.798 | .718/.843 | .713/.946 | .812/.813 |
| Single-hop | Multi-hop | ||||||
| Method | PopQA | NQ | TQA | HotpotQA | 2Wiki | MuSiQue | All |
| Random AUPRC (error rate) | .562 | .440 | .266 | .680 | .882 | .930 | .627 |
| BEATRICE | .877 / .912 | .819 / .802 | .883 / .805 | .922 / .963 | .943 / .992 | .892 / .991 | .911 / .950 |
| LLMsKnow | .765/.845 | .696/.690 | .749/.672 | .850/.925 | .847/.977 | .841/.985 | .847/.916 |
| SAPLMA (mean) | .782/.823 | .709/.673 | .739/.488 | .824/.903 | .800/.963 | .727/.970 | .806/.864 |
| SAPLMA (last) | .781/.842 | .665/.620 | .699/.478 | .830/.911 | .791/.965 | .799/.981 | .799/.872 |
| In-domain (multi-hop) | Out-of-distribution | Macro | |||||||
| Method | HotpotQA | 2WikiMQA | MuSiQue | PopQA | NQ | TriviaQA | EM | F1 | Gen |
| No-RAG | .220/.293 | .238/.282 | .034/.114 | .192/.246 | .176/.267 | .574/.617 | .239 | .303 | 1.00 |
| Vanilla RAG | .322/.437 | .260/.309 | .060/.133 | .447 / .526 | .371 / .485 | .757 / .818 | .370 | .451 | 1.00 |
| SKR | .310/.416 | .296/.342 | .052/.126 | .440/.519 | .364/.474 | .738/.796 | .367 | .445 | 1.00 |
| Adaptive-RAG | .322/.431 | .282/.323 | .060/.098 | .433/.496 | .319/.425 | .718/.769 | .356 | .424 | 1.63 |
| DRAGIN | .300/.394 | .298/.343 | .062/.151 | .415/.473 | .354/.467 | .672/.724 | .350 | .425 | 3.55 |
| Question | Closed-book answer (score) | With-context answer (score) | Gold | PGIR action |
| Qwen3.5-9B | ||||
| The Nun is based on what film series? | The Conjuring (.67) | not retrieved () | The Conjuring | Accept, no retrieval |
| What gemstone is The Moonstone in the novel by Wilkie Collins? | Moonstone (.57) | Diamond (.90) | Diamond | Retrieve, then trust |
| Arctic King, Saladin, and Tom Thumb are which vegetable? | Onions (.26) | Cauliflower (.01) | Lettuce | Retrieve; both flagged, none trusted |
| Gemma 4 | ||||
| Who played Clayton Farlowe in Dallas ? | Michael Landon (.05) | Howard Keel (.65) | Howard Keel | Retrieve, then trust |
| Closed-book | With context | |||||||
| ID | OOD | ID | OOD | |||||
| Method | Test | PopQA | NQ | TQA | Test | PopQA | NQ | TQA |
| Sampling-based (10 extra generations per query) | ||||||||
| Semantic Entropy | .339/.894 | .523/.888 | .525/.746 | .756/.717 | .595/.404 | .456/.648 | .384/.726 | .466/.428 |
| Bayesian SE | .323/.891 | .524/.888 | .513/.737 | .749/.713 | .591/.403 | .449/.647 | .367/.720 | .451/.425 |
| SelfCheckGPT-NLI | .199/.880 | .330/.883 | .485/.777 | .693/.729 | .390/.419 | .087/.628 | .169/.740 | .185/.448 |
| In-domain (multi-hop) | Out-of-distribution | Macro | |||||||
| Method | HotpotQA | 2WikiMQA | MuSiQue | PopQA | NQ | TriviaQA | EM | F1 | Gen |
| No-RAG | .050/.069 | .006/.012 | .000/.003 | .048/.052 | .030/.052 | .298/.340 | .072 | .088 | 1.00 |
| Vanilla RAG | .190/.298 | .086/.177 | .030/.078 | .375/ .463 | .284/.396 | .602/.692 | .261 | .351 | 1.00 |
| SKR | .190/.297 | .112/.196 | .028/.074 | .368/.456 | .280/.392 | .602/.690 | .263 | .351 | 1.00 |
| Adaptive-RAG | .222/.329 | .104/.194 | .040 / .097 | .364/.431 | .307 / .415 | .625/.700 | .277 | .361 | 1.70 |
| DRAGIN | .164/.237 | .154/.201 | .020/.062 | .228/.284 | .225/.323 | .514/.584 | .218 | .282 | 2.64 |