Language models expose internal signals that predict whether an answer is correct, readable from a single forward pass of a frozen model without additional generations. Yet existing probes often commit to one signal family or layer and can be brittle under distribution shift; in retrieval-augmented settings, many specialized detectors instead target passage faithfulness, which can diverge from correctness when retrieved evidence is unhelpful or conflicting. We therefore ask where answer correctness is readable, which internal signal families carry it, and how they should be combined. We search over hidden states, token probabilities, residual-stream features, attention, and their fusion, treating the selected readouts as a predictive measurement rather than a mechanistic localization. We run this analysis separately in closed-book and with-context settings, since context can change which readouts are informative. A consistent anatomy emerges: correctness concentrates in the answer span, recovered from the answer tokens even under retrieval, and the families carry it complementarily, so fusing them helps most out of distribution, where a single signal is weakest. The protocol is effective across two backbones and gates a retrieval controller as one downstream use.
Figures & tables
Figure 1: This work in one view. Top: in either setting, one forward pass of the frozen LLM yields four families of internal signals, which the BEATRICE detector reads into a probability that the answer is correct. Bottom: in the PGIR pipeline the closed-book estimate decides whether to retrieve at all, the with-context estimate whether to trust the grounded answer or retrieve again.
Figure 2: The BEATRICE pipeline. (1) Data and labels: the frozen LLM answers once per setting and a judge LLM keeps only unanimously labeled answers. (2) Combination encoders: each of the 15 signal-family combinations trains an encoder that pools token-level signals per segment into a correctness probability (early fusion). (3) Selection: greedy forward selection grows the ensemble on validation until the gain plateaus (late fusion).
Closed-book
With context
ID
OOD
ID
OOD
Method
Test
PopQA
NQ
TQA
Test
PopQA
NQ
TQA
Sampling-based (10 extra generations per query)
Semantic Entropy [ 18 ]
.790/.891
.876/.957
.758/.751
.917 / .808
.773/.527
.751/.811
.561/.684
.836/.567
Bayesian SE [ 34 ]
.801/.893
.884/.957
.758/.767
.921 / .821
.773/.522
.753/.813
.564/.692
.840/.571
SelfCheckGPT [ 21 ]
.805 / .898
.881/.956
.802 / .835
.901/.779
.699/.546
.645/.809
.511/.723
.821/.581
Table 1: Detecting incorrect answers of Qwen3.5-9B. Each cell is AUROC / AUPRC; per column and metric bold is best and underline second, and “–” marks detectors that require context.
Closed-book
With context
ID
OOD
ID
OOD
Variant
Encoder
Test
PopQA
NQ
TQA
Encoder
Test
PopQA
NQ
TQA
Hidden-only
( )
.834/.910
.917 / .973
.795/.828
.859/.735
[ ]
.929/.814
.941/.957
.880/.926
.910/.738
w/o Hidden
( )( )( )( )
.832/.910
.900/.969
.754/.802
.825/.685
[ ]
.910/.780
.937/.958
.854/.916
.891/.663
Early fusion
( )
.839 / .915
.913/.972
.813 / .841
.870 / .755
[ ]
.927/.808
.944/.962
.877/.926
.922 /.739
Late fusion (single)
( )( )( )( )
.837 /.910
.912/.971
.785/.824
.853/.721
[ ][ ][ ][ ]
.929/.819
.942/.959
.878/.928
.917/.736
Table 2: Feature and fusion ablation of BEATRICE. Each cell reports AUROC / AUPRC , 3-seed mean. The Encoder columns list the selected encoders in each setting: a bracket group is one encoder whose glyphs are early-fused, and groups side by side are late-fused; [ ⋅ ] marks a with-context forward pass and ( ⋅ ) a closed-book one. Glyphs: Hidden, Prob, Resid, Attn. Bold = best, underline = second in each column per metric. The greyed Late fusion (all) row late-fuses every encoder; it is a saturation reference, greedy selection matches it to within noise, and it is excluded from the ranking.
Figure 3: Signal complementarity on the in-domain test set. Each cell is the fraction of the column signal’s errors that the row signal classifies correctly, with every detector thresholded to flag its own true number of errors and columns normalised so the two settings compare; the bars give each signal’s own error rate.
Figure 4: Validation AUROC of all fifteen combination encoders, ordered best to worst. The dot matrix marks the signals each encoder fuses, and bars are darker when the encoder includes the hidden-state signal. AUROC axes are cropped to the observed range.
Figure 5: Exhaustive ranking of all three-encoder combinations, best to worst. The ribbon is the local mean of how many of the three encoders carry that signal, split by whether the carrying encoder is with-context (dark) or closed-book (light); for the with-context pool the four wedges track how the closed-book count concentrates along the ranking. The dashed line marks the selected combination.
Figure 6: PGIR control flow. A draft is accepted when its calibrated correctness probability clears τ ; otherwise the controller retrieves and re-answers, up to a budget of B rounds, then returns the highest-scoring draft in the pool.
Figure 7: Efficiency–accuracy frontier: macro F1 against generations per question, a serving-cost proxy (the x -axis is broken for the far-right system). Points sweep PGIR’s budget B ; B=4 is the reported configuration.
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Closed-book
With-context
Split
Dataset
N
Correct (%)
Orig.
Rewr.
Train
HotpotQA
10,000×2
39.8
95.2
95.5
MuSiQue
10,000×2
17.8
60.8
62.5
2WikiMultihopQA
10,000×2
42.8
95.4
96.0
All
30,000×2
33.5
83.8
84.7
Val
HotpotQA
500×2
40.4
92.0
91.6
Appendix
Table 3: Correctness-label distribution across all splits, reported as the fraction of examples the backbone answers correctly (the positive class; the incorrect fraction is the complement). N is the number of examples in that split; for with-context we give the rate on the original passages ( Orig. ) and on the normalization rewrite ( Rewr. ) of Appendix A separately; ×2 marks the splits whose with-context set combines both and is thus twice N .
Figure 8: Validation layer sweep: correctness AUROC of the hidden-state signal read from each layer. Layer 17 is the maximum and is used throughout.
In-domain
Out-of-distribution
Method
MuS
HoP
2Wi
Test
PopQA
NQ
TQA
Random AUPRC (error rate)
.860
.622
.636
.706
.776
.573
.292
BEATRICE
.804 / .955
.844/.866
.792 / .871
.839 / .913
.924 / .976
.814 / .847
.874/.762
SelfCheckGPT-NLI
.756/.944
.860 / .909
.789/.855
.805/.898
.881/.956
.802/.835
.901/.779
Bayesian SE
.708/.924
.848/.883
.776/.862
.801/.893
.884/.957
.758/.767
.921 / .821
LLMsKnow
.735/.934
.802/.848
.757/.853
.795/.891
.859/.947
.749/.787
.828/.646
Appendix
Table 4: Closed-book detection panorama: AUROC / AUPRC on every dataset for all baselines and BEATRICE. Positive class for AUPRC is the incorrect answer; the italic row is the random-guess AUPRC (per-column error rate). Bold marks the best in each column across all rows. In-domain test sizes: MuSiQue 499, HotpotQA and 2WikiMultihopQA 500 each; OOD covers PopQA/NQ/TriviaQA on which the probe is zero-shot.
In-domain
Out-of-distribution
Method
MuS
HoP
2Wi
Test
PopQA
NQ
TQA
Random AUPRC (error rate)
.486
.066
.106
.219
.607
.675
.219
BEATRICE
.874 / .869
.829 / .437
.954 / .800
.932 / .826
.948 / .965
.896 / .941
.930 / .774
SAPLMA (mean)
.819/.822
.726/.209
.925/.738
.892/.744
.924/.937
.847/.899
.883/.637
LLMsKnow
.815/.781
.743/.250
.929/.671
.884/.703
.912/.935
.831/.898
.851/.620
SAPLMA (last)
.794/.775
.760/.269
.924/.626
.883/.696
.908/.930
.805/.875
.856/.601
Appendix
Table 5: Retrieval-augmented (with-context) detection panorama: AUROC / AUPRC on every dataset for all baselines and BEATRICE, including the RAG-specific detectors (Lookback-Lens, ReDeEP variants, LUMINA variants). Conventions as in Table 4 . In-domain test sizes: 500 per multi-hop set.
Setting
Data
Strongest baseline
Δ AUROC [ 95% CI]
p
Closed- book
Test
SelfCheck-NLI (.805)
+.034[+.010,+.060]†
.007
PopQA
SelfCheck-ngram (.906)
+.018[−.003,+.037]
.091
NQ
SelfCheck-NLI (.802)
+.012[−.016,+.041]
.41
TQA
Bayesian SE (.921)
−.046[−.066,−.026]†
<10−4
With context
Test
SAPLMA (.892)
+.034[+.021,+.049]†
<10−4
PopQA
SAPLMA (.924)
+.025[+.014,+.036]†
<10−4
Appendix
Table 6: Paired bootstrap of the AUROC difference between BEATRICE and the strongest of all 23 baselines on each column, sampling-based methods included ( 104 resamples of the shared examples; two-sided p ). The with-context pairing scores the original contexts, the set every baseline scores. A dagger marks differences whose 95% interval excludes zero.
Setting
Data
Champion
Hidden only
Δ (95% CI)
Closed-book
Test
.839
.832
+ .007 † [.000, .014]
PopQA
.924
.917
+ .007 [ − .000, .014]
NQ
.814
.793
+ .021 † [.011, .032]
TQA
.876
.865
+ .010 † [.000, .020]
With context
Test
.926
.921
+ .005 [ − .001, .011]
PopQA
.948
.942
+ .006 [ − .000, .012]
Appendix
Table 7: Paired significance of the fused champion over the strongest single signal, the hidden-state encoder read on its own. For each evaluation set we bootstrap the AUROC difference ( 10,000 resamples over the examples the two models share) and report the 95% interval; a dagger marks the columns whose interval excludes zero. The champion improves on the hidden-state encoder in every column, and the improvement is significant on both held-out single-hop sets (NQ, TriviaQA) in each setting and on the closed-book in-domain test, while it is smallest and not significant on the in-domain aggregate with context. Figures are computed on the shared examples (original passages), so the with-context in-domain value is marginally below Table 1 .
Closed-book
With context
Dropped span
Test
PopQA
NQ
TQA
Test
PopQA
NQ
TQA
None (full)
.839/.913
.924/.976
.814/.848
.874/.762
.932/.826
.948/.965
.896/.941
.930/.774
Context
–
–
–
–
.927/.809
.944/.959
.883/.929
.925/.758
Question
.837/.914
.920/.974
.818/.848
.882/.757
.926/.800
.943/.962
.880/.931
.927/.767
Answer
.831/.912
.911/.971
.803/.841
.842/.714
.919/.801
.944/.963
.884/.932
.903/.711
Appendix
Table 8: Token-span ablation: encoders retrained with one span’s features zeroed at training time (evaluation protocol unchanged; each cell reports AUROC / AUPRC with the incorrect answer as the AUPRC positive class). Dropping the answer span is consistently the most damaging cut in both settings (with context: −.013 Test / −.027 TQA AUROC; closed-book: −.013 PopQA / −.032 TQA), while dropping the question span is nearly free — and even helps slightly on closed-book OOD — indicating the correctness signal concentrates on the answer tokens. Dropping the retrieved context costs little for detection even in the with-context setting, consistent with the probe reading the model’s processing of the context off the answer tokens rather than the context tokens themselves.
Closed-book
With context
ID
OOD
ID
OOD
Predictor
Test
PopQA
NQ
TQA
Test
PopQA
NQ
TQA
Text-only classifiers (question, answer, and retrieved passage as text)
Majority class
.500
.500
.500
.500
.500
.500
.500
.500
TF-IDF + logistic regr.
.785
.774
.676
.643
.901
.863
.819
.807
TF-IDF + linear SVM
.746
.760
.658
.629
.869
.836
.787
.765
Appendix
Table 9: Predicting answer correctness of Qwen3.5-9B from surface text alone, against BEATRICE. Each cell is AUROC. Every text classifier is trained on the same training split and selected on the same validation split as BEATRICE; its input is the question, the model’s final answer, and, in the with-context setting, the retrieved passage. In-domain Test is the multi-hop aggregate; PopQA, NQ, and TQA are out-of-distribution. The two figures after BEATRICE are its closed-book and with-context champion sizes; both are far smaller than the fine-tuned encoder yet lead every column. Per column, bold is best.
Figure 9: When the auxiliary families change a decision, with-context branches on Qwen3.5-9B, full test split. Each panel plots the encoder score with the family ablated ( x ) against the family’s contribution ( y ); in the shaded wedge the contribution changes the decision, and colored points are the flipped examples. Prob and Attn are measured in [ ] , Resid in [ ] .
Figure 10: Cross-backbone structure of the fifteen combination encoders, full test split. Top: each encoder’s AUROC on Qwen3.5-9B ( x ) against Gemma 4 ( y ), one panel per setting; dark points read hidden states, light ones do not. Bottom: the AUROC of each two-family encoder minus the AUROC of the better of its two single-family members ( ×100 ).
Figure 11: Attribution share of each layer–head cell of the attention family on Qwen3.5-9B: absolute integrated-gradients attribution of the encoder score, summed over that head’s channels and all spans, n=400 test examples. Top: the with-context combination branch [ ] ; bottom: the single-family encoder [ ] . The mark on the colorbar is the uniform share.
Closed-book
With-context
flip rate
flip precision
flip rate
flip precision
Qwen3.5-9B
Hidden
.385
.598 [.560, .641]
.668
.964 [.956, .972]
Prob
.039
.746 [.644, .864]
.028
.583 [.488, .690]
Resid
.120
.556 [.489, .622]
.074
.897 [.857, .933]
Attn
.034
.569 [.431, .706]
.065
.809 [.758, .861]
Appendix
Table 10: Decision flips under family ablation, full test split. One signal family’s input channels are replaced by their means over the evaluation data; a flip means the encoder’s decision, P(answer correct) above or below one half, changes. Flip rate is the fraction of examples flipped; flip precision is the fraction of flips on which the intact encoder decides correctly, with 95% bootstrap intervals. Closed-book rows are measured in ( ) , with-context rows in [ ] ; the Resid rows use ( ) and [ ] . Bold marks the most precise auxiliary family per setting. Gemma closed-book precisions should be read against that split’s 9% positive rate.
test
NQ
TriviaQA
Generator answer accuracy (%)
closed-book
29.4
42.7
70.8
with the retrieved passage
78.4
32.5
78.1
Wrong answers present in passage
34.0%
45.1%
34.7%
Attention flip precision
.809
.672
.607
answer in passage, but wrong
.111
.000
.000
Appendix
Table 11: The attention family against retrieval quality, on the with-context branch [ ] of Qwen3.5-9B. The first three rows characterise the data: the generator’s answer accuracy without and with the retrieved passage, and the share of wrong answers whose text nevertheless appears verbatim in the passage (normalised substring match). The last three rows are the decision-flip precision of Table 10 , computed on all examples and then restricted to the two subsets on which verbatim support and correctness disagree.
Single-hop
Multi-hop
Method
PopQA
NQ
TQA
HotpotQA
2Wiki
MuSiQue
All
Random AUPRC (error rate)
.504
.396
.158
.492
.706
.882
.523
BEATRICE
.866 / .868
.836 / .781
.899/ .696
.864 / .866
.832 / .909
.798/.964
.891 / .894
SAPLMA (mean)
.813/.809
.770/.644
.910 /.643
.826/.780
.757/.865
.811 / .969
.852/.836
SAPLMA (last)
.811/.799
.778/.653
.844/.566
.790/.774
.775/.892
.766/.958
.836/.827
LLMsKnow
.800/.789
.756/.656
.815/.562
.795/.798
.718/.843
.713/.946
.812/.813
Appendix
Table 12: Real-retrieval detection on Qwen3.5-9B : AUROC / AUPRC on each dataset under real retrieved context (E5 over Wikipedia-2018, k=3 ), 500 questions each. AUPRC positive class is the incorrect answer; the italic row is the random-guess AUPRC (per-column error rate). Bold marks the best in each column per metric. Every trained detector reuses its with-context weights and is evaluated zero-shot.
Single-hop
Multi-hop
Method
PopQA
NQ
TQA
HotpotQA
2Wiki
MuSiQue
All
Random AUPRC (error rate)
.562
.440
.266
.680
.882
.930
.627
BEATRICE
.877 / .912
.819 / .802
.883 / .805
.922 / .963
.943 / .992
.892 / .991
.911 / .950
LLMsKnow
.765/.845
.696/.690
.749/.672
.850/.925
.847/.977
.841/.985
.847/.916
SAPLMA (mean)
.782/.823
.709/.673
.739/.488
.824/.903
.800/.963
.727/.970
.806/.864
SAPLMA (last)
.781/.842
.665/.620
.699/.478
.830/.911
.791/.965
.799/.981
.799/.872
Appendix
Table 13: Real-retrieval detection on Gemma 4 , conventions as in Table 12 . On this backbone the sampling detectors fall to or below chance (the consistent-refusal effect of Appendix K ).
In-domain (multi-hop)
Out-of-distribution
Macro
Method
HotpotQA
2WikiMQA
MuSiQue
PopQA
NQ
TriviaQA
EM
F1
Gen
No-RAG
.220/.293
.238/.282
.034/.114
.192/.246
.176/.267
.574/.617
.239
.303
1.00
Vanilla RAG
.322/.437
.260/.309
.060/.133
.447 / .526
.371 / .485
.757 / .818
.370
.451
1.00
SKR
.310/.416
.296/.342
.052/.126
.440/.519
.364/.474
.738/.796
.367
.445
1.00
Adaptive-RAG
.322/.431
.282/.323
.060/.098
.433/.496
.319/.425
.718/.769
.356
.424
1.63
DRAGIN
.300/.394
.298/.343
.062/.151
.415/.473
.354/.467
.672/.724
.350
.425
3.55
Appendix
Table 14: PGIR as a downstream controller, against the confidence-control family: systems that gate or verify retrieval with a self-assessed signal, all sharing one backbone and retriever. Cells are exact-match / F1. Macro is the six-set mean and Gen is generations per query. Per column and metric, bold = best and underline = second.
Question
Closed-book answer (score)
With-context answer (score)
Gold
PGIR action
Qwen3.5-9B
The Nun is based on what film series?
The Conjuring (.67)
not retrieved ()
The Conjuring
Accept, no retrieval
What gemstone is The Moonstone in the novel by Wilkie Collins?
Moonstone (.57)
Diamond (.90)
Diamond
Retrieve, then trust
Arctic King, Saladin, and Tom Thumb are which vegetable?
Onions (.26)
Cauliflower (.01)
Lettuce
Retrieve; both flagged, none trusted
Gemma 4
Who played Clayton Farlowe in Dallas ?
Michael Landon (.05)
Howard Keel (.65)
Howard Keel
Retrieve, then trust
Appendix
Table 15: Case studies from the PGIR runs on both backbones (E5 over Wikipedia). Each score is the calibrated P(answer correct) the controller acts on; the acceptance threshold is τ=0.6 . A high closed-book score is accepted without retrieval; a low one triggers retrieval, after which a draft clearing τ is trusted even when it overturns the closed-book guess, while a pool that stays low is left flagged and the best candidate is returned without claiming it is right.
Figure 12: Exhaustive ranking of all three-encoder combinations on Gemma 4, best to worst, in the format of Figure 5 . All-closed-book combinations sit at the bottom in the with-context pool; the greedy-selected combination is marked by the dashed line.
Closed-book
With context
ID
OOD
ID
OOD
Method
Test
PopQA
NQ
TQA
Test
PopQA
NQ
TQA
Sampling-based (10 extra generations per query)
Semantic Entropy
.339/.894
.523/.888
.525/.746
.756/.717
.595/.404
.456/.648
.384/.726
.466/.428
Bayesian SE
.323/.891
.524/.888
.513/.737
.749/.713
.591/.403
.449/.647
.367/.720
.451/.425
SelfCheckGPT-NLI
.199/.880
.330/.883
.485/.777
.693/.729
.390/.419
.087/.628
.169/.740
.185/.448
Appendix
Table 16: Detecting incorrect answers of Gemma 4 E4B (instruction-tuned) in the closed-book and retrieval-augmented settings, mirroring Table 1 : each cell reports AUROC / AUPRC with “the answer is wrong” as the AUPRC positive class. The entire pipeline is re-run on the new base model under the protocol of Table 1 : same data scale and splits, consensus-purified labels regenerated from this model’s own greedy answers, all layers and hyperparameters of every method re-selected on the validation split only, and the with-context evaluation again merging original and rewritten contexts. The selected probes are ( )( )( ) (closed-book) and [ ][ ][ ] (with context). Sampling-based scores fall to or below chance on this base model; see the discussion in the text. Per column and metric, bold = best and underline = second; “–” marks detectors that require retrieved context and are undefined closed-book.
In-domain (multi-hop)
Out-of-distribution
Macro
Method
HotpotQA
2WikiMQA
MuSiQue
PopQA
NQ
TriviaQA
EM
F1
Gen
No-RAG
.050/.069
.006/.012
.000/.003
.048/.052
.030/.052
.298/.340
.072
.088
1.00
Vanilla RAG
.190/.298
.086/.177
.030/.078
.375/ .463
.284/.396
.602/.692
.261
.351
1.00
SKR
.190/.297
.112/.196
.028/.074
.368/.456
.280/.392
.602/.690
.263
.351
1.00
Adaptive-RAG
.222/.329
.104/.194
.040 / .097
.364/.431
.307 / .415
.625/.700
.277
.361
1.70
DRAGIN
.164/.237
.154/.201
.020/.062
.228/.284
.225/.323
.514/.584
.218
.282
2.64
Appendix
Table 17: PGIR on the Gemma backbone, mirroring Table 14 : all systems share the same backbone (Gemma 4 E4B instruction-tuned, greedy decoding) and the same retriever (E5 over Wikipedia-2018, k=3 ), evaluated on the identical splits as the Qwen table. Cells are exact-match / F1; Macro is the six-set mean and Gen is generations per query. The trained systems (SKR, Adaptive-RAG, Probing-RAG, PGIR) are retrained for this backbone, and PGIR’s acceptance threshold and budget are re-tuned on the dev split only. The confidence-control family ranking is stable across backbones: PGIR leads macro EM and F1 with Probing-RAG second on both, and PGIR again trades individual sets to specialists (SeaKR on HotpotQA at 17.5 generations per query, Probing-RAG on TriviaQA). Per column and metric, bold = best and underline = second.
Test-time reasoning has become a significant field of study since the introduction of chain-of-thought reasoning in large language models (LLMs). However, the mechanisms of this reasoning process are still under-explored -- from the same input prompt, and even the same partial solution, LLMs can produce varied answers if sampled multiple times. We propose to leverage question-asking as an inference-time intervention that articulates information about the model's hidden state. To achieve that, we present a student-teacher setting where a student asks questions to a teacher. We train a probe on the student's hidden state before and after asking a question and find it is predictive of the trajectory's final correctness, even before generating the teacher's answer. This suggests there is a meaningful signal from the self-diagnosis that occurs during question generation rather than information transfer from the teacher. We then frame question-asking as a sequential decision problem, using this probe as a quality score, and define a gating policy to ask questions that maximize likelihood of correctness. We find that the success of question-asking as an intervention is largely dependent on the model's self-consistency. Our empirical results show a gap between detection and recovery; while our gating policy captures model correctness and uncertainty, interventions are equally likely to harm correct trajectories as they are to recover incorrect ones. This gap between diagnosis and correction has broader implications on language models' capacity for self-refinement under uncertainty.
Chu Fei Luo, Samuel Dahan, Xiaodan Zhu
Department of Electrical and Computer Engineering & Ingenuity Labs, Queen’s University · Conflict Analytics Lab, Queen’s University · Cornell Law School +1
Large language models encode rich information in their hidden states. This work asks whether the correctness of code that Qwen3-4B-Instruct-2507 has not yet generated is already legible in its hidden states, evaluated on a set of 444 tasks from LiveCodeBench. The correctness of the model's first-attempt code is linearly decodable from the hidden state at the final prompt token, captured before any output token is generated, with a leakage-free held-out AUC of 0.881 +/- 0.008 across 50 outer splits. To assess whether this signal is explained by prompt length, each hidden state dimension is residualized with respect to its linear effect. The probe still achieves an AUC of 0.842 +/- 0.010, substantially above a logistic prompt-length baseline of 0.657 +/- 0.014, and none of the nonlinear models tested improves upon it. A companion question about whether self-repair leaves a geometric signature in the model's hidden states could not be answered, because successful repairs following a failed first attempt are too rare in this setting to support the analysis. The contribution is both empirical and methodological, providing evidence that pre-generation hidden states contain a robust signal of eventual code correctness, together with a confound-control diagnostic that quantifies how much of that signal survives adjustment for prompt length.
Code generated by modern language models often reads naturally. Yet, it also often fails to implement what was asked. This should be no surprise, as research shows the models' own confidence signals are poorly calibrated with actual correctness. A promising way to assess correctness looks inside the model: by contrasting the hidden states of correct and incorrect programs, recent work captured an internal signal of code correctness that is able to judge candidate solutions better than the model's token-level or stated confidence, with no test execution. However, this signal was captured under one particular way, leaving open an important question: whether it reflects a robust property of the model or an artifact of that choice. We study this question systematically, varying how the signal is extracted from the model internals. Besides this, we also ask if the signal's quality is limited by the data used to extract it, by constructing program pairs that differ only in the fault that makes them incorrect. Our results show that no single configuration is best, and that isolating the fault does not help.
Francisco Ribeiro, Sohaila Abdulsattar, Renata Gonzalez +2
New York University Abu Dhabi Abu Dhabi, United Arab Emirates