Large language models routinely answer questions that cannot be answered from the information given, and in dialogue they answer before enough has been said. Unanswerability is linearly decodable from hidden states, but it is unclear which of its forms share a representation and whether the signal is useful in dialogue. We contribute a turn-labeled multi-turn benchmark (423 conversations, 1,661 labeled turn-states) and an evaluation harness with a simulated user who answers clarifying questions, and use them with six datasets and six open-weight LLMs to test how far probes for unanswerability carry. Probes transfer robustly between datasets that share a ground of unanswerability: missing information in math (AUROC 0.77-0.97) and in a passage (SQuAD 2.0<->MuSiQue, 0.77-0.90). Probes for epistemic "known-unknowns" transfer poorly to math, but this separation weakens under lexical controls and changes with layer and coordinate system, so it remains unresolved. Single-turn probes fail zero-shot to detect when a conversation becomes answerable; in-structure probes recover it, but no better than a bag-of-words classifier. A gate on the calibrated probe, with no model fine-tuning, fires on underspecified turns far more precisely than chance, and its end-task success comes within 0.08 of a gate given the true labels. Yet across four models it does not reliably beat vanilla generation or prompted consolidation. The remaining gap lies mostly in how models use a clarification, not in detection.
Figures & tables
four datasets (common layer)
six datasets
model
diag
cross [95% CI]
drop [95% CI]
SQ → MQ
MQ → SQ
Qwen3.5-9B
0.948
0.732 [.721,.743]
0.216 [.195,.236]
0.88
0.81
Qwen3.6-27B
0.979
0.834 [.825,.843]
0.145 [.129,.159]
0.88
0.77
Gemma4-12B
0.962
0.808 [.799,.818]
0.153 [.138,.169]
0.89
0.83
Gemma4-31B
0.976
0.757 [.747,.767]
0.219 [.202,.235]
0.90
0.90
Ministral3-8B
0.939
0.727 [.715,.738]
0.213 [.191,.233]
0.84
0.86
Table 1: RQ1 transfer (AUROC, mean over seeds {13,42,77} ). Left: the four-core-dataset matrix at the common best layer. Diag is within-dataset; cross averages all off-diagonal cells (including SUM ↔ UMWP); drop = diag − cross. CIs come from a seed and sampling bootstrap. Right: extractive-pair transfer between SQuAD 2.0 (SQ) and MuSiQue (MQ) at the six-dataset common layer. Per-model LODO and source-only-layer results are in the appendix.
Figure 1: Train × test transfer matrices on the four core datasets for the four 8–12B models (AUROC, mean over seeds). All six models on all six datasets are in Appendix Figure 3 .
SA → SUM
KUQ → SUM
SA ↔ KUQ /
D (deg)
model
raw / res.
L 6 / L 4
SA → math, res.
Δμ
whitened
Qwen3.5-9B
.47 / .61
.52 / .68
.66 / .64
+5.2∗
−2.1∗
Qwen3.6-27B
.55 / .77
.65 / .65
.81 / .79
+8.2∗
−0.4
Gemma4-12B
.51 / .80
.74 / .75
.72 / .86
+5.8∗
−1.8∗
Gemma4-31B
.52 / .79
.45 / .45
.78 / .86
+6.2∗
−2.4∗
Ministral3-8B
.57 / .63
.60 / .60
.72 / .70
+15.9∗
+0.6
Table 2: Robustness of the known-unknowns separation (SA = SelfAware). Transfer is AUROC; res. means after removing the TF-IDF-predictable component of the hidden state. KUQ → SUM is shown at the six-dataset (L 6 ) and four-dataset (L 4 ) common layers. The fourth column compares residualized SA ↔ KUQ transfer (mean of both directions) with residualized SA → {SUM, UMWP}; underlined where known-unknowns transfers no better within its ground than to math. D is the angle between the {math, extractive} subspace and the known-unknowns direction minus the math–extractive angle; ∗ : 95% bootstrap CI excludes 0. Bold: above the 0.65 threshold.
model
zero-shot
ceiling
within-turn
∠ to ST
qwen3.5-9b
0.784
0.927
0.76–0.84
73.9 ∘
ministral-3-8b
0.774
0.935
0.79–0.86
56.7 ∘
granite-4.1-8b
0.789
0.925
0.76–0.84
81.2 ∘
gemma-4-12b
0.845
0.934
0.79–0.86
58.7 ∘
Table 3: RQ2 turn-level AUROC. Zero-shot: best layer of the frozen single-turn probe (mid/late layers are at or below chance, 0.20–0.55). Ceiling: probe trained on turn-states with conversation-grouped CV. Within-turn: ceiling restricted to fixed turn indices 2–4, min–max over turns (absolute position is fixed; relative position still varies with conversation length). ∠ to ST: principal angle of the multi-turn Δμ to the whole single-turn subspace.
condition
Qwen3.5
Granite
Gemma4-12B
Ministral
multi-turn success, answering user
vanilla
0.764
0.598
0.506
0.811
ask-when-unsure
0.759
0.811
0.459
0.766
consolidation (RECAP)
0.837
0.825
0.631
0.856
probe gate ( τ=0.5 )
0.693
0.681
0.584
0.853
oracle gate (true labels)
0.721
0.681
0.660
0.861
Table 4: RQ3 on N-matched items (423 conversations; 2,368 single-turn items). Multi-turn success is objective per source: numeric match, differential code execution, or slot coverage. The user simulator answers clarifying questions from the not-yet-revealed constraints. † The position gate fires while any scripted constraint is unrevealed; it needs the script length and is not deployable. Firing precision is the share of fired turns that are underspecified, and recall the share of underspecified turns on which the gate fires; random precision is the mean over three seeds (oracle: 1/1 by construction). The scripted-user results and the cost-sensitive gate are in the appendix.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
rule
commit
committed (UTC)
first result (UTC)
RQ1 dissociation rule
3600fc0
07-11 21:32
07-11 21:44
LODO pooled probe
c577bc9
07-12 11:34
07-12 11:35
Direction geometry
462b8dd
07-12 16:12
07-12 16:33
Benchmark repair protocol
bc2049f
07-17 14:23
07-17 14:25
RQ3 endpoint + metrics
d80bc2c
07-20 09:49
07-20 10:20
Robustness datasets, text baselines
1bd8f6d
09-25 07:07
09-25 07:07
Appendix
Table 5: Rule commits and first outputs (all 2026, UTC). Rows below the line cover the robustness analyses and the answering-user evaluation. The two-sided headline (unified mechanism vs dissociation) is fixed in the project specification committed on 07-10. The answering-user scoring commit restates the cf5450f endpoint and makes two implicit details explicit: Holm correction across models, and the per-model best prompted baseline. Before that commit, exactly one generation output (Granite vanilla) had been inspected, as a harness sanity check. The rule for the second known-unknowns set was amended by the authors before its run, to read SelfAware ↔ KUQ transfer only after lexical residualization.
model
τ
correct
abst unans
false-abst
acc ans
qwen
0.1
0.575
0.888
0.726
0.262
qwen
0.3
0.582
0.869
0.676
0.294
qwen
0.5
0.589
0.861
0.643
0.317
qwen
0.7
0.600
0.849
0.585
0.351
qwen
0.9
0.603
0.781
0.475
0.424
granite
0.1
0.548
0.841
0.730
0.254
Appendix
Table 6: Single-turn τ sweep (pooled, n=2368 ): correct behavior, abstention on unanswerable, false abstention on answerable, accuracy on answerable. Full grid (nine values of τ ) in results/rq3_tau_single.csv .
t
user turn (source)
label
p
gate
model reply (excerpt)
1
“…I made 6 trays of them.” (script)
1
0.995
ask
“…I need to know how many muffins are on each tray. How many muffins are on each tray?”
2
“Each tray holds 12 muffins.” (answer)
1
0.793
ask
“…the most important piece of information missing is the selling price per muffin …How much are you planning to sell each muffin for?”
“Before selling any, I set aside one quarter …to donate …” (script)
0
0.403
answer
“… 41 of 72 = 18 muffins …the amount you will actually raise …is $108.”
Appendix
Table 7: Worked example (qwen3.5-9b, probe gate, answering user). Source: script is the next scripted constraint; answer is the simulated user’s reply to the model’s question. Label: 1 = underspecified (sufficiency set not yet revealed). p : Platt-calibrated gate probability.
Figure 2: Train × test transfer matrices at each model’s common best layer (AUROC, mean over seeds {13,42,77} ) for all six models; the color scale is clipped at 0.5. Rows are the training dataset, columns the test dataset. Generated from results/transfer_bestlayer.csv by scripts/plot_transfer_heatmap.py .
LODO held-out
source-only layers
model
math
extr
k-unk
cross / drop
verdict
Qwen3.5-9B
0.649
0.878
0.888
0.680 / 0.272
dissoc.
Qwen3.6-27B
0.796
0.936
0.850
0.766 / 0.213
dissoc. (was mixed)
Gemma4-12B
0.731
0.920
0.856
0.720 / 0.244
dissoc. (was mixed)
Gemma4-31B
0.780
0.939
0.812
0.670 / 0.307
dissoc.
Ministral3-8B
0.671
0.869
0.680
0.682 / 0.261
dissoc.
Appendix
Table 8: Left: LODO, a descriptive upper bound reported at the best held-out-domain layer. Right: transfer when each training dataset’s layer is chosen from that dataset alone. Selecting the layer on seed 13 and evaluating on seeds 42 and 77 gives the same 6/6.
Qwen3.5
Granite
Gemma4-12B
Ministral
vanilla, scripted user
0.759
0.589
–
–
probe gate, scripted user
0.596
0.492
–
–
vanilla, answering user
0.764
0.598
0.506
0.811
probe gate, answering user
0.693
0.681
0.584
0.853
cost gate, answering user
0.669
0.690
0.407
0.856
user turns, vanilla → probe
4.3 → 4.8
4.0 → 4.9
4.1 → 4.9
4.2 → 5.3
Appendix
Table 9: Multi-turn success under the scripted user (two models) and the answering user. The cost-sensitive gate does not improve on τ=0.5 . The last row is the share of probe-gated conversations in which the model asked at least one question that the hidden constraints could not answer.
model
Δ ceiling
Δ zero-shot
Qwen3.5-9B
+0.069±0.010
+0.071
Ministral3-8B
+0.101±0.007
+0.066
Granite4.1-8B
+0.025±0.009
−0.050
Gemma4-12B
−0.005±0.025
−0.153
Appendix
Table 10: Δ = AUROC(clean context) − AUROC(context with a truncated reply), within turns 2–4; mean ± sd over seeds. The zero-shot probe is seed-deterministic.
Figure 3: Train × test transfer on all six datasets at each model’s six-dataset common best layer (AUROC, mean over seeds {13,42,77} ; color clipped at 0.5). Generated from results/rq1_v2/transfer_bestlayer.csv by scripts/plot_transfer_heatmap.py .
A model should refuse two different things: answers it would get wrong, and questions it should not answer at all, such as unanswerable ones or ones resting on a false premise. The usual recipe thresholds a single confidence score, which cannot tell these apart. Across five instruction-tuned models from three families (2B to 14B), we find they are separate axes. Ordinary answer-confidence tracks whether an answer is right but is nearly blind to whether the question is answerable; a linear probe on hidden states does the reverse. The blind spot does not shrink with scale. It is worst on naturally occurring false-premise questions (CREPE). There, answer-confidence, P(IK), P(True), and even asking the model outright whether a premise is false all stay near chance, while a hidden-state probe reaches 0.69 to 0.77 AUROC: the model represents a problem it will not report. This turns out to be fixable. Instructing a model to check premises backfires, because it then disputes sound and false premises alike (57% false challenges), unable to tell them apart; routing the same instruction with the probe roughly triples challenge precision. We turn the two axes into a calibrated policy that answers only when an answerability score and a correctness score each clear a separately certifies behave differently: the unanswerable-answer rate is controllable at every scale, while the wrong-answer rate is capped by model accuracy, so the guarantee tightens as threshold policy certifies both budgets at 0.75 coverage of correct answers, against 0.31 for a single threshold; at 14B it is the only policy that certifies at all.
Reliable evaluation of large language models should separate supported answering from unsupported guessing without conflating either with data contamination, prompt idiosyncrasy, or generic refusal behavior. We present a contamination-aware, multi-zone benchmark for measuring the transition from answerable knowledge to abstention-expected unknowns under frozen build-time labels. The benchmark contains 1,200 items across five domains, explicit abstention expectations, contamination-risk metadata, and dual parsing with an official strict parser plus a normalized robustness parser. We evaluate FLAN-T5, Qwen2.5-Instruct, and Llama-3-Instruct models under locked answer-or-abstain prompts, answer-only controls, and prompt-template variants. The benchmark is not solved by generic non-answer behavior: FLAN baselines remain weak on productive abstention, while stronger instruction-tuned models expose a selective but incomplete transition from answering to abstaining. Qwen2.5-3B-Instruct achieves the best overall reliability, but answer-expected zones remain difficult, calibration remains poor, and benign-item refusal persists. Prompt and parser robustness analyses preserve the main ranking and qualitative conclusions. The benchmark therefore provides a reproducible protocol for auditing answerability, abstention, refusal, and contamination as distinct but interacting dimensions of LLM reliability.The dataset is publicly available at https://github.com/renweimeng/Know2Guess-A-Contamination-Aware-Multi-Zone-Benchmark.
User queries are often underspecified and may admit multiple valid interpretations. Rather than silently making assumptions about the user's intent, a helpful assistant should surface such ambiguity by asking a clarifying question. Doing so requires two abilities: recognizing that a query is ambiguous, and acting on that recognition by seeking clarification instead of answering directly. To study these abilities, we evaluate models on ambiguous, unambiguous, and disambiguated questions in three settings: standard question answering, explicit ambiguity judgment, and behavioral analysis, where a judge model classifies responses as direct answers, refusals, or clarifying questions. We find a clear gap between recognition and behavior: models often identify ambiguity when explicitly asked to judge it, yet in the QA setting they overwhelmingly default to direct answers. Retrieved context further widens this gap by improving answerability while making models even less likely to ask clarifying questions.