Large language models routinely answer questions that cannot be answered from the information given, and in dialogue they answer before enough has been said. Unanswerability is linearly decodable from hidden states, but it is unclear which of its forms share a representation and whether the signal is useful in dialogue. We contribute a turn-labeled multi-turn benchmark (423 conversations, 1,661 labeled turn-states) and an evaluation harness with a simulated user who answers clarifying questions, and use them with six datasets and six open-weight LLMs to test how far probes for unanswerability carry. Probes transfer robustly between datasets that share a ground of unanswerability: missing information in math (AUROC 0.77-0.97) and in a passage (SQuAD 2.0<->MuSiQue, 0.77-0.90). Probes for epistemic "known-unknowns" transfer poorly to math, but this separation weakens under lexical controls and changes with layer and coordinate system, so it remains unresolved. Single-turn probes fail zero-shot to detect when a conversation becomes answerable; in-structure probes recover it, but no better than a bag-of-words classifier. A gate on the calibrated probe, with no model fine-tuning, fires on underspecified turns far more precisely than chance, and its end-task success comes within 0.08 of a gate given the true labels. Yet across four models it does not reliably beat vanilla generation or prompted consolidation. The remaining gap lies mostly in how models use a clarification, not in detection.
Figures & tables
four datasets (common layer)
six datasets
model
diag
cross [95% CI]
drop [95% CI]
SQ → MQ
MQ → SQ
Qwen3.5-9B
0.948
0.732 [.721,.743]
0.216 [.195,.236]
0.88
0.81
Qwen3.6-27B
0.979
0.834 [.825,.843]
0.145 [.129,.159]
0.88
0.77
Gemma4-12B
0.962
0.808 [.799,.818]
0.153 [.138,.169]
0.89
0.83
Gemma4-31B
0.976
0.757 [.747,.767]
0.219 [.202,.235]
0.90
0.90
Ministral3-8B
0.939
0.727 [.715,.738]
0.213 [.191,.233]
0.84
0.86
Table 1: RQ1 transfer (AUROC, mean over seeds {13,42,77} ). Left: the four-core-dataset matrix at the common best layer. Diag is within-dataset; cross averages all off-diagonal cells (including SUM ↔ UMWP); drop = diag − cross. CIs come from a seed and sampling bootstrap. Right: extractive-pair transfer between SQuAD 2.0 (SQ) and MuSiQue (MQ) at the six-dataset common layer. Per-model LODO and source-only-layer results are in the appendix.
Figure 1: Train × test transfer matrices on the four core datasets for the four 8–12B models (AUROC, mean over seeds). All six models on all six datasets are in Appendix Figure 3 .
SA → SUM
KUQ → SUM
SA ↔ KUQ /
D (deg)
model
raw / res.
L 6 / L 4
SA → math, res.
Δμ
whitened
Qwen3.5-9B
.47 / .61
.52 / .68
.66 / .64
+5.2∗
−2.1∗
Qwen3.6-27B
.55 / .77
.65 / .65
.81 / .79
+8.2∗
−0.4
Gemma4-12B
.51 / .80
.74 / .75
.72 / .86
+5.8∗
−1.8∗
Gemma4-31B
.52 / .79
.45 / .45
.78 / .86
+6.2∗
−2.4∗
Ministral3-8B
.57 / .63
.60 / .60
.72 / .70
+15.9∗
+0.6
Table 2: Robustness of the known-unknowns separation (SA = SelfAware). Transfer is AUROC; res. means after removing the TF-IDF-predictable component of the hidden state. KUQ → SUM is shown at the six-dataset (L 6 ) and four-dataset (L 4 ) common layers. The fourth column compares residualized SA ↔ KUQ transfer (mean of both directions) with residualized SA → {SUM, UMWP}; underlined where known-unknowns transfers no better within its ground than to math. D is the angle between the {math, extractive} subspace and the known-unknowns direction minus the math–extractive angle; ∗ : 95% bootstrap CI excludes 0. Bold: above the 0.65 threshold.
model
zero-shot
ceiling
within-turn
∠ to ST
qwen3.5-9b
0.784
0.927
0.76–0.84
73.9 ∘
ministral-3-8b
0.774
0.935
0.79–0.86
56.7 ∘
granite-4.1-8b
0.789
0.925
0.76–0.84
81.2 ∘
gemma-4-12b
0.845
0.934
0.79–0.86
58.7 ∘
Table 3: RQ2 turn-level AUROC. Zero-shot: best layer of the frozen single-turn probe (mid/late layers are at or below chance, 0.20–0.55). Ceiling: probe trained on turn-states with conversation-grouped CV. Within-turn: ceiling restricted to fixed turn indices 2–4, min–max over turns (absolute position is fixed; relative position still varies with conversation length). ∠ to ST: principal angle of the multi-turn Δμ to the whole single-turn subspace.
condition
Qwen3.5
Granite
Gemma4-12B
Ministral
multi-turn success, answering user
vanilla
0.764
0.598
0.506
0.811
ask-when-unsure
0.759
0.811
0.459
0.766
consolidation (RECAP)
0.837
0.825
0.631
0.856
probe gate ( τ=0.5 )
0.693
0.681
0.584
0.853
oracle gate (true labels)
0.721
0.681
0.660
0.861
Table 4: RQ3 on N-matched items (423 conversations; 2,368 single-turn items). Multi-turn success is objective per source: numeric match, differential code execution, or slot coverage. The user simulator answers clarifying questions from the not-yet-revealed constraints. † The position gate fires while any scripted constraint is unrevealed; it needs the script length and is not deployable. Firing precision is the share of fired turns that are underspecified, and recall the share of underspecified turns on which the gate fires; random precision is the mean over three seeds (oracle: 1/1 by construction). The scripted-user results and the cost-sensitive gate are in the appendix.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
rule
commit
committed (UTC)
first result (UTC)
RQ1 dissociation rule
3600fc0
07-11 21:32
07-11 21:44
LODO pooled probe
c577bc9
07-12 11:34
07-12 11:35
Direction geometry
462b8dd
07-12 16:12
07-12 16:33
Benchmark repair protocol
bc2049f
07-17 14:23
07-17 14:25
RQ3 endpoint + metrics
d80bc2c
07-20 09:49
07-20 10:20
Robustness datasets, text baselines
1bd8f6d
09-25 07:07
09-25 07:07
Appendix
Table 5: Rule commits and first outputs (all 2026, UTC). Rows below the line cover the robustness analyses and the answering-user evaluation. The two-sided headline (unified mechanism vs dissociation) is fixed in the project specification committed on 07-10. The answering-user scoring commit restates the cf5450f endpoint and makes two implicit details explicit: Holm correction across models, and the per-model best prompted baseline. Before that commit, exactly one generation output (Granite vanilla) had been inspected, as a harness sanity check. The rule for the second known-unknowns set was amended by the authors before its run, to read SelfAware ↔ KUQ transfer only after lexical residualization.
model
τ
correct
abst unans
false-abst
acc ans
qwen
0.1
0.575
0.888
0.726
0.262
qwen
0.3
0.582
0.869
0.676
0.294
qwen
0.5
0.589
0.861
0.643
0.317
qwen
0.7
0.600
0.849
0.585
0.351
qwen
0.9
0.603
0.781
0.475
0.424
granite
0.1
0.548
0.841
0.730
0.254
Appendix
Table 6: Single-turn τ sweep (pooled, n=2368 ): correct behavior, abstention on unanswerable, false abstention on answerable, accuracy on answerable. Full grid (nine values of τ ) in results/rq3_tau_single.csv .
t
user turn (source)
label
p
gate
model reply (excerpt)
1
“…I made 6 trays of them.” (script)
1
0.995
ask
“…I need to know how many muffins are on each tray. How many muffins are on each tray?”
2
“Each tray holds 12 muffins.” (answer)
1
0.793
ask
“…the most important piece of information missing is the selling price per muffin …How much are you planning to sell each muffin for?”
“Before selling any, I set aside one quarter …to donate …” (script)
0
0.403
answer
“… 41 of 72 = 18 muffins …the amount you will actually raise …is $108.”
Appendix
Table 7: Worked example (qwen3.5-9b, probe gate, answering user). Source: script is the next scripted constraint; answer is the simulated user’s reply to the model’s question. Label: 1 = underspecified (sufficiency set not yet revealed). p : Platt-calibrated gate probability.
Figure 2: Train × test transfer matrices at each model’s common best layer (AUROC, mean over seeds {13,42,77} ) for all six models; the color scale is clipped at 0.5. Rows are the training dataset, columns the test dataset. Generated from results/transfer_bestlayer.csv by scripts/plot_transfer_heatmap.py .
LODO held-out
source-only layers
model
math
extr
k-unk
cross / drop
verdict
Qwen3.5-9B
0.649
0.878
0.888
0.680 / 0.272
dissoc.
Qwen3.6-27B
0.796
0.936
0.850
0.766 / 0.213
dissoc. (was mixed)
Gemma4-12B
0.731
0.920
0.856
0.720 / 0.244
dissoc. (was mixed)
Gemma4-31B
0.780
0.939
0.812
0.670 / 0.307
dissoc.
Ministral3-8B
0.671
0.869
0.680
0.682 / 0.261
dissoc.
Appendix
Table 8: Left: LODO, a descriptive upper bound reported at the best held-out-domain layer. Right: transfer when each training dataset’s layer is chosen from that dataset alone. Selecting the layer on seed 13 and evaluating on seeds 42 and 77 gives the same 6/6.
Qwen3.5
Granite
Gemma4-12B
Ministral
vanilla, scripted user
0.759
0.589
–
–
probe gate, scripted user
0.596
0.492
–
–
vanilla, answering user
0.764
0.598
0.506
0.811
probe gate, answering user
0.693
0.681
0.584
0.853
cost gate, answering user
0.669
0.690
0.407
0.856
user turns, vanilla → probe
4.3 → 4.8
4.0 → 4.9
4.1 → 4.9
4.2 → 5.3
Appendix
Table 9: Multi-turn success under the scripted user (two models) and the answering user. The cost-sensitive gate does not improve on τ=0.5 . The last row is the share of probe-gated conversations in which the model asked at least one question that the hidden constraints could not answer.
model
Δ ceiling
Δ zero-shot
Qwen3.5-9B
+0.069±0.010
+0.071
Ministral3-8B
+0.101±0.007
+0.066
Granite4.1-8B
+0.025±0.009
−0.050
Gemma4-12B
−0.005±0.025
−0.153
Appendix
Table 10: Δ = AUROC(clean context) − AUROC(context with a truncated reply), within turns 2–4; mean ± sd over seeds. The zero-shot probe is seed-deterministic.
Figure 3: Train × test transfer on all six datasets at each model’s six-dataset common best layer (AUROC, mean over seeds {13,42,77} ; color clipped at 0.5). Generated from results/rq1_v2/transfer_bestlayer.csv by scripts/plot_transfer_heatmap.py .