A language model can fail a syntactic test in two distinct ways: by not encoding the relevant structure, or by encoding it but failing to use it at the output. Behavioral evaluation alone cannot tell these apart. We propose a three-level evaluation framework (behavioral deployment, LM-head readout, and probe recoverability) measured on the same items under the same binary decision. Using a compact trilingual (English, Chinese, German) control-dependency benchmark, we find that probe recoverability exceeds or equals LM-head readout, which in turn exceeds or equals behavioral deployment, across seven models and all three languages in the aggregate. The recoverability surplus is never negative across all 14 (model, task) conditions. The disconnect concentrates in subject-control, where a nearest-noun heuristic gives the wrong answer. The single largest gap (0.653) appears on Qwen3-0.6B Instruct in question answering. The gap persists at Qwen3-14B Instruct. Instruction tuning degrades deployment more than encoding in percentage terms. We rule out option-position bias, late-layer erasure, output-formatting artifacts, and probe-training variance. The pattern is consistent with decoding that favors surface shortcuts, and the behavior-probe gap measures the strength of that preference. Activation patching shows the gap is layer-localized. Under instruction tuning, the LM-head-decoded layer shifts approximately ten layers later than the probe-decoded layer. These findings argue that behavioral evaluation understates what models encode, while probing alone overstates what they deploy.
Figures & tables
Model
QA
Paraphrase
Mean Behavior
Best LM-head
Best Probe
Surplus
B/R Ratio
Qwen3-0.6B Base
0.542
0.646
0.594
0.708
0.938
0.344
0.633
Qwen3-0.6B Instruct
0.417
0.604
0.510
0.667
0.833
0.323
0.612
Qwen3-14B Instruct
0.812
0.771
0.792
0.792
1.000
0.208
0.792
Llama-3.2-1B Base
0.521
0.625
0.573
0.688
0.812
0.240
0.705
Llama-3.2-1B Instruct
0.479
0.542
0.510
0.583
0.625
0.115
0.817
Gemma-4 E4B
0.521
0.604
0.562
0.604
0.604
0.042
0.931
Table 1: Main trilingual model comparison across behavioral deployment, LM-head readout, and probe recoverability.
Figure 1: Three-level gap across all seven models (right: recoverability surplus).
Figure 2: Three-level framework on the hardest case (Qwen3-0.6B Instruct, QA, subject-control), where behavioral deployment (0.250) falls far below LM-head readout (0.667), which falls far below probe recoverability (0.903).
Figure 3: Probe accuracy by relative layer depth across Qwen3 models.
Figure 4: Three-level separation for Qwen3-0.6B base and instruct, by task, on the same items.
Figure 5: Subject-control deployment gap (probe − behavior) across scale.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
QA Format
Paraphrase Format
Sentence: “John promised Mary to leave early.”
Sentence: “John promised Mary to leave early.”
Prompt: Q: Who is understood to leave early? A: John ← log-prob scored A: Mary ← log-prob scored
Prompt: Which sentence better describes the situation? (A) John is understood to leave early. ← scored (B) Mary is understood to leave early. ← scored
Gold controller: John ( promise = subject-control)
Appendix
Table 2: The two behavioral evaluation formats, both using pairwise log-probability scoring over the same candidate controller nouns and differing only in prompt structure.
Mode
Qwen3-0.6B Base
Qwen3-0.6B Instruct
Qwen3- 14B Instruct
ab
0.924
0.847
0.958
bv
0.889
0.792
0.979
av
0.618
0.458
0.931
abv
0.847
0.701
0.944
v
0.757
0.521
0.986
Appendix
Table 3: Best per-mode probe accuracy across layers, mean over three seeds.
Model
Seed 1729
Seed 2718
Seed 3141
Mean
Std
Qwen3-0.6B Base
0.938
0.938
0.896
0.924
0.020
Qwen3-0.6B Instruct
0.833
0.812
0.896
0.847
0.035
Qwen3-14B Instruct
1.000
0.979
0.979
0.986
0.010
Llama-3.2-1B Base
0.812
0.833
0.833
0.826
0.010
Llama-3.2-1B Instruct
0.625
0.625
0.625
0.625
0.000
Gemma-4 E4B Base
0.604
0.479
0.625
0.569
0.064
Appendix
Table 4: Probe accuracy across three training seeds at the globally-best (layer, feature mode) selected by three-seed mean.
Model
Language
Behavior
LM-head
Probe
Surplus
Qwen3-0.6B Base
English
0.656
0.875
0.938
+0.281
Chinese
0.656
0.625
1.000
+0.344
German
0.469
0.625
0.875
+0.406
Qwen3-0.6B Instruct
English
0.688
0.875
0.812
+0.125
Chinese
0.469
0.562
1.000
+0.531
German
0.375
0.562
0.688
+0.312
Appendix
Table 5: Per-language breakdown for all seven models.
Model
Variant
Summary
Accuracy
Object
Subject
Order Gap
Qwen3-0.6B Base
debiased_forced_choice_ab
debiased
0.531
0.604
0.458
0.062
Qwen3-0.6B Base
debiased_label_choice
debiased
0.552
0.542
0.562
0.021
Qwen3-0.6B Base
contrastive_scoring
single
0.479
0.667
0.292
nan
Qwen3-0.6B Instruct
debiased_forced_choice_ab
debiased
0.500
0.500
0.500
0.000
Qwen3-0.6B Instruct
debiased_label_choice
debiased
0.562
0.479
0.646
0.083
Qwen3-0.6B Instruct
contrastive_scoring
single
0.500
0.708
0.292
nan
Appendix
Table 6: Readout-cleaning results after counterbalancing option order and label order.
Figure 6: Raw accuracy under both option orders and the debiased result, per model.
Figure 7: Patching causal lift versus relative layer depth on QA subject-control (vertical dotted lines mark probe peaks).
Figure 8: Patching causal lift across all four cells per model.
Model
Task
Type
Behavior
Best Probe
Gap
Probe Layer
Probe Mode
Qwen3-0.6B Base
qa
object-control
0.375
0.917
0.542
18
ab
Qwen3-0.6B Base
qa
subject-control
0.708
0.931
0.222
18
ab
Qwen3-0.6B Base
paraphrase
object-control
0.750
0.917
0.167
18
ab
Qwen3-0.6B Base
paraphrase
subject-control
0.542
0.931
0.389
18
ab
Qwen3-0.6B Instruct
qa
object-control
0.583
0.847
0.264
16
ab
Qwen3-0.6B Instruct
qa
subject-control
0.250
0.903
0.653
19
bv
Appendix
Table 7: Type-level behavior-versus-probe comparison for the most relevant models.
Mode
Type
Item-out
Verb-out
Gap
ab
subj
0.819±0.05
0.681±0.13
+0.139
bv
subj
0.958±0.04
0.944±0.06
+0.014
av
subj
0.972±0.02
0.889±0.09
+0.083
abv
subj
0.903±0.05
0.889±0.10
+0.014
v
subj
1.000±0.00
0.889±0.05
+0.111
ab
obj
0.806±0.02
0.792±0.00
+0.014
Appendix
Table 8: Item-out LOOCV vs. verb-out CV at the Qwen3-14B Instruct v -mode probe peak (layer 17, mean ± SD across three seeds).
Figure 9: Subject-control verb-out gap (item-out LOOCV − verb-out CV) across the full layer trajectory in bv mode.
Model
Layer
Item-out
Verb-out
Gap
Verdict
Qwen3-0.6B Base
18
0.917±0.00
0.500±0.04
+0.417
LEAK
Qwen3-0.6B Instruct
19
0.903±0.02
0.472±0.06
+0.431
LEAK
Llama-3.2-1B Base
7
0.917±0.00
0.250±0.00
+0.667
LEAK
Llama-3.2-1B Instruct
2
0.625±0.00
0.083±0.00
+0.542
LEAK
Gemma-4 E4B
2
0.611±0.06
0.153±0.02
+0.458
LEAK
Gemma-4 E4B Instruct
4
0.750±0.00
0.306±0.06
+0.444
LEAK
Appendix
Table 9: Subject-control full-trajectory probe peak in bv mode across all seven models, with verb-out CV at the same layer (mean ± standard deviation across three seeds; gap criterion as in Figure 9 ).