Encoded but Not Decoded: Layer-Localized Evidence for a Three-Level Gap in LLM Syntax
Organizations: College of International Studies, National University of Defense Technology, Nanjing, China
Abstract
A language model can fail a syntactic test in two distinct ways: by not encoding the relevant structure, or by encoding it but failing to use it at the output. Behavioral evaluation alone cannot tell these apart. We propose a three-level evaluation framework (behavioral deployment, LM-head readout, and probe recoverability) measured on the same items under the same binary decision. Using a compact trilingual (English, Chinese, German) control-dependency benchmark, we find that probe recoverability exceeds or equals LM-head readout, which in turn exceeds or equals behavioral deployment, across seven models and all three languages in the aggregate. The recoverability surplus is never negative across all 14 (model, task) conditions. The disconnect concentrates in subject-control, where a nearest-noun heuristic gives the wrong answer. The single largest gap (0.653) appears on Qwen3-0.6B Instruct in question answering. The gap persists at Qwen3-14B Instruct. Instruction tuning degrades deployment more than encoding in percentage terms. We rule out option-position bias, late-layer erasure, output-formatting artifacts, and probe-training variance. The pattern is consistent with decoding that favors surface shortcuts, and the behavior-probe gap measures the strength of that preference. Activation patching shows the gap is layer-localized. Under instruction tuning, the LM-head-decoded layer shifts approximately ten layers later than the probe-decoded layer. These findings argue that behavioral evaluation understates what models encode, while probing alone overstates what they deploy.
Figures & tables
| Model | QA | Paraphrase | Mean Behavior | Best LM-head | Best Probe | Surplus | B/R Ratio |
| Qwen3-0.6B Base | 0.542 | 0.646 | 0.594 | 0.708 | 0.938 | 0.344 | 0.633 |
| Qwen3-0.6B Instruct | 0.417 | 0.604 | 0.510 | 0.667 | 0.833 | 0.323 | 0.612 |
| Qwen3-14B Instruct | 0.812 | 0.771 | 0.792 | 0.792 | 1.000 | 0.208 | 0.792 |
| Llama-3.2-1B Base | 0.521 | 0.625 | 0.573 | 0.688 | 0.812 | 0.240 | 0.705 |
| Llama-3.2-1B Instruct | 0.479 | 0.542 | 0.510 | 0.583 | 0.625 | 0.115 | 0.817 |
| Gemma-4 E4B | 0.521 | 0.604 | 0.562 | 0.604 | 0.604 | 0.042 | 0.931 |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| QA Format | Paraphrase Format |
| Sentence: “John promised Mary to leave early.” | Sentence: “John promised Mary to leave early.” |
| Prompt: Q: Who is understood to leave early? A: John log-prob scored A: Mary log-prob scored | Prompt: Which sentence better describes the situation? (A) John is understood to leave early. scored (B) Mary is understood to leave early. scored |
| Decision: | Decision: |
| Gold controller: John ( promise = subject-control) | |
| Mode | Qwen3-0.6B Base | Qwen3-0.6B Instruct | Qwen3- 14B Instruct |
| ab | 0.924 | 0.847 | 0.958 |
| bv | 0.889 | 0.792 | 0.979 |
| av | 0.618 | 0.458 | 0.931 |
| abv | 0.847 | 0.701 | 0.944 |
| v | 0.757 | 0.521 | 0.986 |
| Model | Seed 1729 | Seed 2718 | Seed 3141 | Mean | Std |
| Qwen3-0.6B Base | 0.938 | 0.938 | 0.896 | 0.924 | 0.020 |
| Qwen3-0.6B Instruct | 0.833 | 0.812 | 0.896 | 0.847 | 0.035 |
| Qwen3-14B Instruct | 1.000 | 0.979 | 0.979 | 0.986 | 0.010 |
| Llama-3.2-1B Base | 0.812 | 0.833 | 0.833 | 0.826 | 0.010 |
| Llama-3.2-1B Instruct | 0.625 | 0.625 | 0.625 | 0.625 | 0.000 |
| Gemma-4 E4B Base | 0.604 | 0.479 | 0.625 | 0.569 | 0.064 |
| Model | Language | Behavior | LM-head | Probe | Surplus |
| Qwen3-0.6B Base | English | 0.656 | 0.875 | 0.938 | +0.281 |
| Chinese | 0.656 | 0.625 | 1.000 | +0.344 | |
| German | 0.469 | 0.625 | 0.875 | +0.406 | |
| Qwen3-0.6B Instruct | English | 0.688 | 0.875 | 0.812 | +0.125 |
| Chinese | 0.469 | 0.562 | 1.000 | +0.531 | |
| German | 0.375 | 0.562 | 0.688 | +0.312 |
| Model | Variant | Summary | Accuracy | Object | Subject | Order Gap |
| Qwen3-0.6B Base | debiased_forced_choice_ab | debiased | 0.531 | 0.604 | 0.458 | 0.062 |
| Qwen3-0.6B Base | debiased_label_choice | debiased | 0.552 | 0.542 | 0.562 | 0.021 |
| Qwen3-0.6B Base | contrastive_scoring | single | 0.479 | 0.667 | 0.292 | nan |
| Qwen3-0.6B Instruct | debiased_forced_choice_ab | debiased | 0.500 | 0.500 | 0.500 | 0.000 |
| Qwen3-0.6B Instruct | debiased_label_choice | debiased | 0.562 | 0.479 | 0.646 | 0.083 |
| Qwen3-0.6B Instruct | contrastive_scoring | single | 0.500 | 0.708 | 0.292 | nan |
| Model | Task | Type | Behavior | Best Probe | Gap | Probe Layer | Probe Mode |
| Qwen3-0.6B Base | qa | object-control | 0.375 | 0.917 | 0.542 | 18 | ab |
| Qwen3-0.6B Base | qa | subject-control | 0.708 | 0.931 | 0.222 | 18 | ab |
| Qwen3-0.6B Base | paraphrase | object-control | 0.750 | 0.917 | 0.167 | 18 | ab |
| Qwen3-0.6B Base | paraphrase | subject-control | 0.542 | 0.931 | 0.389 | 18 | ab |
| Qwen3-0.6B Instruct | qa | object-control | 0.583 | 0.847 | 0.264 | 16 | ab |
| Qwen3-0.6B Instruct | qa | subject-control | 0.250 | 0.903 | 0.653 | 19 | bv |
| Mode | Type | Item-out | Verb-out | Gap |
| ab | subj | |||
| bv | subj | |||
| av | subj | |||
| abv | subj | |||
| v | subj | |||
| ab | obj |
| Model | Layer | Item-out | Verb-out | Gap | Verdict |
| Qwen3-0.6B Base | 18 | LEAK | |||
| Qwen3-0.6B Instruct | 19 | LEAK | |||
| Llama-3.2-1B Base | 7 | LEAK | |||
| Llama-3.2-1B Instruct | 2 | LEAK | |||
| Gemma-4 E4B | 2 | LEAK | |||
| Gemma-4 E4B Instruct | 4 | LEAK |