cs.CLSep 24, 2026

Encoded but Not Decoded: Layer-Localized Evidence for a Three-Level Gap in LLM Syntax

Authors: Zhenyan Lu, He Wang, Xiaohui Huang

Organizations: College of International Studies, National University of Defense Technology, Nanjing, China

Abstract

A language model can fail a syntactic test in two distinct ways: by not encoding the relevant structure, or by encoding it but failing to use it at the output. Behavioral evaluation alone cannot tell these apart. We propose a three-level evaluation framework (behavioral deployment, LM-head readout, and probe recoverability) measured on the same items under the same binary decision. Using a compact trilingual (English, Chinese, German) control-dependency benchmark, we find that probe recoverability exceeds or equals LM-head readout, which in turn exceeds or equals behavioral deployment, across seven models and all three languages in the aggregate. The recoverability surplus is never negative across all 14 (model, task) conditions. The disconnect concentrates in subject-control, where a nearest-noun heuristic gives the wrong answer. The single largest gap (0.653) appears on Qwen3-0.6B Instruct in question answering. The gap persists at Qwen3-14B Instruct. Instruction tuning degrades deployment more than encoding in percentage terms. We rule out option-position bias, late-layer erasure, output-formatting artifacts, and probe-training variance. The pattern is consistent with decoding that favors surface shortcuts, and the behavior-probe gap measures the strength of that preference. Activation patching shows the gap is layer-localized. Under instruction tuning, the LM-head-decoded layer shifts approximately ten layers later than the probe-decoded layer. These findings argue that behavioral evaluation understates what models encode, while probing alone overstates what they deploy.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Beyond Decodability: Reconstructing Language Model Representations with an Encoding Probe

    May 1, 2026Gaofei Shen, Martijn Bentum, Tomas O. Lentz +2Language ModelingDecoding

  2. Linear Probes Detect Task Format, Not Reasoning Mode in Language Model Hidden States

    Jun 1, 2026Subramanyam Sahoo, Vinija Jain, Aman Chadha +1LLM Reasoning StrategiesHidden States

  3. Implicit Representations of Grammaticality in Language Models

    May 6, 2026Yingshan Susan Wang, Linlu Qiu, Zhaofeng Wu +2GrammaticalityText Corpora