Understanding how Large Language Models (LLMs) encode linguistic structures remains a fundamental challenge in interpretability research. While diagnostic classifiers (or "probes") are widely used for this task, they face significant methodological criticism: training auxiliary classifiers introduces capacity confounds and calibration issues, often making it difficult to distinguish the model's intrinsic representations from the probe's ability to learn the task. To address these limitations, we introduce a probe-free framework for localizing linguistic selectivity at the individual neuron level. Leveraging the controlled contrasts of linguistic minimal pairs, we propose a Neuron Separability Index (NSI), a metric that directly quantifies how reliably single neurons differentiate grammatical from ungrammatical constructions without parameter updates. Applying NSI across 68 linguistic paradigms and seven checkpoints reveals three main patterns: 1) raw separability reaches near-peak levels earlier for morphological and syntactic distinctions than for syntax-semantics interface and conceptual distinctions. 2) after permutation normalization, single-unit selectivity is sparse, weak, and narrowly tuned: only a small fraction of units are sensitive to an average paradigm, and strongly selective "grandmother neurons" are rare. 3) whole-vector linear separability, single-neuron selectivity, and behavioral competence are largely dissociated, and targeted ablations further separate activation selectivity from causal reliance.
Figures & tables
Figure 1: Minimal-pair neuron separability pipeline. Minimal pairs from BLiMP/COMPS are fed into a transformer LM (here, Qwen3-0.6B). For each minimal pair, we extract last-token activations, then focus on a single neuron and assemble paired activation vectors across items for the grammatical (positive) and ungrammatical (negative) sentences. We compute a raw separability score from the correlation between the two paired vectors. To control for lexical effects, we construct a null distribution by randomly swapping the grammatical/ungrammatical labels within half of the pairs and recomputing separability across many permutations; the observed raw score is then converted into a neuron separability index (NSI) value.
Figure 2: Colors indicate linguistic domains: Concept , Syntax–Semantics Interface , Syntax , and Morphology . Layer-wise average separability curves across linguistic phenomena. Each curve reports mean neuron separability for one phenomenon, averaged across its constituent paradigms.
Figure 3: (a) Proportion of sensitive neurons ( NSI>0 ) across layers, broken down by linguistic phenomenon. Lines show the mean proportion across paradigms within each phenomenon; shaded bands indicate variability across paradigms. Overall, neurons with positive sensitivity to the targeted contrasts are sparse throughout the network. Mean recruitment differs across the four domains: Concept ( ≈8.27% ), Syntax–Semantics Interface ( ≈3.31% ), Syntax ( ≈2.50% ), and Morphology ( ≈3.88% ). Sensitivity also tends to decrease in deeper layers, suggesting that later layers allocate relatively less capacity to local grammatical well-formedness and more to higher-level information (e.g., semantics, reasoning, and discourse). The right panel of Appendix Figure 11 provides the paradigm-level distributions. (b) Distribution of neuron breadth across paradigms. For each neuron (a unit at a specific layer), we count the paradigms for which it is responsive ( NSI>0 ). Left: histogram of the number of responsive paradigms per neuron. Right: empirical CDF of the same quantity. Vertical lines mark the mean and the 90th/95th percentiles, showing that most responsive neurons participate in only a small number of paradigms and that broad, multi-paradigm responsiveness is rare. Corresponding response distributions at the domain and phenomenon levels appear in Appendix Figures 14 and 14 .
Figure 4: (a) Mean NSI across layers computed over sensitive neurons only (NSI >0 ), grouped by linguistic phenomenon. Even within this subset, the average NSI remains small, indicating that sensitivity is generally weak rather than sharply selective. (b) For each BLiMP/COMPS paradigm, we compute the maximum NSI across neurons. Bar and point colors indicate each paradigm’s parent phenomenon; the outlined Others bar shows the mean over the remaining 65 paradigms. Only three paradigms contain a neuron with NSI >2 . Appendix Figure 12 gives paradigm- and neuron-level detail.
Figure 5: Layer-wise average tuning breadth at the paradigm, phenomenon, and domain levels. A neuron is counted as responsive to a phenomenon or domain when it has positive selectivity for at least one constituent paradigm. Shaded regions show standard error.
Figure 6: Functional modularity in tuning across linguistic phenomena. The heatmap displays mean phenomenon-level NSI for neurons responsive to at least three phenomena. Rows are grouped by domain, and columns are hierarchically clustered neurons. The block structure shows that even poly-selective neurons retain dominant within-domain preferences rather than behaving as uniform cross-domain grammar detectors.
Figure 7: Robustness and generalization of neuron-level selectivity. (a) Across four Qwen3 scales and three additional checkpoints, only 2.5 – 3.9% of units have positive NSI for an average paradigm. (b) The average neuron responds to only 1.7 – 2.7 of 68 paradigms; diamonds mark the 95th percentile. (c) For Qwen3-0.6B, the fraction of units above threshold falls sharply, and only 3 of 68 paradigms contain any unit from NSI ≥1.645 through NSI >3 . (d) Whole-vector probe accuracy is essentially uncorrelated with minimal-pair behavior across 67 BLiMP paradigms.
Paradigm
k
Top ΔM [95% CI]
Random mean
Bottom ΔM
prand
Det.–noun agr.
5
−0.005 [ −0.014,0.003 ]
−0.003
0.147
0.446
10
0.000 [ −0.011,0.011 ]
−0.009
0.054
0.634
20
0.111 [ 0.084,0.138 ]
−0.034
−0.022
0.911
Det.–noun agr.+adj.
5
−0.006 [ −0.015,0.004 ]
−0.005
0.161
0.545
10
−0.005 [ −0.020,0.012 ]
−0.012
0.177
0.604
20
0.026 [ 0.004,0.048 ]
−0.009
0.351
0.733
Table 1: Primary targeted group-ablation results. ΔM is the change in mean logp(x+)−logp(x−) ; brackets give paired-bootstrap 95% CIs, and prand compares each top group with 100 same-size random groups.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Paradigm-first 90% saturation layers summarized by phenomenon. For each paradigm, we identify the earliest layer at which its median raw-separability score reaches 90% of its own maximum. Bars show the mean of these paradigm-level saturation layers within each phenomenon. The legend reports means computed directly over all constituent paradigms in each domain: 19.0 for Concept, 10.3 for Syntax–Semantics Interface, 9.5 for Syntax, and 8.4 for Morphology.
Figure 9: Correspondence between raw-separability and probing 90% saturation layers across the 68 paradigms. For each paradigm and measure, saturation is the earliest layer reaching 90% of that paradigm’s own maximum. Colors encode phenomena, marker shapes encode domains, the dashed diagonal indicates equal layers, and the solid line is the least-squares fit. The association is weakly positive (Pearson r=0.15 ; Spearman ρ=0.23 ), indicating limited paradigm-level agreement between the two localization measures. Probing embedding index 0 is excluded, and sampled hidden-state indices 2, 4, …, 28 are mapped to model layers 1, 3, …, 27.
Model
Layers × width
NSI >0 (%)
NSI >2 (%)
Paradigms >2
Mean B
P95/max B
B≥10 (%)
Qwen3-0.6B
28×1024
3.568
1.54×10−4
3
2.43
6/13
0.18
Qwen3-1.7B
28×2048
3.261
2.82×10−4
6
2.22
5/13
0.09
Qwen3-4B
36×2560
2.524
1.12×10−4
2
1.72
5/13
0.06
Qwen3-8B
36×4096
2.593
1.92×10−3
4
1.76
5/15
0.24
Pythia-410M
24×1024
3.942
5.39×10−4
2
2.68
7/16
0.86
TinyLlama-1.1B
22×2048
2.939
6.20×10−4
2
2.00
5/17
0.35
Appendix
Table 2: Cross-model selectivity and tuning breadth over 68 paradigms. “Paradigms >2 ” is the number of paradigms containing any NSI >2 unit. Breadth B counts the number of paradigms with NSI >0 for each neuron.
Figure 10: Cumulative distribution of paradigm-level tuning breadth across models. Even under the lenient NSI >0 definition, the high-breadth tail is small: the 95th percentile is 5–7 paradigms and fewer than 1% of neurons respond to 10 or more paradigms in every model.
Figure 11: Left : Distributions of non-zero NSI scores (i.e., over sensitive neurons with NSI>0 ) aggregated across all layers, shown separately for each paradigm. Across paradigms, the mean NSI is typically around ∼0.5 , indicating that most sensitive neurons exhibit only weak separability. Right : Paradigm-wise proportion of sensitive neurons across all 68 BLiMP/COMPS paradigms for Qwen3-0.6B.
Figure 12: Paradigm-wise maximum NSI scores across all 68 BLiMP/COMPS paradigms for Qwen3-0.6B. Only three paradigms contain a unit with NSI >2 ; the bottom panels show that each case is realized by one unit across the network.
Figure 15
Threshold
Units above (%)
Paradigms with any
Mean count
Max count
0
3.7010
68
1061.15
5229
1
0.0944
68
27.06
120
1.645
1.54×10−4
3
0.044
1
2
1.54×10−4
3
0.044
1
2.326
1.54×10−4
3
0.044
1
2.58
1.54×10−4
3
0.044
1
Appendix
Table 3: Threshold sweep over 68 paradigms and 28×1024 units per paradigm.
Condition
Mean raw
NSI >0 (%)
NSI >2 (%)
Paradigms >2
Median max
Matched good–bad
0.1438
3.469
6.67×10−4
4
1.128
Random good–good
0.5008
100.000
0
0
0.980
Random bad–bad
0.5009
100.000
0
0
0.989
Cross-item good–bad
0.5004
16.464
0.0454
50
3.325
Appendix
Table 4: Pairing controls over 68 paradigms.
Subset
Condition
Paradigms
Mean raw
Median raw
One-prefix
Original
10
0.0753
0.0757
Critical deleted
10
0.0011
0.0007
Random deleted
10
0.0746
0.0772
Two-prefix
Original
6
0.1007
0.1069
Shared word deleted
6
0.1802
0.1500
Random deleted
6
0.0971
0.0933
Appendix
Table 5: Critical-token deletion controls. We emphasize raw separability because deletion changes the geometry and variance of the permutation null.
Comparison
n
Pearson r
Spearman ρ
Behavior vs. layer-14 probe
67
-0.032
0.023
Behavior vs. peak probe
67
-0.030
0.008
Behavior vs. mean NSI
67
-0.021
-0.022
Behavior vs. maximum NSI
67
0.169
0.177
Layer-14 probe vs. mean NSI
67
-0.557
-0.578
Layer-14 probe vs. maximum NSI
67
0.092
-0.178
Appendix
Table 6: Correlations across 67 BLiMP paradigms.
Figure 15: Additional comparisons among behavior, whole-vector probing, and neuron-level NSI. (a) Mean NSI is uncorrelated with behavioral accuracy. (b) Whole-vector probe accuracy and mean NSI are negatively correlated, further showing that distributed linear separability and average single-neuron selectivity capture different properties.
Domain
Paradigms
Behavior
Probe
Mean NSI
Max NSI
Syntax–Semantics Interface
23
0.727
0.907
0.018
1.133
Syntax
26
0.736
0.964
0.014
1.424
Morphology
18
0.831
0.861
0.024
1.932
Appendix
Table 7: Domain-level averages over the 67 BLiMP paradigms. The legacy semantics , syntax_semantics , and syntax/semantics source labels are merged into the Syntax–Semantics Interface domain.
Paradigm
Domain
Behavior
Probe
Mean NSI
Max NSI
principle_A_case_1
Syntax–Semantics Interface
0.999
0.936
0.005
1.134
anaphor_number_agreement
Morphology
0.986
0.951
0.011
1.118
principle_A_domain_1
Syntax–Semantics Interface
0.985
1.000
0.010
1.224
wh_vs_that_with_gap_long_distance
Syntax
0.221
0.973
0.006
1.098
sentential_subject_island
Syntax
0.238
1.000
0.009
1.105
wh_vs_that_with_gap
Syntax
0.298
0.996
0.005
1.064
Appendix
Table 8: Representative paradigms ranked by behavioral accuracy. Near-perfect linear probing can coexist with poor model behavior.
Figure 16: Targeted ablation at each frozen candidate’s actual layer. Red and gray lines show the top-score and bottom signed-score sets; the blue line and band show the mean and 2.5th–97.5th percentile interval across 100 same-layer random sets. Negative values indicate a reduced grammatical-over-ungrammatical log-probability margin.
Paradigm
k
Top ΔM
Top CI
Random mean
Random interval
Bottom ΔM
ΔA
prand
Det.–noun agr.
1
0.001
[ −0.004,0.006 ]
−0.001
[ −0.038,0.017 ]
0.159
0.001
0.554
5
−0.005
[ −0.014,0.003 ]
−0.003
[ −0.152,0.119 ]
0.147
−0.001
0.446
10
0.000
[ −0.011,0.011 ]
−0.009
[ −0.188,0.263 ]
0.054
0.002
0.634
20
0.111
[ 0.084,0.138 ]
−0.034
[ −0.236,0.218 ]
−0.022
−0.002
0.911
Det.–noun agr.+adj.
1
−0.003
[ −0.008,0.001 ]
−0.005
[ −0.077,0.021 ]
0.174
0.001
0.396
5
−0.006
[ −0.015,0.004 ]
−0.005
[ −0.106,0.110 ]
0.161
0.000
0.545
Appendix
Table 9: Complete targeted-ablation results. Top CI is the paired-bootstrap 95% interval over items; Random interval is the 2.5th–97.5th percentile interval across 100 random sets. ΔA is the change in total-log-probability minimal-pair accuracy.
Large Language Models (LLMs) achieve strong linguistic performance, yet their internal mechanisms for producing these predictions remain unclear. We investigate the hypothesis that LLMs encode representations of linguistic constraint violations within their parameters, which are selectively activated when processing ungrammatical sentences. To test this, we use sparse autoencoders to decompose polysemantic activations into sparse, monosemantic features and recover candidates for violation-related features. We introduce a sensitivity score for identifying features that are preferentially activated on constraint-violated versus well-formed inputs, enabling unsupervised detection of potential violation-specific features. We further propose a conjunctive falsification framework with three criteria evaluated jointly. Overall, the results are negative in two respects: (1) the falsification criteria are not jointly satisfied across linguistic phenomena, and (2) no features are consistently shared across all categories. While some phenomena show partial evidence of selective causal structure, the overall pattern provides limited support for a unified set of grammatical violation detectors in current LMs.
Hardy, Sebastian Padó
IMS, University of Stuttgart, Stuttgart, Germany · Universitas Mikroskil, Medan, Indonesia.
Whether neural language models (NLMs) possess the ability to distinguish strings on the basis of their grammaticality remains a debated topic in the computational linguistics literature. Existing evidence has largely relied on probability-based measures, testing whether models assign higher probabilities to grammatical than ungrammatical strings. However, probability comparisons have been criticized as a measure for grammatical knowledge based on the assumption that grammaticality is inherently entangled with likelihood. Model-assigned probability is a function of many related sentence properties, such as lexical frequency, plausibility, and world knowledge. In this work, we move beyond probability-based evaluations and investigate whether grammaticality is encoded in the internal representations of NLMs. Using mass-mean probing, we test whether grammatical and ungrammatical sentences are systematically separated in representational space. We further examine the extent to which these representations are independent of sentence properties that are correlated with grammaticality, as well as their generalization across grammatical phenomena and languages. Our results provide evidence that grammaticality is robustly encoded in sentence representations of a wide range of pretrained NLMs, yielding clear representational separation on the dimension of grammaticality that cannot be fully explained by alternative sentence-level factors. Moreover, this encoding generalizes across a broad range of grammatical phenomena and to some degree, across languages, suggesting that grammaticality constitutes a coherent representational dimension in contemporary NLMs. These findings contribute new evidence to debates about the nature of syntactic knowledge in language models and offer a complementary framework for evaluating grammatical competence that is not dependent on string probabilities alone.
Jane Li, Najoung Kim
Department of Cognitive Science Johns Hopkins University Baltimore, MD · Department of Linguistics Boston University Boston, MA
Grammaticality and likelihood are distinct notions in human language. Pretrained language models (LMs), which are probabilistic models of language fitted to maximize corpus likelihood, generate grammatically well-formed text and discriminate well between grammatical and ungrammatical sentences in tightly controlled minimal pairs. However, their string probabilities do not sharply discriminate between grammatical and ungrammatical sentences overall. But do LMs implicitly acquire a grammaticality distinction distinct from string probability? We explore this question through studying internal representations of LMs, by training a linear probe on a dataset of grammatical and (synthetic) ungrammatical sentences obtained by applying perturbations to a naturalistic text corpus. We find that this simple grammaticality probe generalizes to human-curated grammaticality judgment benchmarks and outperforms LM probability-based grammaticality judgments. When applied to semantic plausibility benchmarks, in which both members of a minimal pair are grammatical and differ in only plausibility, the probe however performs worse than string probability. The English-trained probe also exhibits nontrivial cross-lingual generalization, outperforming string probabilities on grammaticality benchmarks in numerous other languages. Additionally, probe scores correlate only weakly with string probabilities. These results collectively suggest that LMs acquire to some extent an implicit grammaticality distinction within their hidden layers.