Neural Language Models Learn the Contextual Distributions of Dependency Structures: a statistical learning theory to compositionality
Authors: Wang Bojun, Junjie Chen, Holly Jenkins, Elizabeth Wonnacott
Organizations: Department of Education, University of Oxford, Oxford, United Kingdom · Department of Statistics, University of Oxford, Oxford, United Kingdom
It is unclear how Neural Language Models (NLMs) acquire the structural meaning encoded by grammatical structures that is independent of lexical semantics. We propose a statistical learning process in which learned dependency structures themselves become new distributional units for subsequent statistical learning. Under this account, once a dependency structure is acquired, the model tracks its contextual distributions. These contextual features reflect the semantic properties of a composite structure. To test this hypothesis, we design a synthetic grammar in which each grammatical structure has distinct contextual distributions that cannot be recovered from the distributional statistics of their component tokens alone. We train a series of BERT-style masked language models on this grammar and examine their developmental trajectory. The results show that models can successfully learn the contextual distributions of composite dependency structures even though they cannot be inferred from token statistics alone. Developmental analysis further reveals a clear developmental trajectory. The learning of the dependency relations that define a grammatical structure consistently precedes the learning of its contextual features. These findings suggest that statistical learning in NLMs is not merely the accumulation of token co-occurrence statistics, but a process in which learned dependency structures become new units of distributional learning. We argue that this process provides a statistical-learning account of how NLMs solve the compositionality problem in language. Finally, we discuss the possibility that this statistical learning process provides an explanatory theory on how language cognition could emerge from pure distributional statistics.
Figures & tables
Dataset subset
Proportion of complete dataset
Sequence type
Example seven-token sequence
Core set
20%
BAA with M
M U B A A U U
Core set
20%
ABA with N
U A B A U U N
Compensation set
16%
M and N co-occur
U N U U M U U
Compensation set
4%
M only
U U U U U M U
Compensation set
4%
N only
N U U U U U U
Compensation set
36%
Noise only
U U U U U U U
Table 1: Composition of the artificial-language dataset
Figure 1: Developmental change in the base-language model’s first-layer V-vector representations: categories overlap at initialization, A and B separate by step 4,000, and M and N separation develops by step 39,200.
Figure 2: First-layer V-vector projections for the control-language model at the initial and minimum-validation-loss checkpoint; A and B are separated, whereas M and N remain overlapping.
Figure 3: Probability mass analysis through the learning path. Solid lines show the mean probability mass, ribbons show the 95% confidence interval.
Figure 4: Controlled probability-mass difference across 100 models for M, N, and random-U contexts. (Panel A) Full training path; (Panel B) close-up for the first 3,000 updates. Solid lines show model means; ribbons show 95% confidence intervals.
Figure 5: Probability mass analysis for 4L-4H-16D models, 8L-8H-32D models, and 8L-8H-64D up to 5,000 steps.
Figure 6: (Panel A) Full training paths; (Panel B) early learning, zoomed to 10,000, 3,000, and 2,000 updates, respectively. Lines show model means; ribbons show 95% confidence intervals.
Figure 7: Static first-layer V-vector trajectories of the 4L-4H-16D base-language model. The panels show training steps 0, 3,600, and 32,800.
Figure 8: 4L-4H-16D model trained on the control language. Initial stage clustering and lowest validation-loss clustering.
Figure 9: Static first-layer V-vector trajectories of the 8L-8H-64D base-language model. The panels show training steps 0, 6,400 and 48,600.
Figure 10: 8L-8H-64D model trained on the control language. Initial stage clustering and lowest validation-loss clustering.
Figure 11: recursion of internal-external dependency nesting
In this study, we use a developmental approach to investigate the statistical learning and mental representation of neural language models (NLM). A series of Generative Transformer models are trained on a synthetic grammar. The model states are saved at multiple stages in the course of training. Through analyzing how the internal representations of these models change in the developmental path, we found that NLMs acquire the most abstract global statistical knowledge at the beginning of learning and later acquire the relatively local statistical dependencies. This learning path contains many over-generalizations from the very beginning and these over-generalizations are gradually constrained in the later stage of learning. Based on this observation, we propose a new framework to explain the statistical learning and language cognition of NLMs.
Whether neural language models (NLMs) possess the ability to distinguish strings on the basis of their grammaticality remains a debated topic in the computational linguistics literature. Existing evidence has largely relied on probability-based measures, testing whether models assign higher probabilities to grammatical than ungrammatical strings. However, probability comparisons have been criticized as a measure for grammatical knowledge based on the assumption that grammaticality is inherently entangled with likelihood. Model-assigned probability is a function of many related sentence properties, such as lexical frequency, plausibility, and world knowledge. In this work, we move beyond probability-based evaluations and investigate whether grammaticality is encoded in the internal representations of NLMs. Using mass-mean probing, we test whether grammatical and ungrammatical sentences are systematically separated in representational space. We further examine the extent to which these representations are independent of sentence properties that are correlated with grammaticality, as well as their generalization across grammatical phenomena and languages. Our results provide evidence that grammaticality is robustly encoded in sentence representations of a wide range of pretrained NLMs, yielding clear representational separation on the dimension of grammaticality that cannot be fully explained by alternative sentence-level factors. Moreover, this encoding generalizes across a broad range of grammatical phenomena and to some degree, across languages, suggesting that grammaticality constitutes a coherent representational dimension in contemporary NLMs. These findings contribute new evidence to debates about the nature of syntactic knowledge in language models and offer a complementary framework for evaluating grammatical competence that is not dependent on string probabilities alone.
Jane Li, Najoung Kim
Department of Cognitive Science Johns Hopkins University Baltimore, MD · Department of Linguistics Boston University Boston, MA
Neural language models are typically trained on next-token prediction, although linguistic structure spans multiple temporal scales. Successor representations (SRs) make this horizon explicit by encoding discounted distributions over future states. Here, we ask whether such predictive representations can recover not only word classes, but also finer functional and construction-like structure from natural language. A residual network trained on WikiText-103 predicts SR distributions at three horizons without part-of-speech supervision. At the shortest horizon, unsupervised clustering robustly recovers nouns, verbs, and adjectives, while directed inter-cluster transitions reproduce familiar syntactic asymmetries. At finer resolutions and across 13 part-of-speech categories, the same geometry reveals semantic-functional groupings that cross category boundaries and directed relations tracing candidate date, measurement, and title-name constructions. Part-of-speech agreement declines as the predictive horizon lengthens. These results suggest that word classes are coarse regions within a richer predictive geometry in which categorical and construction-like linguistic structure emerge from future-word distributions.
Mathis Immertreu, Achim Schilling, Thomas Kinfe +1
Cognitive Computational Neuroscience Group, Friedrich-Alexander-Universität Erlangen–Nürnberg (FAU), Germany. · Mannheim Center for Neuromodulation and Neuroprosthetics (MCNN), University Hospital Mannheim, University Heidelberg, Germany.