Some properties of languages, e.g., subject-object-verb (SOV) word order, are more prevalent than others among the thousands of attested natural languages (NLs). Such typological commonality is often attributed to learning biases. Computational simulations, recently with language models (LMs), have facilitated the exploration of this theory. In this paper, we extend existing analyses of the relationship between LMs' learning biases and typological commonality on both data and model sides, focusing on: (i) cross-serial dependencies, the upper limit of attested syntactic complexity, and (ii) stack-based LMs (SLMs), potentially facilitating learning of hierarchical patterns. We first evaluate generalization of SLMs on cross-serial dependencies across diverse artificial languages and confirm that they struggle with such constructions. However, SLMs with limited working memory generalize better suggesting a possible basis for such inductive bias and thus the typological commonality of some word order configurations.
Figures & tables
GCG Lexical Syntactic Category El-Naggar et al. (2025a) ; El-Naggar et al. (2025b)
Verb with Complement (VCOMP) – (S ∣1 NP SUBJ ) ∣1 SCOMP
Kim ga believed that Sandy o lied
Table 1: Lexical syntactic categories, their derivations, and their examples, where a word corresponding to the category is in bold. The top half shows the categories defined by El-Naggar et al. (2025a) ; El-Naggar et al. (2025b) . The bottom half shows the new categories that we added to extend the grammar. The vertical bars “ ∣ ” in the GCG lexical syntactic categories represent either forward or backward slashes, and those with the same index are controlled by the same word order parameter (see Table 2 ).
Digit
0
1
1 (Base)
SOV
VOS
2 (COMP)
Preposed complementizer
Postposed complementizer
3 (PP)
Postposition
Preposition
4 (ADJ)
Prenominal adjective
Postnominal adjective
5 (REL)
Preposed relativizer
Postposed relativizer
Table 2: Word order parameters and the constructions introduced by their assignment (0/1). The exact implementation is shown in Table 1 .
Figure 1: Partial dependency representations of CSDs with embedded relative clauses. The examples are for SOV with prenominal relative clauses ( 0XXX0 ); the pattern will be, for example, fully mirrored for VOS with postnominal relative clauses ( 1XXX1 ). Colored arcs indicate cross-serial noun–verb dependencies. Black arcs indicate relative-clause attachment. Relative clauses are bracketed.
Figure 2: PPL distribution on the Medium test set (average ± standard deviation). The bar shows the standard deviation among five runs with different random seeds. Note that all training was conducted and stopped under the same early-stopping setting, and the high variance indicates training instability across different random seeds (see section 6 ).
Model
Short
Medium
Long
Rel
Conj-1
Conj-2
Nested
Avg.
RNN
33.40 ± 0.14
100.69 ± 13.91
345.68 ± 510.02
376.50 ± 223.62
552.94 ± 483.43
731.62 ± 1033.28
538.97 ± 574.27
382.83
RNN SUP
33.32 ± 0.18
88.29 ± 16.40
249.02 ± 75.80
299.37 ± 97.15
404.95 ± 126.49
485.74 ± 160.56
437.99 ± 171.39
285.53
RNN ND
33.43 ± 0.13
94.98 ± 16.35
333.45 ± 134.35
328.11 ± 101.49
497.27 ± 165.78
615.10 ± 241.20
504.19 ± 245.97
343.79
LSTM
33.49 ± 0.22
72.72 ± 13.04
313.71 ± 86.15
277.15 ± 100.83
501.66 ± 208.33
688.17 ± 305.38
522.85 ± 257.08
344.25
LSTM SUP
33.49 ± 0.20
73.80 ± 11.75
415.24 ± 190.50
364.23 ± 154.01
661.19 ± 298.72
1057.02 ± 601.37
772.05 ± 377.53
482.43
LSTM ND
34.52 ± 7.51
88.03 ± 24.14
516.53 ± 226.21
422.71 ± 188.28
773.97 ± 388.83
1159.86 ± 664.59
882.36 ± 526.66
554.00
Table 3: PPL for each model and evaluation data (average ± standard deviation). The standard deviation here indicates how the average PPL varies among 32 languages. PPL values obtained from SLMs, which are better than those of the corresponding base model, are shown in bold . The best PPLs among the LMs with the same base architecture are shown with an underline .
Model
Short
Med.
Long
Rel
Conj1
Conj2
Nest
RNN
0.57
− 0.66 †
− 0.57 †
− 0.31 †
− 0.48 †
− 0.59 †
− 0.25
RNN SUP
− 0.24
− 0.58 †
− 0.43 †
− 0.60 †
− 0.62 †
− 0.45 †
− 0.39 †
RNN ND
0.36
− 0.40 †
− 0.24
− 0.61 †
− 0.26
− 0.12
− 0.36 †
LSTM
0.43
− 0.56 †
− 0.69 †
− 0.57 †
− 0.26
− 0.28
− 0.38 †
LSTM SUP
0.54
− 0.60 †
− 0.66 †
− 0.49 †
− 0.48 †
− 0.56 †
− 0.55 †
LSTM ND
0.59
− 0.25
0.17
− 0.24
0.10
0.20
0.12
Table 4: Spearman Correlation to Word Order Frequency by Model and Test/Evaluation Type. † indicates a significantly negative correlation (p < 0.05).
Feature
Coefficient
Significance
β0
0.0371
βF
− 0.0651
βM(LSTM)
− 0.5231
βM(TF)
− 0.0628
βS(Sup)
− 0.0213
βS(Nd)
− 0.0790
Table 5: Fixed effects from the regression model predicting perplexity. The word order digit effects ( βd , βd:M , βd:S ) are excluded from the Table. Coefficients for each parameter βeffect are shown. Significance markers denote: *** p<0.001, ** p<0.01, * p<0.05.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Artifact
License
Usage
NLTK Bird et al. (2009)
Apache License 2.0
to create ALs and targeted evaluation data
Docker Merkel and others (2014)
Apache 2.0
to train SLMs
El-Naggar et al. (2025b) Datasets
Creative Commons CC-BY 4.0
training and testing SLMs
Codes from DuSell and Cotterell (2025) ( https://github.com/bdusell/bearing-syntactic-fruit/commit/ebe4ee8 )
MIT License
to implement models
Rau ( https://github.com/bdusell/rau )
MIT Licence
to implement models
WALS Dryer and Haspelmath (2013)
Creative Commons CC-BY 4.0
to find word order statistics in NLs
Appendix
Table 6: Artifacts used in this paper.
S
82A Order of Subject and Verb Dryer (2013e)
VP
83A Order of Object and Verb Dryer (2013c)
O
81A Order of Subject, Object and Verb Dryer (2013f)
COMP
Feature GB421: Is there a preposed complementizer in complements of verbs of thinking and/or knowing? Skirgård et al. (2023)
Feature GB422: Is there a postposed complementizer in complements of verbs of thinking and/or knowing? Skirgård et al. (2023)
PP
85A Order of Adposition and Noun Phrase Dryer (2013b)
ADJ
87A Order of Adjective and Noun Dryer (2013a)
Appendix
Table 7: WALS and Grambank chapters used in this paper.
Model
Short
Medium
Long
Rel
Conj-1
Conj-2
Nested
Avg.
RNN
33.40 ± 0.14
100.69 ± 13.91
345.68 ± 510.02
376.50 ± 223.62
552.94 ± 483.43
731.62 ± 1033.28
538.97 ± 574.27
382.83
RNN SUP (w/o SC)
33.44 ± 0.18
100.43 ± 13.04
292.64 ± 88.90
352.21 ± 127.53
506.62 ± 189.70
623.84 ± 258.64
533.75 ± 249.39
348.99
RNN ND (w/o SC)
33.45 ± 0.14
100.40 ± 12.66
271.09 ± 86.73
324.98 ± 89.09
428.46 ± 117.36
492.40 ± 150.42
447.41 ± 147.36
299.74
LSTM
33.49 ± 0.22
72.72 ± 13.04
313.71 ± 86.15
277.15 ± 100.83
501.66 ± 208.33
688.17 ± 305.38
522.85 ± 257.08
344.25
LSTM SUP (w/o SC)
33.45 ± 0.18
75.78 ± 12.06
374.29 ± 133.49
301.13 ± 112.04
574.42 ± 283.83
850.53 ± 507.61
615.52 ± 351.33
403.59
LSTM ND (w/o SC)
33.48 ± 0.21
79.94 ± 14.51
359.74 ± 105.87
307.96 ± 106.65
537.33 ± 202.30
753.00 ± 314.25
581.48 ± 258.67
378.99
Appendix
Table 8: PPL for each dataset and evaluation data (average ± standard deviation). The standard deviation here indicates how the average PPL varies among 32 languages. “SC” denotes the shortcut connection. PPL values obtained from SLMs, which are better than those of the corresponding base model, are shown in bold . The best PPLs among the LMs with the same base architecture are shown with an underline .
Model
Short
Medium
Long
Rel
Conj-1
Conj-2
Nested
RNN
0.57
− 0.66 †
− 0.57 †
− 0.31 †
− 0.48 †
− 0.59 †
− 0.25
RNN SUP (w/o shortcut)
0.52
0.04
− 0.07
0.15
0.30
0.39
0.29
RNN ND (w/o shortcut)
0.36
0.18
− 0.31 †
− 0.05
0.03
0.22
0.20
LSTM
0.43
− 0.56 †
− 0.69 †
− 0.57 †
− 0.26
− 0.28
− 0.38 †
LSTM SUP (w/o shortcut)
0.55
− 0.65 †
− 0.56 †
− 0.68 †
− 0.38 †
− 0.34 †
− 0.48 †
LSTM ND (w/o shortcut)
0.61
− 0.53 †
− 0.58 †
− 0.67 †
− 0.39 †
− 0.33 †
− 0.44 †
Appendix
Table 9: Spearman Correlation to Word Order Frequency by Model and Test/Evaluation Type. † indicates significantly negative correlation (p < 0.05).
Model
Short
Medium
Long
Rel
Conj-1
Conj-2
Nested
RNN
0.003
0.089
0.442
0.258
0.379
0.489
0.325
RNN SUP (w/o shortcut)
0.003
0.089
0.216
0.153
0.176
0.203
0.172
RNN SUP
0.004
0.088
0.188
0.184
0.197
0.217
0.205
RNN ND (w/o shortcut)
0.003
0.089
0.190
0.172
0.191
0.220
0.173
RNN ND
0.003
0.114
0.276
0.239
0.278
0.318
0.316
LSTM
0.004
0.089
0.178
0.231
0.249
0.276
0.281
Appendix
Table 10: Coefficient of variations of PPLs among five different random seeds for each model and evaluation data. The value is averaged over 32 languages in each setting.
Many of the thousands of attested languages share common configurations of features, creating a spectrum from typologically very rare (e.g., object-verb-subject word order) or impossible languages to very common combinations of features (e.g., subject-object-verb word order). One central question is under what conditions such typological tendencies can be predicted, and specifically whether the learning bias of language models (LMs) is sufficient to reproduce such patterns. In this study, we add one dimensionality to such analysis -- the learning scenario for LMs -- to explore its interaction with the inductive bias of LMs. Specifically, as a first study, we examine the effect of curriculum learning (CL), as a developmentally motivated learning scenario, i.e., starting with simpler sentences rather than randomly-ordered input. We expand existing LM-based exploration (El-Naggar et al., 2025a,b) with a simple CL variant and find that CL substantially impacts the apparent inductive bias of LMs.
We systematically compare word order preferences in decoder-only language models across 192 artificial languages and typologically diverse natural languages. On artificial languages, models exhibit a left-branching preference that aligns with neither natural language universals nor human word order learning biases. On natural languages, monolingual models show no clear base word order bias at small scales, but as data grows, a preference for right-branching subject-verb-object (SVO) languages emerges while SOV falls behind despite being the most frequent order cross-linguistically. This SVO advantage extends to multilingual models and correlates with language resource level and data quality rather than word order. Thus, the same architecture exhibits opposite preferences on artificial and natural languages, establishing that word order biases observed in practice are data-driven. Since highly-resourced languages are overwhelmingly SVO, these biases risk gradually reducing word order diversity, particularly in languages that productively use multiple word orders, with the widespread adoption of LLMs.
Varvara Arzt, Allan Hanbury, Terra Blevins
Faculty of Informatics, TU Wien · Khoury College of Computer Sciences, Northeastern University
A central question in language acquisition is whether linguistic biases can emerge from general learning mechanisms operating over underdetermined input. Artificial Language Learning (ALL) studies have shown that human learners reliably generalize beyond the evidence provided, including by preferring scope-homomorphic noun phrase modifier orders. In this work, we investigate whether language models exhibit the same bias under similar conditions. We create a controlled learning environment in which models are trained on a corpus where all noun phrases containing multiple modifiers have been removed, eliminating direct evidence about modifier ordering, and are then evaluated on multiple modifier sentences. Across three model sizes, we find that they consistently prefer scope-homomorphic orders despite never observing them during training. These preferences vary in strength by modifier type. To investigate the source of these preferences, we examine noun-modifier association strength using pointwise mutual information (PMI). While PMI reflects known modifier-ordering patterns, it does not explain the models' ordering preferences. These findings demonstrate that LMs can recover human-like linguistic generalizations from impoverished input and provide a controlled framework for investigating the mechanisms underlying such biases.
Amanda Popadich, Shane Steinert-Threlkeld
Department of Linguistics, University of Washington