Some properties of languages, e.g., subject-object-verb (SOV) word order, are more prevalent than others among the thousands of attested natural languages (NLs). Such typological commonality is often attributed to learning biases. Computational simulations, recently with language models (LMs), have facilitated the exploration of this theory. In this paper, we extend existing analyses of the relationship between LMs' learning biases and typological commonality on both data and model sides, focusing on: (i) cross-serial dependencies, the upper limit of attested syntactic complexity, and (ii) stack-based LMs (SLMs), potentially facilitating learning of hierarchical patterns. We first evaluate generalization of SLMs on cross-serial dependencies across diverse artificial languages and confirm that they struggle with such constructions. However, SLMs with limited working memory generalize better suggesting a possible basis for such inductive bias and thus the typological commonality of some word order configurations.
Figures & tables
GCG Lexical Syntactic Category El-Naggar et al. (2025a) ; El-Naggar et al. (2025b)
Verb with Complement (VCOMP) – (S ∣1 NP SUBJ ) ∣1 SCOMP
Kim ga believed that Sandy o lied
Table 1: Lexical syntactic categories, their derivations, and their examples, where a word corresponding to the category is in bold. The top half shows the categories defined by El-Naggar et al. (2025a) ; El-Naggar et al. (2025b) . The bottom half shows the new categories that we added to extend the grammar. The vertical bars “ ∣ ” in the GCG lexical syntactic categories represent either forward or backward slashes, and those with the same index are controlled by the same word order parameter (see Table 2 ).
Digit
0
1
1 (Base)
SOV
VOS
2 (COMP)
Preposed complementizer
Postposed complementizer
3 (PP)
Postposition
Preposition
4 (ADJ)
Prenominal adjective
Postnominal adjective
5 (REL)
Preposed relativizer
Postposed relativizer
Table 2: Word order parameters and the constructions introduced by their assignment (0/1). The exact implementation is shown in Table 1 .
Figure 1: Partial dependency representations of CSDs with embedded relative clauses. The examples are for SOV with prenominal relative clauses ( 0XXX0 ); the pattern will be, for example, fully mirrored for VOS with postnominal relative clauses ( 1XXX1 ). Colored arcs indicate cross-serial noun–verb dependencies. Black arcs indicate relative-clause attachment. Relative clauses are bracketed.
Figure 2: PPL distribution on the Medium test set (average ± standard deviation). The bar shows the standard deviation among five runs with different random seeds. Note that all training was conducted and stopped under the same early-stopping setting, and the high variance indicates training instability across different random seeds (see section 6 ).
Model
Short
Medium
Long
Rel
Conj-1
Conj-2
Nested
Avg.
RNN
33.40 ± 0.14
100.69 ± 13.91
345.68 ± 510.02
376.50 ± 223.62
552.94 ± 483.43
731.62 ± 1033.28
538.97 ± 574.27
382.83
RNN SUP
33.32 ± 0.18
88.29 ± 16.40
249.02 ± 75.80
299.37 ± 97.15
404.95 ± 126.49
485.74 ± 160.56
437.99 ± 171.39
285.53
RNN ND
33.43 ± 0.13
94.98 ± 16.35
333.45 ± 134.35
328.11 ± 101.49
497.27 ± 165.78
615.10 ± 241.20
504.19 ± 245.97
343.79
LSTM
33.49 ± 0.22
72.72 ± 13.04
313.71 ± 86.15
277.15 ± 100.83
501.66 ± 208.33
688.17 ± 305.38
522.85 ± 257.08
344.25
LSTM SUP
33.49 ± 0.20
73.80 ± 11.75
415.24 ± 190.50
364.23 ± 154.01
661.19 ± 298.72
1057.02 ± 601.37
772.05 ± 377.53
482.43
LSTM ND
34.52 ± 7.51
88.03 ± 24.14
516.53 ± 226.21
422.71 ± 188.28
773.97 ± 388.83
1159.86 ± 664.59
882.36 ± 526.66
554.00
Table 3: PPL for each model and evaluation data (average ± standard deviation). The standard deviation here indicates how the average PPL varies among 32 languages. PPL values obtained from SLMs, which are better than those of the corresponding base model, are shown in bold . The best PPLs among the LMs with the same base architecture are shown with an underline .
Model
Short
Med.
Long
Rel
Conj1
Conj2
Nest
RNN
0.57
− 0.66 †
− 0.57 †
− 0.31 †
− 0.48 †
− 0.59 †
− 0.25
RNN SUP
− 0.24
− 0.58 †
− 0.43 †
− 0.60 †
− 0.62 †
− 0.45 †
− 0.39 †
RNN ND
0.36
− 0.40 †
− 0.24
− 0.61 †
− 0.26
− 0.12
− 0.36 †
LSTM
0.43
− 0.56 †
− 0.69 †
− 0.57 †
− 0.26
− 0.28
− 0.38 †
LSTM SUP
0.54
− 0.60 †
− 0.66 †
− 0.49 †
− 0.48 †
− 0.56 †
− 0.55 †
LSTM ND
0.59
− 0.25
0.17
− 0.24
0.10
0.20
0.12
Table 4: Spearman Correlation to Word Order Frequency by Model and Test/Evaluation Type. † indicates a significantly negative correlation (p < 0.05).
Feature
Coefficient
Significance
β0
0.0371
βF
− 0.0651
βM(LSTM)
− 0.5231
βM(TF)
− 0.0628
βS(Sup)
− 0.0213
βS(Nd)
− 0.0790
Table 5: Fixed effects from the regression model predicting perplexity. The word order digit effects ( βd , βd:M , βd:S ) are excluded from the Table. Coefficients for each parameter βeffect are shown. Significance markers denote: *** p<0.001, ** p<0.01, * p<0.05.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Artifact
License
Usage
NLTK Bird et al. (2009)
Apache License 2.0
to create ALs and targeted evaluation data
Docker Merkel and others (2014)
Apache 2.0
to train SLMs
El-Naggar et al. (2025b) Datasets
Creative Commons CC-BY 4.0
training and testing SLMs
Codes from DuSell and Cotterell (2025) ( https://github.com/bdusell/bearing-syntactic-fruit/commit/ebe4ee8 )
MIT License
to implement models
Rau ( https://github.com/bdusell/rau )
MIT Licence
to implement models
WALS Dryer and Haspelmath (2013)
Creative Commons CC-BY 4.0
to find word order statistics in NLs
Appendix
Table 6: Artifacts used in this paper.
S
82A Order of Subject and Verb Dryer (2013e)
VP
83A Order of Object and Verb Dryer (2013c)
O
81A Order of Subject, Object and Verb Dryer (2013f)
COMP
Feature GB421: Is there a preposed complementizer in complements of verbs of thinking and/or knowing? Skirgård et al. (2023)
Feature GB422: Is there a postposed complementizer in complements of verbs of thinking and/or knowing? Skirgård et al. (2023)
PP
85A Order of Adposition and Noun Phrase Dryer (2013b)
ADJ
87A Order of Adjective and Noun Dryer (2013a)
Appendix
Table 7: WALS and Grambank chapters used in this paper.
Model
Short
Medium
Long
Rel
Conj-1
Conj-2
Nested
Avg.
RNN
33.40 ± 0.14
100.69 ± 13.91
345.68 ± 510.02
376.50 ± 223.62
552.94 ± 483.43
731.62 ± 1033.28
538.97 ± 574.27
382.83
RNN SUP (w/o SC)
33.44 ± 0.18
100.43 ± 13.04
292.64 ± 88.90
352.21 ± 127.53
506.62 ± 189.70
623.84 ± 258.64
533.75 ± 249.39
348.99
RNN ND (w/o SC)
33.45 ± 0.14
100.40 ± 12.66
271.09 ± 86.73
324.98 ± 89.09
428.46 ± 117.36
492.40 ± 150.42
447.41 ± 147.36
299.74
LSTM
33.49 ± 0.22
72.72 ± 13.04
313.71 ± 86.15
277.15 ± 100.83
501.66 ± 208.33
688.17 ± 305.38
522.85 ± 257.08
344.25
LSTM SUP (w/o SC)
33.45 ± 0.18
75.78 ± 12.06
374.29 ± 133.49
301.13 ± 112.04
574.42 ± 283.83
850.53 ± 507.61
615.52 ± 351.33
403.59
LSTM ND (w/o SC)
33.48 ± 0.21
79.94 ± 14.51
359.74 ± 105.87
307.96 ± 106.65
537.33 ± 202.30
753.00 ± 314.25
581.48 ± 258.67
378.99
Appendix
Table 8: PPL for each dataset and evaluation data (average ± standard deviation). The standard deviation here indicates how the average PPL varies among 32 languages. “SC” denotes the shortcut connection. PPL values obtained from SLMs, which are better than those of the corresponding base model, are shown in bold . The best PPLs among the LMs with the same base architecture are shown with an underline .
Model
Short
Medium
Long
Rel
Conj-1
Conj-2
Nested
RNN
0.57
− 0.66 †
− 0.57 †
− 0.31 †
− 0.48 †
− 0.59 †
− 0.25
RNN SUP (w/o shortcut)
0.52
0.04
− 0.07
0.15
0.30
0.39
0.29
RNN ND (w/o shortcut)
0.36
0.18
− 0.31 †
− 0.05
0.03
0.22
0.20
LSTM
0.43
− 0.56 †
− 0.69 †
− 0.57 †
− 0.26
− 0.28
− 0.38 †
LSTM SUP (w/o shortcut)
0.55
− 0.65 †
− 0.56 †
− 0.68 †
− 0.38 †
− 0.34 †
− 0.48 †
LSTM ND (w/o shortcut)
0.61
− 0.53 †
− 0.58 †
− 0.67 †
− 0.39 †
− 0.33 †
− 0.44 †
Appendix
Table 9: Spearman Correlation to Word Order Frequency by Model and Test/Evaluation Type. † indicates significantly negative correlation (p < 0.05).
Model
Short
Medium
Long
Rel
Conj-1
Conj-2
Nested
RNN
0.003
0.089
0.442
0.258
0.379
0.489
0.325
RNN SUP (w/o shortcut)
0.003
0.089
0.216
0.153
0.176
0.203
0.172
RNN SUP
0.004
0.088
0.188
0.184
0.197
0.217
0.205
RNN ND (w/o shortcut)
0.003
0.089
0.190
0.172
0.191
0.220
0.173
RNN ND
0.003
0.114
0.276
0.239
0.278
0.318
0.316
LSTM
0.004
0.089
0.178
0.231
0.249
0.276
0.281
Appendix
Table 10: Coefficient of variations of PPLs among five different random seeds for each model and evaluation data. The value is averaged over 32 languages in each setting.