Rethinking the Relationship between the Power Law and Hierarchical Structures
Authors: Kai Nakaishi, Ryo Yoshida, Kohei Kajikawa, Koji Hukushima, Yohei Oseki
Organizations: RIKEN, Japan · National Institute for Japanese Language and Linguistics, Japan · The University of Tokyo, Japan · Georgetown University, USA · National Institute of Informatics, Japan
Statistical analysis of corpora provides an approach to quantitatively investigate natural languages. This approach has revealed that several power laws consistently emerge across different corpora and languages, suggesting universal mechanisms underlying languages. In particular, the power-law decay of correlations has been interpreted as evidence of underlying hierarchical structures in syntax, semantics, and discourse. This perspective has also been extended beyond corpora produced by human adults, including child speech, birdsong, and chimpanzee action sequences. However, the argument supporting this interpretation has not been empirically tested in natural languages. To address this gap, the present study examines the validity of the argument for syntactic structures. Specifically, we test whether the statistical properties of parse trees align with the assumptions in the argument. Using English and Japanese corpora, we analyze the mutual information, deviations from probabilistic context-free grammars (PCFGs), and other properties in natural language parse trees, as well as in the PCFG that approximates these parse trees. Our results indicate that the assumptions do not hold for syntactic structures and that it is difficult to apply the proposed argument not only to sentences by human adults but also to other domains, highlighting the need to reconsider the relationship between the power law and hierarchical structures.
Figures & tables
Figure 1: Sequential and structural distances. (a) If trees are balanced, the sequential distance grows exponentially with the structural distance, rseq∼exp(μrstr) . (b) If trees are strongly biased, the growth is slower. For example, the relation is linear, rseq∼rstr .
Figure 2: Example of a parse tree from BLLIP, preprocessed as described in Section 3.1 . This tree has a strong right-branching bias.
Figure 3: CFIB J is the MI between the children of two nodes whose categories are fixed.
Figure 4: Estimation and fitting of the MI for (a) the exponential model with λ=0.1 and (b) the power-law model with α=2 . Fitting was performed for Ndata=8×106 .
Figure 5: MI IPOS between POS tags as a function of the sequential distance rseq . The fitted exponential and power-law decays for Ndata=2.56×106 are also presented.
Figure 6: MI Itag between POS or phrasal tags as a function of the structural distance rstr . The fitted exponential and power-law decays for Ndata=2.56×106 are also presented.
Figure 7: Number of POS tag pairs such that the structural and sequential distances are rstr and rseq , respectively. Black circles represent the average sequential distance for each structural distance.
Figure 8: CFIB J for (x0,x1)=(NP,NP) as a function of the structural distance rstr . The fitted exponential and power-law decays for Ndata=2.56×106 are also presented.
Figure 9: Statistical properties of syntactic structures in WikiText and NPCMJ. (a1–4) Results for WikiText: (a1) IPOS as a function of rseq ; (a2) Itag as a function of rstr ; (a3) relationship between rseq and rstr ; and (a4) CFIB J for (x0,x1)=(NP,NP) as a function of rstr . Fitting was performed with Ndata=2.56×106 . (b1–b4) Corresponding results for NPCMJ. Fitting was performed with Ndata=6.4×105 for IPOS and Itag , and with Ndata=4.0×104 for J .
Figure 10: Statistical properties of trees generated by the PCFG obtained from the maximum likelihood estimation of syntactic structures. (a) MI IPOS between POS tags. (b) MI Itag between POS or phrasal tags. (c) Distribution and average of sequential distances. (d) CFIB J for (x0,x1)=(NP,NP) . Fitting was performed for Ndata=8×106 in (a) and (b), and for rstr≤10 in (a).
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Corpus
Depth
Branching
Prop. Unary
BLLIP
11.598(±4.354)
1.527(±0.904)
0.222
NPCMJ
6.642(±3.336)
1.547(±1.282)
0.271
Appendix
Table 1: Tree-shape statistics computed on the original parse trees in BLLIP and NPCMJ. The depth and the branching factor are reported as mean ± standard deviation across sentences.
Tags
Part-of-Speech Tags and Phrase Labels
NP
NP NN NNP NNS CD PRP TMP POS QP NNPS EX NX FW
VP
VP VBD VB VBN VBG VBZ VBP TO
S
S S1 SBAR SINV FRAG SQ SBARQ RRC
PP
IN PP RP PRT
DT
DT PRP$ PDT
JJ
JJ ADJP JJR JJS NAC ADJ
Appendix
Table 2: Eleven tags used for BLLIP and WikiText and the corresponding part-of-speech tags and phrase labels.
Tags
Part-of-Speech Tags and Phrase Labels
NP
NP NUMCLP PRN NML PNLP NUMCLPSYM N NPR NUM CL ADJN PRO ADJI D Q FN WPRO PNL PRN WNUM WD FW
IP
IP VB AX AXD VB0 VB2 PASS PASS2 MD
PP
PP P
ADVP
ADVP ADV NEG WADV
CONJP
CONJP CONJ
PP
PP P
Appendix
Table 3: Seven tags used for NPCMJ and the corresponding part-of-speech tags and phrase labels.
Figure 11: (a1, 2) Results for BLLIP under the unbinarized setting: (a1) MI Itag between POS or phrasal tags, where fitting was performed for Ndata=2.56×106 ; (a2) distribution and average of sequential distances. (b1, 2) Results for BLLIP under the phrasal-tag-only setting: (b1) MI Itag between phrasal tags, where fitting was performed for Ndata=2.56×106 ; (a2) CFIB J , where fitting was pwerformed for Ndata=6.4×105 . (c1, 2) Results for NPCMJ under the unbinarized setting: (a1) MI Itag between POS or phrasal tags, where fitting was performed for Ndata=2.56×106 ; (a2) distribution and average of sequential distances.
Figure 12: Results for BLLIP obtained with different estimators. (a) MI IPOS between POS tags. (b) MI Itag between POS or phrasal tags. (c) CFIB J for (x0,x1)=(NP,NP) . In all cases, Ndata=2.56×106 .
Figure 13: Results for the PCFG obtained with different estimators. (a) MI IPOS between POS tags. (b) MI Itag between POS or phrasal tags. In both (a) and (b), Ndata=8×106 . (c-e) CFIB J for (x0,x1)=(NP,NP) obtained with (c) MM, and (d) NSB, and (e) ZA.
Figure 14: CFIB J for BLLIP with (a) (x0,x1)=(NP,VP) and (b) (x0,x1)=(VP,VP) . Fitting was performed for Ndata=2.56×106 in (a) and Ndata=6.4×105 in (b).
Despite their widespread use, the principles governing the organisation of syntactic dependency trees remain poorly understood. I analyse dependency trees from 124 typologically, genetically, and geographically diverse languages. Their topology departs systematically from randomness. Relative to uniformly sampled random trees, dependency trees exhibit greater structural robustness and lower branching heterogeneity. I propose that these universal regularities emerge naturally from incremental grammatical encoding. I model this process using sublinear preferential attachment. The model accurately reproduces the observed topology. More generally, the results demonstrate how a universal statistical property of syntax can emerge from a simple, cognitively motivated generative process. They further illustrate a broader principle of efficiency by construction: communicatively efficient syntactic structures can emerge without direct optimisation for communication.
Fermín Moscoso del Prado Martín
Department of Language and Communication & Center for Language Studies Radboud University Erasmuslaan 1, 6525 NL, Nijmegen, The Netherlands
The probabilities of syntactic structures in human languages are assumed to emerge fully from language-specific experience. Here, I show that a universal prior over syntactic structures emerges from a model of human language production, in which words are progressively integrated into syntactic structure. Without fitting any parameters to specific language data, the resulting prior assigns higher probabilities to attested than to random dependency trees in all 138 typologically diverse languages examined. These prior probabilities correlate positively with those estimated from corpora in 33 of 34 languages. The results indicate that part of the probability structure of syntax can arise independently of language-specific learning. This identifies human language production as a possible cognitive source of universal statistical structure in language, while providing a data-independent structural bias for probabilistic models, including large language models.
Fermín Moscoso del Prado Martín
Department of Computer Science and Technology, University of Cambridge, United Kingdom
Understanding how the structure of language can be learned from sentences alone is a central question in both cognitive science and machine learning. Studies of the internal representations of Large Language Models (LLMs) support their ability to parse text when predicting the next word, while representing semantic notions independently of surface form. Yet, which data statistics make these feats possible, and how much data is required, remain largely unknown. Probabilistic context-free grammars (PCFGs) provide a tractable testbed for studying these questions. However, prior work has focused either on the post-hoc characterization of the parsing-like algorithms used by trained networks; or on the learnability of PCFGs with fixed syntax, where parsing is unnecessary. Here, we (i) introduce a tunable class of PCFGs in which both the degree of ambiguity and the correlation structure across scales can be controlled; (ii) provide a learning mechanism -- an inference algorithm inspired by the structure of deep convolutional networks -- that links learnability and sample complexity to specific language statistics; and (iii) validate our predictions empirically across deep convolutional and transformer-based architectures. Overall, we propose a unifying framework where correlations at different scales lift local ambiguities, enabling the emergence of hierarchical representations of the data.
Jack T. Parley, Francesco Cagnetta, Matthieu Wyart
Institute of Physics, École Polytechnique Fédérale de Lausanne (EPFL), Lausanne, Switzerland · Theoretical and Scientific Data Science, SISSA, Trieste, Italy · Department of Physics and Astronomy, Johns Hopkins University, Baltimore, Maryland.