Rethinking the Relationship between the Power Law and Hierarchical Structures
Organizations: RIKEN, Japan · National Institute for Japanese Language and Linguistics, Japan · The University of Tokyo, Japan · Georgetown University, USA · National Institute of Informatics, Japan
Abstract
Statistical analysis of corpora provides an approach to quantitatively investigate natural languages. This approach has revealed that several power laws consistently emerge across different corpora and languages, suggesting universal mechanisms underlying languages. In particular, the power-law decay of correlations has been interpreted as evidence of underlying hierarchical structures in syntax, semantics, and discourse. This perspective has also been extended beyond corpora produced by human adults, including child speech, birdsong, and chimpanzee action sequences. However, the argument supporting this interpretation has not been empirically tested in natural languages. To address this gap, the present study examines the validity of the argument for syntactic structures. Specifically, we test whether the statistical properties of parse trees align with the assumptions in the argument. Using English and Japanese corpora, we analyze the mutual information, deviations from probabilistic context-free grammars (PCFGs), and other properties in natural language parse trees, as well as in the PCFG that approximates these parse trees. Our results indicate that the assumptions do not hold for syntactic structures and that it is difficult to apply the proposed argument not only to sentences by human adults but also to other domains, highlighting the need to reconsider the relationship between the power law and hierarchical structures.
Figures & tables
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Corpus | Depth | Branching | Prop. Unary |
|---|---|---|---|
| BLLIP | |||
| NPCMJ |
| Tags | Part-of-Speech Tags and Phrase Labels |
|---|---|
| NP | NP NN NNP NNS CD PRP TMP POS QP NNPS EX NX FW |
| VP | VP VBD VB VBN VBG VBZ VBP TO |
| S | S S1 SBAR SINV FRAG SQ SBARQ RRC |
| PP | IN PP RP PRT |
| DT | DT PRP$ PDT |
| JJ | JJ ADJP JJR JJS NAC ADJ |
| Tags | Part-of-Speech Tags and Phrase Labels |
|---|---|
| NP | NP NUMCLP PRN NML PNLP NUMCLPSYM N NPR NUM CL ADJN PRO ADJI D Q FN WPRO PNL PRN WNUM WD FW |
| IP | IP VB AX AXD VB0 VB2 PASS PASS2 MD |
| PP | PP P |
| ADVP | ADVP ADV NEG WADV |
| CONJP | CONJP CONJ |
| PP | PP P |