Organizations: School of Natural Sciences, Institute for Advanced Study, Princeton, NJ, 08540, USA · Department of Brain Sciences, Weizmann Institute of Science, Rehovot, 76100, Israel
Humans and large language models can predict next letter or word from its prior context much better than random guessing, indicating strong redundancy of language viewed as a stochastic process. Quantitatively this redundancy was estimated by Shannon to be around 80%, which means that every letter of a printed English text conveys approximately 1 bit of information and not 4.8 bits that 27 letters (including spaces) could potentially carry. This estimate was later confirmed by using autoregressive token probabilies computed by large language models. However, the statistical organization of language that give rise to such a large redundancy remains unclear. Here we introduce a statistical framework of language linking its redundancy to the hierarchical semantic organization of text. To this end, we use large language models to recursively segment any given text into semantically coherent chunks, inducing a semantic tree'' that spans the whole range of text organization, beginning from its main idea to individual tokens (words). For a large corpus of texts of a particular type, say fiction stories, the resulting ensemble of semantic trees is characterized by specific statistical regularities, giving rise to a structural'' entropy rate defined in this study. Surprisingly, we discovered that for several datasets considered in this work, semantic tree entropy rate was quite close to LLM-measured quantity and exhibited a similar trend across corpus. In particular, simpler texts like children stories exhibit lower branching in their semantic trees and correspondingly lower entropy rates, whereas fiction and poetry exhibit progressively larger branching factors and greater entropy rates. These results suggest that hierarchical semantic organization of texts is an important factor in their overall information transmission rates.
Figures & tables
Figure 1: Semantic tree entropy contributes significantly to the overall predictability of text. (a) Cumulative LLM surprisal, HN=−∑i=1NlogP(ti∣t<i) , estimated using Llama-3.3-70B [ 15 ] ) for 925 fiction stories from RedditStories [ 16 ] . Color indicates text length N . The blue dash-dotted line is a linear fit whose slope gives the LLM-based entropy-rate estimate hLLM . Slope of the red dashed line correspond to the entropy computed from the random tree ensemble at K=4 . Intercept of the red dashed line is the same as the blue dashed-dotted line. (b) Entropy rate hK of the random tree ensemble, as an approximation for the semantic trees, versus the maximum number of semantic chunks K per level. For K=4 , corresponding to the typical human working memory (WM) capacity, the predicted entropy rate lies within the range of Shannon’s estimate of approximately 1 bit per character for printed English (shaded band). This scale agreement motivates the hypothesis that hierarchical semantic organization contributes substantially to language entropy. (c) The chunking procedure applies recursive semantic segmentation on a text until reaching the single-token level, thereby parsing the document into a hierarchical tree of spans whose leaves are tokens (a “semantic tree”).
Figure 2: Random K -ary trees capture the chunk-size statistics of empirical semantic trees. (a) Empirical chunk-size distributions at level L=7 for 20 fiction stories from RedditStories [ 16 ] , compared with theoretical distributions PL(n∣N) computed from the random tree model for each text length N . (b) After normalizing chunk sizes by the root length N , empirical distributions pooled across 925 texts from RedditStories collapse onto the large- N scaling predictions fL at multiple levels L .
Table 1: Goodness-of-fit, measured by the average KL divergence ⟨DKL⟩ , across corpora [ 30 , 31 , 16 , 32 ] for different values of K . For each corpus and model ( Llama-4-Maverick [ 33 ] and DeepSeek-V3.1 [ 34 ] ), the optimal K⋆=argminK⟨DKL⟩ is highlighted in green.
Figure 3: Fluctuation vanishes asymptotically for both LLM- and tree-based estimates of text entropy. (a) Entropy-rate estimates from individual random tree realizations (simulated using Eq. ( 4 )) concentrate around the predicted value as N increases, consistent with the emergence of typical trees. (b) Per-token entropy estimates for 925 fiction stories from RedditStories [ 16 ] , computed in two ways: from LLM perplexity (green dots) and from the likelihood of empirical semantic trees obtained via chunking (blue dots). As N increases, both estimates fluctuate around the predicted entropy rate. (c) Data from (b) binned by text length ( N ) uniformly on a log-scale. Solid curves indicate the mean within each bin, and the shaded regions indicate one standard deviation.
Figure 4: Semantic tree entropy predicts entropy rates across corpora. Entropy rate of different corpora estimated in two ways: theoretical values correspond to hK (Eq. ( 6 )) evaluated at optimal K for each corpus (Table 1 ); hLLM are obtained from cross-entropy estimate (Eq. ( 1 )) on each corpus.
Figure 5: Semantic trees exhibit universal scaling across levels and corpora. (a) Theoretical scaling functions fL across levels L , shown on log–log axes. (b) When replotted using the O(1) lognormal variable x=(lns−μL)/σL , the curves in (a) collapse onto the universal N(0,1) distribution as L increases. (c) Empirical chunk size distributions f^L (same as in Fig. 2 (b), rebinned to be uniform in logs ) across levels L . (d) The empirical curves in (c) likewise collapse under the transformation x=(lns−μ^L)/σ^L , consistent with the theoretical prediction. (e) Empirical distributions from all four corpora collapse onto the same universal curve after standardization, suggesting that semantic trees share common scaling behavior despite differences in genre and effective branching factor. Each marker shows, for one corpus, the across-level mean histograms of the standardized variable as in (d), computed on a shared grid of 15 equal-width bins.
Figure S1: Entropy of trees. Theory and enumeration agrees for K=4 , and gives H(N)≈2.5 nats×N .
Figure S2: Entropy scaling . Numerics is computed with the Markov model, large K expansion is Eq. ( S.108 ).
Figure S3: Semantic tree of RedditStories 32721. The tree was originally plotted as a long vertical figure and is adapted here into five panels. Read the tree from top to bottom and left to right across panels (a)–(e). The root appears in panel (c) as the green node (248).
Figure S4: Semantic tree of Tiny Story (198810). The tree was originally plotted as a long vertical figure and is adapted here into three panels. Read the tree from top to bottom and left to right across panels (a)–(c). The root appears in panel (b) as the green node (135).
Figure S5: Semantic tree of Modern Poetry (8443). The tree was originally plotted as a long vertical figure and is adapted here into three panels. Read the tree from top to bottom and left to right across panels (a)–(c). The root appears in panel (b) as the green node (135).
Figure S6: Semantic chunking using DeepSeek-V3.1 . (a) Comparison between the theoretical and empirical chunk size distributions at an intermediate level (7) for 20 randomly selected texts from RedditStories of different lengths. (b) Comparison between the theoretical and empirical normalized chunk size distributions (pooled over 100 randomly selected texts from RedditStories ) for level 3-11. (c) The same normalized chunk size distribution at differrent levels as in (b) but shown on a log-log plot. (d) The normalized chunk size distribution at different levels collapse to a universal standard normal distribution upon the transformation x=(lns−μ^L)/σ^L . (e) The same data collapse is observed for different corpora.
Figure S7: Similar chunk-size statistics coexist with distinct linguistic structure. (a) Random chunking produces normalized chunk-size distributions similar to those of semantic chunking [main text Fig. 2 (b)]. (b) Exact clause recovery by semantic and random chunking. A clause is recovered when its complete span matches a single internal node at any tree level. Top: 26 Labov Stories with 1,633 labeled clauses [ 69 ] . Bottom: 100 Reddit Stories with 8,017 clauses segmented by GPT-5.4-mini , independently of the semantic trees generated using Llama-4-Maverick . Recovery fractions are pooled over clauses; random-tree results are averaged over 100 realizations per story. Error bars indicate 95% confidence intervals obtained by resampling stories.
Figure S8: Robustness of random tree scaling across chunking procedures. Normalized chunk-size distributions are shown on log–log axes for (a) random chunking, (b) semantic chunking using Llama-4-Maverick , and (c) semantic chunking using DeepSeek-V3.1 .
Evaluating whether large language models (LLMs) capture the structure of natural language beyond local fluency remains an open challenge. Existing evaluation methods, largely based on task performance or short-context behavior, provide limited insight into the long-range statistical organization of generated text. We propose a complementary evaluation framework based on repeated subsequences. By analyzing their distribution across scales and relating it to higher-order Rényi entropies, we probe how texts reuse previously established structure under finite-length conditions. Experiments on human-written texts and length-matched GPT-generated texts show that, while power-law models can describe restricted ranges of block length, the observed entropy growth is often equally or better characterized by logarithmic--power forms. Across datasets, natural language exhibits stable entropy-growth patterns over accessible ranges, with consistent average behavior despite variability across individual texts. In contrast, GPT-generated texts show systematic and statistically significant shifts in estimated exponents with model size. These results demonstrate that repeated-subsequence entropy provides a quantitative structural diagnostic that reveals systematic differences in long-range organization, distinguishing natural language from state-of-the-art LLM outputs beyond surface-level fluency.
Large language models (LLMs) are commonly prompted and interfaced with human-readable natural language, even when the intended reader is another model. This paper investigates whether semantic information can be encoded in compact, non-standard textual forms that sacrifice human readability while remaining recoverable by LLMs. We refer to this class of model-centric textual representations as BabelTele, approached here not as a fixed protocol but as an empirical probe into LLMs' capacity to generate and interpret such representations. Through readability diagnostics, model likelihood measures, human questionnaires, and downstream task evaluations, we find that BabelTele can substantially depart from ordinary natural language while preserving core semantics for instruction-tuned LLMs. As a task-agnostic representational paradigm, BabelTele demonstrates high information density, maintaining 99.5% semantic fidelity even when the text volume is condensed to 27.9% of its original length. We further evaluate its semantic robustness in cross-model transfer, agent memory, and multi-agent communication. Results suggest that BabelTele can reduce context overhead while generally maintaining reliable downstream performance, although its effectiveness depends on the compressor-reader pair and task setting. These findings indicate that human readability, natural-language typicality, and model-side semantic recoverability can be partially decoupled, opening a path toward model-native representations in future exploration of LLM systems.
Jiayi Zhu, Haoxuan Peng, Junxi Wang +3
Shanghai Jiao Tong University · The University of Sydney · Hefei University of Technology +2
In natural language processing, the entropy of a language is a measure of its unpredictability and complexity. The first study on this subject was conducted by Claude Shannon in 1951. By having participants predict the next character in a sentence, he was able to approximate the entropy of the English language. Several follow-up studies by other authors have since been conducted for English, and one for Hebrew. However, to date, Shannon's experiment has never been conducted for Ukrainian. In this paper, we perform this experiment for Ukrainian by recruiting 184 volunteers using social media channels. We rely on techniques used for English to approximate the entropy value of Ukrainian. The final result is an upper bound of Hupper≈1.201 bits per character. We compare this to the performance of current Large Language Models. The methods and code used are also documented and published, along with a discussion of the main challenges encountered.
Anton Lavreniuk, Mykyta Mudryi, Markiian Chaklosh
Polish-Japanese Academy of Information Technology · University of the National Education Commission in Kraków