cs.CLFeb 13, 2026

Semantic Chunking and the Entropy of Natural Language

Authors: Weishun Zhong, Doron Sivan, Tankut Can, Mikhail Katkov, Misha Tsodyks

Organizations: School of Natural Sciences, Institute for Advanced Study, Princeton, NJ, 08540, USA · Department of Brain Sciences, Weizmann Institute of Science, Rehovot, 76100, Israel

Abstract

Humans and large language models can predict next letter or word from its prior context much better than random guessing, indicating strong redundancy of language viewed as a stochastic process. Quantitatively this redundancy was estimated by Shannon to be around 80%, which means that every letter of a printed English text conveys approximately 1 bit of information and not 4.8 bits that 27 letters (including spaces) could potentially carry. This estimate was later confirmed by using autoregressive token probabilies computed by large language models. However, the statistical organization of language that give rise to such a large redundancy remains unclear. Here we introduce a statistical framework of language linking its redundancy to the hierarchical semantic organization of text. To this end, we use large language models to recursively segment any given text into semantically coherent chunks, inducing a semantic tree'' that spans the whole range of text organization, beginning from its main idea to individual tokens (words). For a large corpus of texts of a particular type, say fiction stories, the resulting ensemble of semantic trees is characterized by specific statistical regularities, giving rise to a structural'' entropy rate defined in this study. Surprisingly, we discovered that for several datasets considered in this work, semantic tree entropy rate was quite close to LLM-measured quantity and exhibited a similar trend across corpus. In particular, simpler texts like children stories exhibit lower branching in their semantic trees and correspondingly lower entropy rates, whereas fiction and poetry exhibit progressively larger branching factors and greater entropy rates. These results suggest that hierarchical semantic organization of texts is an important factor in their overall information transmission rates.

Figures & tables

Explore similar work

CardsList
  1. Large Language Models Do Not Always Need Readable Language

    Jun 18, 2026Jiayi Zhu, Haoxuan Peng, Junxi Wang +3Readability

  2. Entropy of Ukrainian

    Apr 30, 2026Anton Lavreniuk, Mykyta Mudryi, Markiian ChakloshEntropyUkrainian