cs.CLAug 28, 2026

A Formal Limitation on Learning Human Language From Textual Corpora

Authors: Emily Cheng, Ryan Cotterell

Organizations: Universitat Pompeu Fabra · ETH Zürich

Abstract

Can a listener recover what a speaker means from the form of an utterance alone? We answer this question information-theoretically, and for a listener given by any featurizer of text, including the hidden states of contemporary large language models. Modeling language use as a joint distribution over meanings, contexts, and utterances, we derive upper bounds on the probability that a decoder recovers a speaker's intended meaning from a representation of the utterance. The bounds are governed by the uncertainty that form leaves about meaning, which splits into an irreducible part and a part that only (extralinguistic) context, but never the utterance alone, can resolve. Because these quantities are intrinsic to language, no representation, however much text or supervision produced it, can surpass them. The bounds apply, moreover, to meaning spaces that are discrete or continuous. We provide empirical evidence in support of the theory through experiments on artificial languages, Mandarin zero-pronoun resolution, and color reference.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. The Information Shadow: Measuring Structural Limits on What Language Models Can Learn

    Jul 17, 2026Priyansh Srivastava, Romit ChatterjeeFundamental LimitsLanguage Modeling

  2. On the Limits of Prompt-Conditioned Language Models as General-Purpose Learners

    Jun 22, 2026David Mguni, Julian Ma, Jun WangLarge Language Models Fail

  3. Implicit Representations of Grammaticality in Language Models

    May 6, 2026Yingshan Susan Wang, Linlu Qiu, Zhaofeng Wu +2GrammaticalityText Corpora