cs.CLMar 4, 2026

World Properties without World Models: Distributional Associations and the Interpretation of Decoding Results from Language Models

Authors: Elan Barenholtz

Organizations: Department of Psychology & Center for Complex Systems and Brain Sciences Florida Atlantic University

Abstract

A growing literature shows that variables can be linearly decoded from the activations of large language models (LLMs). These range from properties of the world, such as the locations of cities and the lifetimes of historical figures, to emotions and pain. Such findings are often taken as evidence that language models go beyond surface text statistics and form internal models of the world. We show that static word embeddings (fixed, context-insensitive representations learned from corpus statistics) of the same or matched stimuli support much of the same decoding. Across four published cases (place, time, pain and emotion), static vectors predict coordinates and year of death (R^2 = 0.42-0.59), separate pain from matched control sentences (held-out AUC 0.85-0.88), and classify twelve emotions in stories written to avoid naming them (AUC 0.84-0.88). Because static embeddings assign each word a single, context-independent vector, these results are a lower bound on what word associations alone can support. The LLMs retain clear advantages on representational tests, and causal and behavioral findings remain outside the scope of the baseline. On the original authors' entities, where we reproduce their Llama-2 results, the transformer's advantage lies mostly in placing historical figures in the right century and places in the right country, coarse sorting that richer word associations would be expected to improve; within those groups every representation orders items poorly. Static vectors for disambiguated Wikipedia entities, which carry the associations of a particular place or person rather than of the words in its name, close most of the remaining gap, matching Pythia-2.8B on coordinates and Llama-2-7B on year of death. These results indicate that decodability alone cannot distinguish a representation of a property from information already available in fixed distributional associations.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs

    Jun 25, 2026Sinie van der Ben, Raphaël Baur, Yannick Metz +1Valence-Arousal EstimationEmotion

  2. LMEnt: A Suite for Analyzing Knowledge in Language Models from Pretraining Data to Representations

    Sep 3, 2025Daniela Gottesman, Alon Gilae-Dotan, Ido Cohen +4Large Language Model PretrainingToken Co-Occurrence Graphs

  3. Tokens, the oft-overlooked appetizer: Large language models, the distributional hypothesis, and meaning

    Dec 14, 2024Julia Witte Zimmerman, Denis Hudon, Kathryn Cramer +9Single-Token Output DistributionsDistributional Information