cs.CLSep 28, 2026

LLMs are not stochastic parrots: Evidence for meaning-mediated abstraction from conlang-like tasks

Authors: Julia Witte Zimmerman, Calla G. Beauregard, Tabia Tanzin Prama, Parisa Suchdev, Kathryn Cramer, Elisabeth Kollrack

Organizations: Computational Story Lab · Computational Ethics Lab · Vermont Complex Systems Institute University of Vermont Burlington, VT 05405, USA

Abstract

The strong version of the stochastic parrot argument claims that, although large language models (LLMs) may exceed rote regurgitation, they cannot move beyond statistical pattern matching into abstraction or reasoning, remaining ontologically near the lower bound of pattern reuse despite producing alluringly fluent text. We test this hypothesis using conlang-like tasks. Several LLMs are given only natural-language descriptions of fictional languages that subvert prominent superficial patterns in training data by combining statistically uncommon and unattested features. Crucially, no example outputs are given. We argue that if the models exhibit rule-following behaviour, they cannot be relying solely on superficial statistical patterns; such patterns often work against the correct output. Instead, successful performance requires representations of the constraints specified in the prompt. Across three complementary task families, models systematically move in the meaning-predicted direction: they distinguish prompt exposure from instructed use, alter semantic relationships in response to novel constraints, and sometimes produce exact matches to complex translation answer keys. Although performance varies across the spectrum of models used, these results provide evidence for meaning-mediated abstraction in LLMs and refute the strong stochastic parrot hypothesis. Our work shows that, under appropriate architectural and contextual constraints, statistical learning can produce meaning-mediated abstractions, although generation remains strongly constrained by superficial plausibility. We discuss implications for model development and for understanding how increasingly abstract representations may emerge from plausible-text-generation objectives.

Figures & tables

Appendix figures & tables27 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. An Incomplete Loop: Deductive, Inductive, and Abductive Reasoning in Language Models

    Apr 3, 2024Emmy Liu, Graham Neubig, Jacob AndreasLLM Reasoning StrategiesDeduction

  2. On the Limits of Prompt-Conditioned Language Models as General-Purpose Learners

    Jun 22, 2026David Mguni, Julian Ma, Jun WangLarge Language Models Fail

  3. Don't Let Me Ask for It: LLMs Show Deficiencies in Active Multi-Turn Information Acquisition for Abductive Inference

    Aug 4, 2026Shahrukh Mohiuddin, Chalamalasetti Kranti, Sherzod Hakimov +1Abductive ReasoningLarge Language Models Fail