cs.CLOct 6, 2026

Pseudowords as probes: Large Language Models show little of the sublexical sensitivity that governs human pseudoword processing

Authors: Jing Chen, Giulia Loca, Simona Amenta, Marco Marelli

Organizations: Department of Psychology, University of Milano-Bicocca, Piazza dell’Ateneo Nuovo 1, 20126 Milano, Italy · Department of Informatics, Systems and Communication – DISCo, University of Milano-Bicocca, Milano, Italy

Abstract

Systematicity, the probabilistic mapping of form to meaning, permeates language at all levels, and sublexical cues have been shown to govern human pseudoword processing. Yet whether LLMs exhibit comparable sensitivity to these cues remains unclear. We tested five LLMs on two Italian two-alternative forced-choice pseudoword experiments and compared their responses with a human behavioural baseline. LLMs aligned more reliably with humans when real-word options provided a lexical familiarity cue than in the pseudoword-only condition, where they fell substantially below fastText, a character-n-gram model. In addition, the sublexical cosine-similarity cue that reliably drove human--fastText agreement did not consistently transfer to human--LLM alignment, and reasoning-token expenditure bore no consistent relation to human processing difficulty. These findings suggest that LLMs do not necessarily share the sublexical cues that govern human pseudoword processing; we discuss tokenization and training-data coverage as candidate explanations.

Figures & tables

Explore similar work

CardsList
  1. Production and Perception in LLMs: A Token Probability Approach

    Jul 13, 2026Anna Marklová, Jiří Milička, Martina Vokáčová +1Single-Token Output DistributionsLinguistics

  2. Tracing the ongoing emergence of human-like reasoning in Large Language Models

    May 20, 2026Paolo Morosi, Nikoleta Pantelidou, Fritz Günther +2LLM Reasoning StrategiesEmergence

  3. Dual Alignment Between Language Model Layers and Human Sentence Processing

    Apr 20, 2026Tatsuki Kuribayashi, Alex Warstadt, Yohei Oseki +1Natural LanguageSurprisal