cs.CLMay 7, 2026

The Frequency Confound in Language-Model Surprisal and Metaphor Novelty

Authors: Omar MomenSina Zarrieß

Organizations: CRC 1646 – Linguistic Creativity in Communication · Bielefeld University, Germany

Abstract

Language-model (LM) surprisal is widely used as a proxy for contextual predictability and has been reported to correlate with metaphor novelty judgments. However, surprisal is tightly intertwined with lexical frequency. We explore this interaction on metaphor novelty ratings using two different word frequency measures. We analyse surprisal estimates from eight Pythia model sizes and 154 training checkpoints. Across settings, word frequency is a stronger predictor of metaphor novelty than surprisal. Across training stages, the surprisal--novelty association peaks at an early stage and then falls again, mirroring a similarly timed increase in the surprisal--frequency association. These results suggest that the often-reported optimal LM surprisal settings may incorrectly associate contextual predictability with metaphor novelty and processing difficulty, whereas lexical frequency may be the major underlying factor.

Explore similar work

CardsList
  1. surprisal is Not a Theory

    Jul 22, 2026Andrés Buxó-Lugo, Aniello De Santo, Morgan Grobol +2SurprisalLinguistics

  2. On the Proper Treatment of Units in Surprisal Theory

    Apr 30, 2026Samuel Kiegeland, Vésteinn Snæbjarnarson, Tim Vieira +1SurprisalTheoretical Foundations