cs.CLFeb 4, 2026

LLM surprisal is necessary but not sufficient to capture English garden-path effects: Evidence from joint latent modeling of reading paradigms

Authors: Dario Paape, Tal Linzen, Shravan Vasishth

Organizations: Department of Linguistics, University of Potsdam · Department Linguistik, Universität Potsdam, Karl-Liebknecht-Straße 24–25, D-14476, Germany. · Center for Data Science/Department of Linguistics, New York University

Abstract

Temporarily ambiguous garden-path sentences ("While the team trained the striker wondered... ") are known to cause processing difficulty, which can manifest itself in a variety of reading behaviors (in-situ slowdowns, rereading), as well as in miscomprehension or outright rejection of the sentence as ungrammatical. Which types of reading behavior are observed critically depends on the experimental method used to collect the data, which makes comparing results between reading paradigms difficult. To address this problem, we present a latent-process multinomial processing tree (MPT) model of human reading and comprehension/judgment behavior in garden-path sentences that we fit to combined data from four different reading paradigms (eye tracking, uni- and bidirectional self-paced reading, Maze). The model distinguishes between the probability of adopting an incorrect initial analysis, the cost of encountering an incompatible continuation, and the cost of syntactic reanalysis. By taking into account trials with inattentive reading, more realistic estimates of the cost parameters are obtained. Cross-validation reveals that the MPT model has a better predictive fit to human reading patterns and end-of-trial task data than a model based solely on LLM-derived surprisal values. We also test several models that assume an influence of surprisal within the MPT architecture, and find that adding surprisal as an additional predictor or reading time and/or garden-path cost further improves predictive fit.

Figures & tables

Explore similar work

CardsList
  1. An Existence Proof for Neural Language Models That Can Explain Garden-Path Effects via Surprisal

    Apr 20, 2026Ryo Yoshida, Shinnosuke Isono, Taiga Someya +2SurprisalLinguistics

  2. Why are language models less surprised than humans? Testing the Parse Multiplicity Mismatch Hypothesis

    May 14, 2026William Timkey, Brian Dillon, Tal LinzenSurprisalNatural Language

  3. Syntactic Belief Update as the Driver of Garden Path Processing Difficulty

    Jun 25, 2026Alan Zhou, Miloš Stanojević, John T. HaleSurprisalSyntactic Structure