We introduce STRUCTURALCOST, a self-paced reading dataset of 475 participants and 40,800 observations isolating the processing cost of long-distance subject-verb dependency resolution. We replicate a low-powered psycholinguistic finding at NLP scale, namely that human reading times at the main verb increase with dependency length, driven by syntactic embedding beyond linear distance. Different language models -- spanning n-gram models, SSMs, and transformers -- partially mirror this graded difficulty profile, yet underestimate the integration cost humans incur, with a gap that persists across architectures and model sizes. This suggests these models capture the predictive component of human processing but not the full integration cost that working memory imposes. STRUCTURALCOST provides data needed to drive progress toward evaluating the cognitive plausibility of language models.
Figures & tables
Condition
Dep.
Example Sentence
Critical sentences
Baseline
0
The violinist leaves the stage before the audience loudly applauds the great performance.
PP
4
The violinist in the large orchestra leaves the stage before the audience loudly applauds.
SRC
4
The violinist that followed the conductor leaves the stage before the audience loudly applauds.
ORC
4
The violinist that the conductor followed leaves the stage before the audience loudly applauds.
2 × SRC
9
The violinist that followed the conductor that worked in the orchestra leaves the stage.
Table 1: Example sets of critical and control sentences. Dep. (dependency length) = number of intervening words between the main subject and main verb . Critical sentences cross syntactic structure (SRC/ORC) with embedding depth (single/double), plus a 0-word baseline and a 4-word linear-distance control (PP). Control sentences manipulate linear distance via prepositional phrases, with no syntactic embedding, providing a reference class for dissociating distance from structural complexity.
Condition
Dep.
N
Mean RT
SD
Median RT
Cohen’s d (vs. baseline)
(words)
d
95% CI
Baseline
0
4,574
490
308
422
—
—
PP
4
4,562
512
321
435
−0.08
[−0.12,−0.04]
SRC
4
4,549
539
370
439
−0.15
[−0.19,−0.12]
ORC
4
4,496
583
408
467
−0.19
[−0.24,−0.15]
2 × SRC
9
4,545
504
302
436
−0.11
[−0.15,−0.08]
Table 2: RTs (ms) at the main verb by condition, with Cohen’s d relative to the baseline (0-word dependency).
Figure 1: Average log-transformed RTs by condition with 95% CIs (shaded ribbon). Curve colour indicates dependency length between subject noun (S) and main verb (V) : dark grey = 0-word; medium grey = 4-word; light grey = 9-word.
Figure 2: Fixed-effect estimates from M1 fitted separately to Natural Stories, UCL-SPR, and StructuralCost (at the integration site). All predictors are z -scored; estimates are on the log-RT scale. Error bars are 95% CIs.
Figure 3: Spearman correlations between model surprisal and reading times, by dependency length (0, 4, 9 words) and fit level: all words (left) vs. verb only (right). Error bars show 95% CIs. See Appendix F for the table.
M0 → M1
M1 → M2
Model
Global
Local
Global
Local
N -gram baselines
KenLM (GPT-2)
940∗∗∗
0.02
44∗∗∗
20.5∗∗∗
KenLM (Pythia)
993∗∗∗
0.08
44∗∗∗
20.6∗∗∗
GPT-2
Small
1082∗∗∗
27.8∗∗∗
33∗∗∗
8.0∗∗
Table 3: χ2(1) from likelihood ratio tests, globally and locally at the main verb. M0: baseline (word log frequency + word length); M1: + surprisal; M2: + dependency length. ∗p<.05 , ∗∗p<.01 , ∗∗∗p<.001 . Global here tests dependency length as a sentence-level covariate across all positions; Local restricts to the integration site.
Condition
Observed EOI (ms)
Predicted EOI (ms)
See Table 2
KenLM (GPT2)
Mamba-130M
Pythia-1.4B
Pythia-70M
PP
22
1.04
0.10
− 0.07
− 1.84
SRC
49
0.02
− 1.18
− 1.25
− 2.13
ORC
93
1.99
0.10
0.03
− 1.42
2 × SRC
14
− 0.21
− 1.49
− 1.54
− 3.20
2 × ORC
63
1.69
0.31
0.41
− 0.89
Table 4: Predicted versus observed EOIs (relative to Baseline) at the main verb. Predicted values apply spillover-adjusted conversion factors, estimated on Natural Stories, to each model’s surprisal on StructuralCost critical items; negative values indicate a predicted speedup. Models without a significant conversion factor are omitted.
Dataset
Method
Uniq. Sent.
Subj.
Subj./Item
C1
C2
C3
C4
Naturalistic datasets
Natural Stories
Futrell et al.,2021
SPR
485
181
∼ 100
✗
✗
✓
✓
Dundee
Kennedy et al.,2003
ET
2,368
10
10
✗
✗
✓
✗
Provo
Luke and Christianson,2018
ET
∼ 138
84
∼ 8
✗
✗
−
✗
UCL Corpus
Frank et al.,2013
SPR + ET
361 + 205
117 + 43
✓
✗
✓
MECO
Kuperman et al.,2024
ET
104
50
✗
✗
−
Table 5: Comparison of reading time datasets against four criteria required to study LLM replication of human processing difficulty (non-exhaustive list). C1 : isolated sentences; C2 : systematic dependency length manipulation; C3 : sufficient items for model evaluation; C4 : sufficient participants per item. ET = eye-tracking; SPR = self-paced reading. ✓ = satisfied; ✗ = not satisfied; − = partially satisfied; * = information not found.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Dep. 0
Dep. 4
Dep. 9
Total
Critical
10
30
20
60
Control
10
10
10
30
Total
20
40
30
90
Appendix
Table 6: Items per participant by dependency length.
OSpan score
N
%
0
38
8
1
0
0
2
32
7
3
84
18
4
131
28
5
116
24
Appendix
Table 7: Distribution of operation-span (OSpan) scores across participants. Protocol: Conway et al. (2005) ; implementation adapted from von der Malsburg’s Python version.
Spearman r[95% CI ]
Model
Global
Dep=0
Dep=4
Dep=9
N -gram
KenLM_G
.51 [.49,.53]
.11 [-.14,.35]
.14 [-.01,.28]
.00 [-.18,.19]
KenLM_P
.51 [.49,.53]
.16 [-.09,.39]
.15 [.01,.29]
.01 [-.17,.19]
GPT-2
Small
.52 [.50,.54]
.24 [-.02,.47]
.31 [.17,.45]
.34 [.16,.51]
Appendix
Table 8: Spearman correlations between model surprisal and mean per-word RT, globally and locally at the main verb by dependency length. Bold marks the single best local score. CIs computed via bootstrap ( k=1000 ).
A recent study (Kuribayashi et al., 2025) has shown that human sentence processing behavior, typically measured on syntactically unchallenging constructions, can be effectively modeled using surprisal from early layers of large language models (LLMs). This raises the question of whether such advantages of internal layers extend to more syntactically challenging constructions, where surprisal has been reported to underestimate human cognitive effort. In this paper, we begin by exploring internal layers that better estimate human cognitive effort observed in syntactic ambiguity processing in English. Our experiments show that, in contrast to naturalistic reading, later layers better estimate such a cognitive effort, but still underestimate the human data. This dual alignment sheds light on different modes of sentence processing in humans and LMs: naturalistic reading employs a somewhat weak prediction akin to earlier layers of LMs, while syntactically challenging processing requires more fully-contextualized representations, better modeled by later layers of LMs. Motivated by these findings, we also explore several probability-update measures using shallow and deep layers of LMs, showing a complementary advantage to single-layer's surprisal in reading time modeling.
Tatsuki Kuribayashi, Alex Warstadt, Yohei Oseki +1
Probing has shown that language model representations encode rich linguistic information, but it remains unclear whether they also capture cognitive signals about human processing. In this work, we probe language model representations for human reading times. Using regularized linear regression on two eye-tracking corpora spanning five languages (English, Greek, Hebrew, Russian, and Turkish), we compare the representations from every model layer against scalar predictors -- surprisal, information value, and logit-lens surprisal. We find that the representations from early layers outperform surprisal in predicting early-pass measures such as first fixation and gaze duration. The concentration of predictive power in the early layers suggests that human-like processing signatures are captured by low-level structural or lexical representations, pointing to a functional alignment between model depth and the temporal stages of human reading. In contrast, for late-pass measures such as total reading time, scalar surprisal remains superior, despite its being a much more compressed representation. We also observe performance gains when using both surprisal and early-layer representations. Overall, we find that the best-performing predictor varies strongly depending on the language and eye-tracking measure.
Eleftheria Tsipidi, Samuel Kiegeland, Francesco Ignazio Re +4
ETH Zürich · Toyota Technological Institute at Chicago · University College London
Human language comprehension unfolds sequentially: each word is processed in the context of those that came before, and the interpretation builds incrementally over time. Surprisal, the negative log probability of a word given its context, has been the dominant predictor of incremental processing cost. But surprisal reduces rich sequential representations to a single scalar at each word, discarding information about the direction in which the interpretation has been evolving. Dynamical-systems approaches suggest that the trajectory of the evolving interpretive state, not just its position at each moment,should shape processing, and language itself may have local momentum, since speakers plan utterances a few words at a time. We introduce trajectory extrapolation error: at each word, we fit a linear trajectory to the preceding hidden states of a transformer language model and measure deviation from the extrapolated path. On the Natural Stories corpus, this measure is nearly orthogonal to surprisal (r = .044) and independently predicts self-paced reading times. The effect is especially pronounced in garden-path sentences, strengthens with model scale (GPT-2 Small to Large), and replicates across architectures with different positional encoding schemes (GPT-2 vs. Pythia/RoPE). A displacement control shows the effect is not reducible to representational change magnitude: displacement and extrapolation error predict in opposite directions. These findings reveal two dissociable components of processing cost: word-level prediction error (surprisal) and sensitivity to the local momentum of the unfolding interpretation (trajectory extrapolation error).
Elan Barenholtz
Machine Perception & Cognitive Robotics Laboratory Department of Psychology / Center for Complex Systems Florida Atlantic University