We introduce STRUCTURALCOST, a self-paced reading dataset of 475 participants and 40,800 observations isolating the processing cost of long-distance subject-verb dependency resolution. We replicate a low-powered psycholinguistic finding at NLP scale, namely that human reading times at the main verb increase with dependency length, driven by syntactic embedding beyond linear distance. Different language models -- spanning n-gram models, SSMs, and transformers -- partially mirror this graded difficulty profile, yet underestimate the integration cost humans incur, with a gap that persists across architectures and model sizes. This suggests these models capture the predictive component of human processing but not the full integration cost that working memory imposes. STRUCTURALCOST provides data needed to drive progress toward evaluating the cognitive plausibility of language models.
Figures & tables
Condition
Dep.
Example Sentence
Critical sentences
Baseline
0
The violinist leaves the stage before the audience loudly applauds the great performance.
PP
4
The violinist in the large orchestra leaves the stage before the audience loudly applauds.
SRC
4
The violinist that followed the conductor leaves the stage before the audience loudly applauds.
ORC
4
The violinist that the conductor followed leaves the stage before the audience loudly applauds.
2 × SRC
9
The violinist that followed the conductor that worked in the orchestra leaves the stage.
Table 1: Example sets of critical and control sentences. Dep. (dependency length) = number of intervening words between the main subject and main verb . Critical sentences cross syntactic structure (SRC/ORC) with embedding depth (single/double), plus a 0-word baseline and a 4-word linear-distance control (PP). Control sentences manipulate linear distance via prepositional phrases, with no syntactic embedding, providing a reference class for dissociating distance from structural complexity.
Condition
Dep.
N
Mean RT
SD
Median RT
Cohen’s d (vs. baseline)
(words)
d
95% CI
Baseline
0
4,574
490
308
422
—
—
PP
4
4,562
512
321
435
−0.08
[−0.12,−0.04]
SRC
4
4,549
539
370
439
−0.15
[−0.19,−0.12]
ORC
4
4,496
583
408
467
−0.19
[−0.24,−0.15]
2 × SRC
9
4,545
504
302
436
−0.11
[−0.15,−0.08]
Table 2: RTs (ms) at the main verb by condition, with Cohen’s d relative to the baseline (0-word dependency).
Figure 1: Average log-transformed RTs by condition with 95% CIs (shaded ribbon). Curve colour indicates dependency length between subject noun (S) and main verb (V) : dark grey = 0-word; medium grey = 4-word; light grey = 9-word.
Figure 2: Fixed-effect estimates from M1 fitted separately to Natural Stories, UCL-SPR, and StructuralCost (at the integration site). All predictors are z -scored; estimates are on the log-RT scale. Error bars are 95% CIs.
Figure 3: Spearman correlations between model surprisal and reading times, by dependency length (0, 4, 9 words) and fit level: all words (left) vs. verb only (right). Error bars show 95% CIs. See Appendix F for the table.
M0 → M1
M1 → M2
Model
Global
Local
Global
Local
N -gram baselines
KenLM (GPT-2)
940∗∗∗
0.02
44∗∗∗
20.5∗∗∗
KenLM (Pythia)
993∗∗∗
0.08
44∗∗∗
20.6∗∗∗
GPT-2
Small
1082∗∗∗
27.8∗∗∗
33∗∗∗
8.0∗∗
Table 3: χ2(1) from likelihood ratio tests, globally and locally at the main verb. M0: baseline (word log frequency + word length); M1: + surprisal; M2: + dependency length. ∗p<.05 , ∗∗p<.01 , ∗∗∗p<.001 . Global here tests dependency length as a sentence-level covariate across all positions; Local restricts to the integration site.
Condition
Observed EOI (ms)
Predicted EOI (ms)
See Table 2
KenLM (GPT2)
Mamba-130M
Pythia-1.4B
Pythia-70M
PP
22
1.04
0.10
− 0.07
− 1.84
SRC
49
0.02
− 1.18
− 1.25
− 2.13
ORC
93
1.99
0.10
0.03
− 1.42
2 × SRC
14
− 0.21
− 1.49
− 1.54
− 3.20
2 × ORC
63
1.69
0.31
0.41
− 0.89
Table 4: Predicted versus observed EOIs (relative to Baseline) at the main verb. Predicted values apply spillover-adjusted conversion factors, estimated on Natural Stories, to each model’s surprisal on StructuralCost critical items; negative values indicate a predicted speedup. Models without a significant conversion factor are omitted.
Dataset
Method
Uniq. Sent.
Subj.
Subj./Item
C1
C2
C3
C4
Naturalistic datasets
Natural Stories
Futrell et al.,2021
SPR
485
181
∼ 100
✗
✗
✓
✓
Dundee
Kennedy et al.,2003
ET
2,368
10
10
✗
✗
✓
✗
Provo
Luke and Christianson,2018
ET
∼ 138
84
∼ 8
✗
✗
−
✗
UCL Corpus
Frank et al.,2013
SPR + ET
361 + 205
117 + 43
✓
✗
✓
MECO
Kuperman et al.,2024
ET
104
50
✗
✗
−
Table 5: Comparison of reading time datasets against four criteria required to study LLM replication of human processing difficulty (non-exhaustive list). C1 : isolated sentences; C2 : systematic dependency length manipulation; C3 : sufficient items for model evaluation; C4 : sufficient participants per item. ET = eye-tracking; SPR = self-paced reading. ✓ = satisfied; ✗ = not satisfied; − = partially satisfied; * = information not found.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Dep. 0
Dep. 4
Dep. 9
Total
Critical
10
30
20
60
Control
10
10
10
30
Total
20
40
30
90
Appendix
Table 6: Items per participant by dependency length.
OSpan score
N
%
0
38
8
1
0
0
2
32
7
3
84
18
4
131
28
5
116
24
Appendix
Table 7: Distribution of operation-span (OSpan) scores across participants. Protocol: Conway et al. (2005) ; implementation adapted from von der Malsburg’s Python version.
Spearman r[95% CI ]
Model
Global
Dep=0
Dep=4
Dep=9
N -gram
KenLM_G
.51 [.49,.53]
.11 [-.14,.35]
.14 [-.01,.28]
.00 [-.18,.19]
KenLM_P
.51 [.49,.53]
.16 [-.09,.39]
.15 [.01,.29]
.01 [-.17,.19]
GPT-2
Small
.52 [.50,.54]
.24 [-.02,.47]
.31 [.17,.45]
.34 [.16,.51]
Appendix
Table 8: Spearman correlations between model surprisal and mean per-word RT, globally and locally at the main verb by dependency length. Bold marks the single best local score. CIs computed via bootstrap ( k=1000 ).