Using LMs to Model the Effects of Context and Coreference during Sentence Comprehension
Authors: Kohei Kajikawa, Lin Ai, Tatsuki Kuribayashi, Ethan Gotlieb Wilcox
Organizations: Department of Linguistics, Georgetown University, USA · Computing and Mathematical Sciences Division, MBZUAI, UAE · Center for Language AI Research, Tohoku University, Japan
Language models (LMs) are often used as a tool to model human language processing. Recent studies suggest that severely restricting LMs' context window improves their fit to human psycholinguistic data by simulating human working memory constraints. However, it is possible that this strict memory-decay approach overlooks humans' reliance on long-range structural representations, such as discourse structre. In this work, we systematically vary the context window size of GPT-2 across four large-scale naturalistic English reading-time datasets and observe a U-shaped relationship: Although restricted contexts (< 20 tokens) successfully capture local memory limitations, expanded contexts (500--1,000 tokens) ultimately yield the highest overall psycholinguistic fit. To investigate the mechanism driving this benefit, we conduct a counterfactual inference-time experiment that disrupts cross-sentential entity chains by pronominalizing repeated discourse entities. Obscuring these structural linkages significantly degrades the predictive power of larger context windows by 20% to 40%. Our experiments demonstrate that tracking long-range coreference relations is one important factor for the alignment between LM surprisal and human reading behavior, and approximate the extent to which human comprehenders use global discourse relations during language processing.
Figures & tables
Dataset
Mean
SD
Range
Brown
707.00
199.59
313–959
Natural Stories
1237.30
68.28
1156–1359
OneStop
751.27
171.57
450–1271
Provo
59.53
8.20
43–80
Table 1: Descriptive statistics (mean, standard deviation, and range) of the number of tokens by the GPT-2 tokenizer per document across datasets. Provo is notably shorter than the others.
Figure 1: Predictive power ( ΔLL ) across different context window sizes. Each point represents the ΔLL with 95% confidence interval. For the Provo dataset, results are shown up to the document’s maximum length. The leftmost column displays zoomed-in plots for Brown, OneStop FF, and Provo FF, focusing on context windows of 1–20 tokens. Abbreviations: SPR = self-paced reading; FF = first fixation; GD = gaze duration; TF = total fixation time.
Figure 2: Localization analysis of context effects across POS tags using Bayesian mixed-effects models. Densities show the central 99% of the posterior distributions. Points represent posterior medians, while the thick and thin horizontal bars denote the 80% and 95% credible intervals, respectively. Red distributions indicate that the 95% credible interval does not contain zero.
Figure 3: Difference of predictive power ( ΔΔLL ) between Experiment 1 and our two target conditions— Singleton-Disrupted and Coref-Disrupted . Results are for Small .
Figure 4: Distribution of the percentage decrease in log-likelihood under Coref-Disrupted and Singleton-Disrupted conditions for Small at a maximum context across all datasets.
Table 2: Pretrained language models used in this paper.
Artifact
License
Brown ( Smith and Levy, 2013 )
CC BY 3.0
Natural Stories ( Futrell et al., 2021 )
CC BY-NC-SA 4.0
OneStop ( Berzak et al., 2025 )
CC BY 4.0
Provo ( Luke and Christianson, 2018 )
CC BY 4.0
Transformers ( Wolf et al., 2020 )
Apache 2.0
Stanza ( Qi et al., 2020 ; Liu et al., 2024 )
Apache 2.0
Appendix
Table 3: Artifacts and their licenses used in this paper.
Figure 5: Mean surprisal as a function of context window size up to maximum length across four reading-time datasets. Each point represents the mean value and error bars show 95% confidence interval.
Figure 6: Predictive power ( ΔLL ) across varying context window sizes for the GPT-2, GPT-Neo, GPT-J, and OPT model series on Natural Stories SPR.
Figure 7: Difference of predictive power between the original version reported in Figure 1 and Singleton-Disrupted and between the original version and Coref-Disrupted across multiple context sizes in Medium .
Figure 8: Difference of predictive power between the original version reported in Figure 1 and Singleton-Disrupted and between the original version and Coref-Disrupted across multiple context sizes in Large .
Figure 9: Difference of predictive power between the original version reported in Figure 1 and Singleton-Disrupted and between the original version and Coref-Disrupted across multiple context sizes in XL .
We reproduce and stress-test the work of Yu et al. (2023), who characterize how language models (LMs) arbitrate between memorized knowledge and contradictory in-context statements. We replicate their world-capitals experiments on 31 models spanning Pythia, GPT-2, Qwen3, and Ministral families, including base and post-trained variants, and extend evaluations to five additional knowledge relation types from the ParaConflict dataset. We empirically confirm most of their original findings: larger models and higher-frequency entities tend to favor memorized answers, with substantial family-level variance. However, several conclusions do not generalize cleanly: entity-frequency effects disappear on Qwen3-14B and 32B; post-training shifts the memory-context trade-off inconsistently across families; question phrasing alone can change a model's reliance on memorized knowledge by up to 80 percentage points; and semantically unrelated prose can mimic coherent supporting context. Our results clarify where Yu et al.'s claims hold and to what extent they generalize to other prompts.
Guilhem Fouilhé, Nicholas Asher, Philippe Muller
IRIT ANITI · IRIT CNRS · IRIT Université de Toulouse
A recent study (Kuribayashi et al., 2025) has shown that human sentence processing behavior, typically measured on syntactically unchallenging constructions, can be effectively modeled using surprisal from early layers of large language models (LLMs). This raises the question of whether such advantages of internal layers extend to more syntactically challenging constructions, where surprisal has been reported to underestimate human cognitive effort. In this paper, we begin by exploring internal layers that better estimate human cognitive effort observed in syntactic ambiguity processing in English. Our experiments show that, in contrast to naturalistic reading, later layers better estimate such a cognitive effort, but still underestimate the human data. This dual alignment sheds light on different modes of sentence processing in humans and LMs: naturalistic reading employs a somewhat weak prediction akin to earlier layers of LMs, while syntactically challenging processing requires more fully-contextualized representations, better modeled by later layers of LMs. Motivated by these findings, we also explore several probability-update measures using shallow and deep layers of LMs, showing a complementary advantage to single-layer's surprisal in reading time modeling.
Tatsuki Kuribayashi, Alex Warstadt, Yohei Oseki +1
Language models (LMs) are trained to excel at predicting the next word in the sequence given prior context, and humans also share this predictability in reading comprehension. Neuroscience research reveals that next-word predictability influences brain response, as recorded at millisecond resolution using electroencephalography (EEG). While our evidence indicates that advanced LMs achieve accuracies closely aligned with human performance at the next-word prediction task, this raises the question: Does higher prediction accuracy necessarily mean that these models adequately capture the cognitive signals associated with human reading comprehension? Here, we generate regressors for both humans and LMs based on two information measures, including top-1 prediction and surprisal, to predict event-related potential (ERP) elicited from EEG recordings which reflect different stages of cognitive processing during reading. We argue that modelling ERP patterns offers fine-grained analysis of the cognitive plausibility of various LMs during reading. Our results indicate that only surprisal potentially correlates with language-processing ERPs, especially for open-class words with high semantic content. Moreover, our findings challenge the assumption that scaling LMs with increased parameters and computational budgets will consistently lead to improved convergence with human-like linguistic processing.
Boi Mai Quach, Binh T. Nguyen, Cathal Gurrin +1
School of Computing, ML-Labs, Department of Computer Science, School of Computing, ADAPT Centre, Dublin City University, Ireland · VNUHCM – University of Science · School of Computing, ADAPT Centre, Dublin City University, Ireland