Using LMs to Model the Effects of Context and Coreference during Sentence Comprehension
Organizations: Department of Linguistics, Georgetown University, USA · Computing and Mathematical Sciences Division, MBZUAI, UAE · Center for Language AI Research, Tohoku University, Japan
Abstract
Language models (LMs) are often used as a tool to model human language processing. Recent studies suggest that severely restricting LMs' context window improves their fit to human psycholinguistic data by simulating human working memory constraints. However, it is possible that this strict memory-decay approach overlooks humans' reliance on long-range structural representations, such as discourse structre. In this work, we systematically vary the context window size of GPT-2 across four large-scale naturalistic English reading-time datasets and observe a U-shaped relationship: Although restricted contexts (< 20 tokens) successfully capture local memory limitations, expanded contexts (500--1,000 tokens) ultimately yield the highest overall psycholinguistic fit. To investigate the mechanism driving this benefit, we conduct a counterfactual inference-time experiment that disrupts cross-sentential entity chains by pronominalizing repeated discourse entities. Obscuring these structural linkages significantly degrades the predictive power of larger context windows by 20% to 40%. Our experiments demonstrate that tracking long-range coreference relations is one important factor for the alignment between LM surprisal and human reading behavior, and approximate the extent to which human comprehenders use global discourse relations during language processing.
Figures & tables
| Dataset | Mean | SD | Range |
|---|---|---|---|
| Brown | 707.00 | 199.59 | 313–959 |
| Natural Stories | 1237.30 | 68.28 | 1156–1359 |
| OneStop | 751.27 | 171.57 | 450–1271 |
| Provo | 59.53 | 8.20 | 43–80 |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | URL | #Params |
|---|---|---|
| GPT2-small ( Small ) | https://huggingface.co/openai-community/gpt2 | 124M |
| GPT2-medium ( Medium ) | https://huggingface.co/openai-community/gpt2-medium | 355M |
| GPT2-large ( Large ) | https://huggingface.co/openai-community/gpt2-large | 774M |
| GPT2-xl ( XL ) | https://huggingface.co/openai-community/gpt2-xl | 1.5B |
| GPT-Neo 125M | https://huggingface.co/EleutherAI/gpt-neo-125m | 125M |
| GPT-Neo 1.3B | https://huggingface.co/EleutherAI/gpt-neo-1.3B | 1.3B |
| Artifact | License |
|---|---|
| Brown ( Smith and Levy, 2013 ) | CC BY 3.0 |
| Natural Stories ( Futrell et al., 2021 ) | CC BY-NC-SA 4.0 |
| OneStop ( Berzak et al., 2025 ) | CC BY 4.0 |
| Provo ( Luke and Christianson, 2018 ) | CC BY 4.0 |
| Transformers ( Wolf et al., 2020 ) | Apache 2.0 |
| Stanza ( Qi et al., 2020 ; Liu et al., 2024 ) | Apache 2.0 |