A Systematic Analysis of the Predictive Power of LM Surprisal in Reading Chinese
Authors: Hongao Zhu, Muxiaoqiao Xu, Yikang Liu, Siyuan Song, Yuxia Wang, Byung-Doh Oh, Hai Hu
Organizations: Department of Linguistics, UC San Diego · School of Foreign Languages, Shanghai Jiao Tong University · School of Computer Science, Shanghai Jiao Tong University · Department of Linguistics, UT Austin · Division of Linguistics and Multilingual Studies, Nanyang Technological University · Dept. of Language Science and Technology/Division of AI and the Humanities, Hong Kong Polytechnic University
This study analyzes the predictive power of LM-derived, token-level surprisal on Mandarin Chinese reading times. We first propose the Shortest Matching Sequence (SMS), an alignment scheme that maps between the word segmentation assumed by eye-tracking corpora and the LMs' subword tokenization, as the two tokenizations often disagree in the context of Mandarin Chinese. Then, using a suite of Chinese-Pythia models (14M-1.4B) trained on scratch with 30B tokens, we examine how well surprisal predicts first fixation duration, gaze duration, and total reading time in three paragraph-level eye-tracking corpora of Mandarin Chinese (GECO-CN, HKP, and MECO). Contrary to previous null findings, our results show that surprisal is predictive of Chinese reading times. However, whether predictive power scales with model size and the amount of training is corpus-specific: bigger models predict better in GECO-CN, whereas inverse scaling emerges in HKP and, at the largest sizes, in MECO. Subsequently, we tested one possible explanation for the inverse scaling in HKP and found that checkpoints whose surprisal remains closer to n-gram statistics are better predictors of reading. All in all, the predictive power of surprisal on Chinese reading time measurements is corpus-specific, which cautions against drawing scaling conclusions from a single corpus.
Figures & tables
TRT
FFD
GD
Corpus
Nsbj
Nwords
PPL
Nobs
ms
Nobs
ms
Nobs
ms
GECO-CN
30
46,680
60.9
348,105
281.8 (196.3)
348,527
206.5 (89.6)
253,298
221.4 (115.7)
HKP
98
5,071
46.6
195,840
289.7 (202.1)
196,048
209.8 (91.8)
158,274
222.4 (114.1)
MECO
39
1,812
31.5
27,192
358.5 (256.1)
26,293
227.0 (101.9)
26,283
248.5 (135.5)
Table 1: Descriptive statistics of the data entering the regression analyses, i.e., after preprocessing (Appendix A ). HKP refers to the passage-reading subset of HKC, and MECO refers to the simplified-Chinese portion. Nwords is the number of distinct word positions in the text (aligned between corpus and model tokenization; Section 3.4 ) with at least one retained TRT observation. Nobs is the number of subject–word observations retained for each reading-time measure; counts are identical across all model sizes and checkpoints. Reading times are reported as mean (standard deviation) in ms over the retained observations. Perplexity is 2 raised to the mean token-level surprisal (bits) under Qwen2.5-0.5B ( Qwen et al., 2025 ) , a fixed reference model independent of the Chinese-Pythia suite under evaluation; as a model-relative quantity, it indexes corpus difficulty only under this reference.
Chinese HKP eye-tracking corpus + Chinese-LLaMA tokenizer
Domestic ∣ company-DE ∣ leaders ∣ PL ∣ also ∣ with ∣ same ∣ enthusiasm ∣ focus-on ∣ these talents
Trans.
Leaders from domestic companies also focus on these talents with the same enthusiasm.
Table 2: Alignment between eye-tracking corpus tokenization and LM tokenization in Chinese. Token boundaries are shown as “ ∣ ”. Colors mark the three types of boundary mismatch; each is resolved by aligning to the smallest span on which both tokenizations agree, with the operations listed in the lower panel.
Corpus
#Subj
#Units
#Char
Unit Len
#Align
No Merge
Corpus Finer
LM Finer
Crossing
C/L Ratio
GECO-CN
30
540
101,233
187.5
47,996
24,204 (50.4%)
7,176 (15.0%)
13,399 (27.9%)
3,217 (6.7%)
0.54
HKP
98
41
9,066
221.1
5,117
4,077 (79.7%)
681 (13.3%)
316 (6.2%)
43 (0.8%)
2.16
MECO
39
12
3,836
319.7
1,836
1,155 (62.9%)
121 (6.6%)
497 (27.1%)
63 (3.4%)
0.24
Table 3: Corpus statistics and tokenization alignment between corpus and LM tokenizer. #Units is the number of text units submitted to alignment (chunks of the GECO-CN novel, passage chunks of HKP, and whole paragraphs of MECO), not the number of linguistic sentences. #Align is the total number of aligned word spans; the four condition counts sum to #Align. C/L Ratio is the ratio of corpus-finer to LM-finer spans. For MECO, #Subj counts the subjects present in the analyzed data.
Figure 1: Training trajectories of LLM–RT and LLM– n -gram correlations in HKP. Each panel corresponds to a Chinese Pythia model size.
Figure 2: Δllh for reading times across training. Positive values indicate that including surprisal improves the prediction of reading times relative to the baseline.
Corpus
Measure
14M
70M
160M
410M
1.4B
GECO-CN
TRT
1216
2219
2506
2610
2642
FFD
194
458
579
661
676
GD
990
1086
1206
1249
1322
HKP
TRT
365
488
508
396
396
FFD
135
128
110
83
78
GD
221
237
227
179
170
Table 4: Mean Δllh of the linear regression models over late-training checkpoints ( ≥ 300k steps; 16 checkpoints per cell), by corpus, measure, and model size. Bold marks the best-performing size per row. Higher values mean that adding surprisal yields a larger improvement in fit.
Figure 3: Correspondence between n -gram-likeness and reading-time fit in HKP. Each point is one model checkpoint. The x-axis shows the Pearson correlation between LM surprisal and n -gram surprisal; the y-axis shows the Pearson correlation between LM surprisal and a reading-time measure.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Corpus
Measure
Merged
Step 1
Step 2
Step 3
GECO-CN
TRT
719,518
353,763
353,328
348,105
FFD
719,518
353,763
353,761
348,527
GD
719,518
257,369
257,347
253,298
HKP
TRT
489,847
197,317
197,141
195,840
FFD
489,847
197,355
197,352
196,048
GD
489,847
159,373
159,360
158,274
Appendix
Table 5: Number of observations after each step of preprocessing, starting from the merged word-level dataset (after tokenization alignment; Section 3.4 ). Step 1 excludes skipped words (reading time of 0 or missing); step 2 excludes words with non-finite surprisal or reading times of 2000 ms or longer; step 3 excludes paragraph-initial words. GD counts drop sharply at step 1 because GD is undefined for multi-corpus-token spans. Counts are identical across all model sizes and checkpoints.
Reading behaviour varies not only with linguistic input, but also with reader proficiency. In this study, we investigate whether the layer-wise relationship between surprisal from large language models (LLMs) and human gaze behaviour differs across readers with different levels of proficiency and across gaze measures. Using eye-tracking data from the MECO L2 corpus, we compare readers with high and low vocabulary proficiency on first-pass gaze duration (FPGD) and total gaze duration (TGD). We quantify the distribution of the predictive power of surprisal across model layers using Predictive Depth. Across 12 tested LLMs, we find that readers with lower vocabulary proficiency tend to show deeper Predictive Depth for FPGD, while this difference is smaller for TGD. Also, TGD itself shows deeper Predictive Depth than FPGD in both proficiency groups. These patterns suggest that where predictive power is concentrated across LLM layers may be related to the timing and breadth of the reading processes captured by different gaze measures, and that this relationship can vary with reader proficiency. Our leave-one-out analysis further shows that the advantage of informative internal layers extends to unseen texts, although the practical improvements in prediction are limited. Overall, our results show that layer-wise LLM surprisal provides a useful perspective on variation in reading behaviour across both reader groups and gaze measures.
Probing has shown that language model representations encode rich linguistic information, but it remains unclear whether they also capture cognitive signals about human processing. In this work, we probe language model representations for human reading times. Using regularized linear regression on two eye-tracking corpora spanning five languages (English, Greek, Hebrew, Russian, and Turkish), we compare the representations from every model layer against scalar predictors -- surprisal, information value, and logit-lens surprisal. We find that the representations from early layers outperform surprisal in predicting early-pass measures such as first fixation and gaze duration. The concentration of predictive power in the early layers suggests that human-like processing signatures are captured by low-level structural or lexical representations, pointing to a functional alignment between model depth and the temporal stages of human reading. In contrast, for late-pass measures such as total reading time, scalar surprisal remains superior, despite its being a much more compressed representation. We also observe performance gains when using both surprisal and early-layer representations. Overall, we find that the best-performing predictor varies strongly depending on the language and eye-tracking measure.
Eleftheria Tsipidi, Samuel Kiegeland, Francesco Ignazio Re +4
ETH Zürich · Toyota Technological Institute at Chicago · University College London
Surprisal theory hypothesizes that the difficulty of human sentence processing increases linearly with surprisal, the negative log-probability of a word given its context. Computational psycholinguistics has tested this hypothesis using language models (LMs) as proxies for human prediction. While surprisal derived from recent neural LMs generally captures human processing difficulty on naturalistic corpora that predominantly consist of simple sentences, it severely underestimates processing difficulty on sentences that require syntactic disambiguation (garden-path effects). This leads to the claim that the processing difficulty of such sentences cannot be reduced to surprisal, although it remains possible that neural LMs simply differ from humans in next-word prediction. In this paper, we investigate whether it is truly impossible to construct a neural LM that can explain garden-path effects via surprisal. Specifically, instead of evaluating off-the-shelf neural LMs, we fine-tune these LMs on garden-path sentences so as to better align surprisal-based reading-time estimates with actual human reading times. Our results show that fine-tuned LMs do not overfit and successfully capture human reading slowdowns on held-out garden-path items; they even improve predictive power for human reading times on naturalistic corpora and preserve their general LM capabilities. These results provide an existence proof for a neural LM that can explain both garden-path effects and naturalistic reading times via surprisal, but also raise a theoretical question: what kind of evidence can truly falsify surprisal theory?