A Systematic Analysis of the Predictive Power of LM Surprisal in Reading Chinese
Authors: Hongao Zhu, Muxiaoqiao Xu, Yikang Liu, Siyuan Song, Yuxia Wang, Byung-Doh Oh, Hai Hu
Organizations: Department of Linguistics, UC San Diego · School of Foreign Languages, Shanghai Jiao Tong University · School of Computer Science, Shanghai Jiao Tong University · Department of Linguistics, UT Austin · Division of Linguistics and Multilingual Studies, Nanyang Technological University · Dept. of Language Science and Technology/Division of AI and the Humanities, Hong Kong Polytechnic University
This study analyzes the predictive power of LM-derived, token-level surprisal on Mandarin Chinese reading times. We first propose the Shortest Matching Sequence (SMS), an alignment scheme that maps between the word segmentation assumed by eye-tracking corpora and the LMs' subword tokenization, as the two tokenizations often disagree in the context of Mandarin Chinese. Then, using a suite of Chinese-Pythia models (14M-1.4B) trained on scratch with 30B tokens, we examine how well surprisal predicts first fixation duration, gaze duration, and total reading time in three paragraph-level eye-tracking corpora of Mandarin Chinese (GECO-CN, HKP, and MECO). Contrary to previous null findings, our results show that surprisal is predictive of Chinese reading times. However, whether predictive power scales with model size and the amount of training is corpus-specific: bigger models predict better in GECO-CN, whereas inverse scaling emerges in HKP and, at the largest sizes, in MECO. Subsequently, we tested one possible explanation for the inverse scaling in HKP and found that checkpoints whose surprisal remains closer to n-gram statistics are better predictors of reading. All in all, the predictive power of surprisal on Chinese reading time measurements is corpus-specific, which cautions against drawing scaling conclusions from a single corpus.
Figures & tables
TRT
FFD
GD
Corpus
Nsbj
Nwords
PPL
Nobs
ms
Nobs
ms
Nobs
ms
GECO-CN
30
46,680
60.9
348,105
281.8 (196.3)
348,527
206.5 (89.6)
253,298
221.4 (115.7)
HKP
98
5,071
46.6
195,840
289.7 (202.1)
196,048
209.8 (91.8)
158,274
222.4 (114.1)
MECO
39
1,812
31.5
27,192
358.5 (256.1)
26,293
227.0 (101.9)
26,283
248.5 (135.5)
Table 1: Descriptive statistics of the data entering the regression analyses, i.e., after preprocessing (Appendix A ). HKP refers to the passage-reading subset of HKC, and MECO refers to the simplified-Chinese portion. Nwords is the number of distinct word positions in the text (aligned between corpus and model tokenization; Section 3.4 ) with at least one retained TRT observation. Nobs is the number of subject–word observations retained for each reading-time measure; counts are identical across all model sizes and checkpoints. Reading times are reported as mean (standard deviation) in ms over the retained observations. Perplexity is 2 raised to the mean token-level surprisal (bits) under Qwen2.5-0.5B ( Qwen et al., 2025 ) , a fixed reference model independent of the Chinese-Pythia suite under evaluation; as a model-relative quantity, it indexes corpus difficulty only under this reference.
Chinese HKP eye-tracking corpus + Chinese-LLaMA tokenizer
Domestic ∣ company-DE ∣ leaders ∣ PL ∣ also ∣ with ∣ same ∣ enthusiasm ∣ focus-on ∣ these talents
Trans.
Leaders from domestic companies also focus on these talents with the same enthusiasm.
Table 2: Alignment between eye-tracking corpus tokenization and LM tokenization in Chinese. Token boundaries are shown as “ ∣ ”. Colors mark the three types of boundary mismatch; each is resolved by aligning to the smallest span on which both tokenizations agree, with the operations listed in the lower panel.
Corpus
#Subj
#Units
#Char
Unit Len
#Align
No Merge
Corpus Finer
LM Finer
Crossing
C/L Ratio
GECO-CN
30
540
101,233
187.5
47,996
24,204 (50.4%)
7,176 (15.0%)
13,399 (27.9%)
3,217 (6.7%)
0.54
HKP
98
41
9,066
221.1
5,117
4,077 (79.7%)
681 (13.3%)
316 (6.2%)
43 (0.8%)
2.16
MECO
39
12
3,836
319.7
1,836
1,155 (62.9%)
121 (6.6%)
497 (27.1%)
63 (3.4%)
0.24
Table 3: Corpus statistics and tokenization alignment between corpus and LM tokenizer. #Units is the number of text units submitted to alignment (chunks of the GECO-CN novel, passage chunks of HKP, and whole paragraphs of MECO), not the number of linguistic sentences. #Align is the total number of aligned word spans; the four condition counts sum to #Align. C/L Ratio is the ratio of corpus-finer to LM-finer spans. For MECO, #Subj counts the subjects present in the analyzed data.
Figure 1: Training trajectories of LLM–RT and LLM– n -gram correlations in HKP. Each panel corresponds to a Chinese Pythia model size.
Figure 2: Δllh for reading times across training. Positive values indicate that including surprisal improves the prediction of reading times relative to the baseline.
Corpus
Measure
14M
70M
160M
410M
1.4B
GECO-CN
TRT
1216
2219
2506
2610
2642
FFD
194
458
579
661
676
GD
990
1086
1206
1249
1322
HKP
TRT
365
488
508
396
396
FFD
135
128
110
83
78
GD
221
237
227
179
170
Table 4: Mean Δllh of the linear regression models over late-training checkpoints ( ≥ 300k steps; 16 checkpoints per cell), by corpus, measure, and model size. Bold marks the best-performing size per row. Higher values mean that adding surprisal yields a larger improvement in fit.
Figure 3: Correspondence between n -gram-likeness and reading-time fit in HKP. Each point is one model checkpoint. The x-axis shows the Pearson correlation between LM surprisal and n -gram surprisal; the y-axis shows the Pearson correlation between LM surprisal and a reading-time measure.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Corpus
Measure
Merged
Step 1
Step 2
Step 3
GECO-CN
TRT
719,518
353,763
353,328
348,105
FFD
719,518
353,763
353,761
348,527
GD
719,518
257,369
257,347
253,298
HKP
TRT
489,847
197,317
197,141
195,840
FFD
489,847
197,355
197,352
196,048
GD
489,847
159,373
159,360
158,274
Appendix
Table 5: Number of observations after each step of preprocessing, starting from the merged word-level dataset (after tokenization alignment; Section 3.4 ). Step 1 excludes skipped words (reading time of 0 or missing); step 2 excludes words with non-finite surprisal or reading times of 2000 ms or longer; step 3 excludes paragraph-initial words. GD counts drop sharply at step 1 because GD is undefined for multi-corpus-token spans. Counts are identical across all model sizes and checkpoints.