cs.CLOct 4, 2026

A Systematic Analysis of the Predictive Power of LM Surprisal in Reading Chinese

Authors: Hongao Zhu, Muxiaoqiao Xu, Yikang Liu, Siyuan Song, Yuxia Wang, Byung-Doh Oh, Hai Hu

Organizations: Department of Linguistics, UC San Diego · School of Foreign Languages, Shanghai Jiao Tong University · School of Computer Science, Shanghai Jiao Tong University · Department of Linguistics, UT Austin · Division of Linguistics and Multilingual Studies, Nanyang Technological University · Dept. of Language Science and Technology/Division of AI and the Humanities, Hong Kong Polytechnic University

Abstract

This study analyzes the predictive power of LM-derived, token-level surprisal on Mandarin Chinese reading times. We first propose the Shortest Matching Sequence (SMS), an alignment scheme that maps between the word segmentation assumed by eye-tracking corpora and the LMs' subword tokenization, as the two tokenizations often disagree in the context of Mandarin Chinese. Then, using a suite of Chinese-Pythia models (14M-1.4B) trained on scratch with 30B tokens, we examine how well surprisal predicts first fixation duration, gaze duration, and total reading time in three paragraph-level eye-tracking corpora of Mandarin Chinese (GECO-CN, HKP, and MECO). Contrary to previous null findings, our results show that surprisal is predictive of Chinese reading times. However, whether predictive power scales with model size and the amount of training is corpus-specific: bigger models predict better in GECO-CN, whereas inverse scaling emerges in HKP and, at the largest sizes, in MECO. Subsequently, we tested one possible explanation for the inverse scaling in HKP and found that checkpoints whose surprisal remains closer to nn-gram statistics are better predictors of reading. All in all, the predictive power of surprisal on Chinese reading time measurements is corpus-specific, which cautions against drawing scaling conclusions from a single corpus.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Reader Proficiency Shapes Layer-wise Surprisal Profiles

    Sep 29, 2026Akio Hayakawa, Horacio SaggionReadabilityProficiency

  2. Probing for Reading Times

    Apr 20, 2026Eleftheria Tsipidi, Samuel Kiegeland, Francesco Ignazio Re +4ReadabilityEye Tracking

  3. An Existence Proof for Neural Language Models That Can Explain Garden-Path Effects via Surprisal

    Apr 20, 2026Ryo Yoshida, Shinnosuke Isono, Taiga Someya +2SurprisalLinguistics