Reading behaviour varies not only with linguistic input, but also with reader proficiency. In this study, we investigate whether the layer-wise relationship between surprisal from large language models (LLMs) and human gaze behaviour differs across readers with different levels of proficiency and across gaze measures. Using eye-tracking data from the MECO L2 corpus, we compare readers with high and low vocabulary proficiency on first-pass gaze duration (FPGD) and total gaze duration (TGD). We quantify the distribution of the predictive power of surprisal across model layers using Predictive Depth. Across 12 tested LLMs, we find that readers with lower vocabulary proficiency tend to show deeper Predictive Depth for FPGD, while this difference is smaller for TGD. Also, TGD itself shows deeper Predictive Depth than FPGD in both proficiency groups. These patterns suggest that where predictive power is concentrated across LLM layers may be related to the timing and breadth of the reading processes captured by different gaze measures, and that this relationship can vary with reader proficiency. Our leave-one-out analysis further shows that the advantage of informative internal layers extends to unseen texts, although the practical improvements in prediction are limited. Overall, our results show that layer-wise LLM surprisal provides a useful perspective on variation in reading behaviour across both reader groups and gaze measures.
Figures & tables
Figure 1: Illustration of the distribution of predictive power along LLM layers. Readers with limited vocabulary knowledge (red) tend to show deeper Predictive Depth for FPGD compared to readers with higher vocabulary knowledge (blue). The difference in Predictive Depth is smaller for TGD, which itself shows deeper Predictive Depth than FPGD in both reader groups.
All
High-Lex
Low-Lex
LexTALE Score
-
≥ 83.75
< 66.25
Group Size
1001
276
228
Total Trials
8597
2519
1845
Table 1: Statistics of reader groups from MECO L2. Trials are the number of times each participant read each text.
Figure 2: Plots of layer-wise ΔLL for both FPGD (top) and TGD (bottom) from each LLM family. ΔLL values on the y-axis are normalized per token and multiplied by 1000 for better visualization.
FPGD
TGD
TGD-FPGD
LLM
Size
N
Low
High
L-H
Low
High
L-H
Low
High
Gemma 3
4B
34
0.668
0.553
+0.115
0.629
0.586
+0.043
−0.039
+0.033
12B
48
0.613
0.548
+0.065
0.678
0.599
+0.079
+0.065
+0.051
27B
62
0.523
0.549
−0.026
0.674
0.673
+0.001
+0.151
+0.124
Qwen 2.5
7B
28
0.558
0.549
+0.009
0.565
0.587
−0.022
+0.000
+0.039
14B
48
0.520
0.507
+0.014
0.558
0.559
−0.001
+0.038
+0.052
Table 2: Predictive Depth for FPGD and TGD in the Low-Lex and High-Lex groups across all tested LLMs. N represents the number of layers in each model. The last two columns show the difference in PD between TGD and FPGD for each group. Positive differences are shown in bold.
Low-Lex
MAE
(×100)
ρ
LLM
Final
Selected
Final
Selected
(Surface)
8.96
-
Gemma 3
4B
8.94
8.87
0.089
0.117
12B
8.95
8.87
−0.030
0.125
Table 3: Leave-one-out evaluation results for Low-Lex and High-Lex readers. Selected denotes the internal layer chosen based on ΔLL on the training texts. MAE values are multiplied by 100 for readability.
log-FPGD
rG,w
Surprisal
Word
Obs.
Surf.
Fin.
Sel.
Surf.
Fin.
Sel.
related
5.82
5.73
5.73
5.71
0.09
1.98
2.63
note,
5.71
5.59
5.57
5.53
0.12
0.14
0.00
sleep
5.48
5.65
5.63
5.66
−0.17
1.32
13.01
also
5.41
5.52
5.51
5.49
−0.11
1.53
0.02
improves
5.63
5.88
5.87
5.87
−0.25
1.57
8.87
Table 4: Illustrative example for Llama 3.1 70B. The selected internal layer is layer 46 out of 80. improves is highlighted as one of the words with the largest residual rG,w under the Surface LMM. Obs. denotes observed log-FPGD, while Surf., Fin., and Sel., denote the Surface, Final, and Selected LMM predictions, respectively. For each word, the prediction closest to the observed value is in bold.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Figure 3: ΔLL and Predictive Depth across all tested LLMs. As in Figure 2 , the color blue refers to High-Lex, and the red refers to Low-Lex.
Probing has shown that language model representations encode rich linguistic information, but it remains unclear whether they also capture cognitive signals about human processing. In this work, we probe language model representations for human reading times. Using regularized linear regression on two eye-tracking corpora spanning five languages (English, Greek, Hebrew, Russian, and Turkish), we compare the representations from every model layer against scalar predictors -- surprisal, information value, and logit-lens surprisal. We find that the representations from early layers outperform surprisal in predicting early-pass measures such as first fixation and gaze duration. The concentration of predictive power in the early layers suggests that human-like processing signatures are captured by low-level structural or lexical representations, pointing to a functional alignment between model depth and the temporal stages of human reading. In contrast, for late-pass measures such as total reading time, scalar surprisal remains superior, despite its being a much more compressed representation. We also observe performance gains when using both surprisal and early-layer representations. Overall, we find that the best-performing predictor varies strongly depending on the language and eye-tracking measure.
Eleftheria Tsipidi, Samuel Kiegeland, Francesco Ignazio Re +4
ETH Zürich · Toyota Technological Institute at Chicago · University College London
A recent study (Kuribayashi et al., 2025) has shown that human sentence processing behavior, typically measured on syntactically unchallenging constructions, can be effectively modeled using surprisal from early layers of large language models (LLMs). This raises the question of whether such advantages of internal layers extend to more syntactically challenging constructions, where surprisal has been reported to underestimate human cognitive effort. In this paper, we begin by exploring internal layers that better estimate human cognitive effort observed in syntactic ambiguity processing in English. Our experiments show that, in contrast to naturalistic reading, later layers better estimate such a cognitive effort, but still underestimate the human data. This dual alignment sheds light on different modes of sentence processing in humans and LMs: naturalistic reading employs a somewhat weak prediction akin to earlier layers of LMs, while syntactically challenging processing requires more fully-contextualized representations, better modeled by later layers of LMs. Motivated by these findings, we also explore several probability-update measures using shallow and deep layers of LMs, showing a complementary advantage to single-layer's surprisal in reading time modeling.
Tatsuki Kuribayashi, Alex Warstadt, Yohei Oseki +1
Predicting comprehension from eye movements could support adaptive reading interfaces. We present LEXIC, a compact recurrent model that predicts response correctness from fixation sequences, word frequency, and character length. It has 41.6K parameters and requires no language-model inference. Mean area under the receiver operating characteristic curve (AUROC) reaches 0.529 for Unseen Text and 0.554 for Unseen Reader on OneStop. Matched comparisons with AhnCNN show gains from both encoder redesign and lexical augmentation in these settings. Separate training and evaluation on SB-SAT also yield higher mean AUROC than AhnCNN. LEXIC occupies only 166 KB of model weights and requires 1.4 ms per trial on a single CPU core, supporting lightweight on-device inference.