Deep learning of longitudinal visual fields predicts glaucoma progression rate and identifies fast progressors
Authors: Taiabur Rahman, Siddiqur Rahman, Muhammad Moniruzzaman, Ummay Kawsar, Sayedatunnessa Ratna, Shadman Siddique, Rafsan Siddique, Tausif Ahmad, +2 more
Organizations: AI-MIQA, www.ai-miqa.eu. · Vision Eye Institute and Hospital, Dhaka, 1205, Bangladesh. · China West Normal University, Nanchong, Sichuan, China.
Glaucoma is the leading cause of irreversible blindness, and timely identification of fast progressors is essential to prevent disability. Current practice estimates progression by ordinary least-squares regression of mean deviation (MD) on time, requiring 6--10 visual field (VF) tests over several years to obtain a reliable slope. We present GLAM (Glaucoma Longitudinal Analysis Model), a deep learning framework that ingests longitudinal Humphrey 24-2 total deviation sequences with five clinical features and predicts MD and visual field index progression rates using attention-based fusion and aleatoric uncertainty. On the open-access University of Washington Humphrey Visual Field dataset (4,276 patient-eyes), GLAM achieved an MD-rate mean absolute error of 0.139 dB yr−1 (R2=0.927; 73.5% reduction over a ridge baseline) and an AUC of 0.990 for fast-progressor detection. VF-only deep learning can match multimodal pipelines for progression prognostication using routinely collected perimetry alone.
Figures & tables
Characteristic
All
Train
Val.
Test
( n=4,276 )
( n=2,992 )
( n=640 )
( n=644 )
Visits per eye, mean ± s.d.
5.2±2.8
5.2±2.8
5.2±2.8
5.2±2.8
Follow-up, years (median, IQR)
4.3 (2.5–7.7)
4.3 (2.5–7.7)
4.3 (2.5–7.7)
4.2 (2.5–7.7)
MD rate, dB yr -1 (mean ± s.d.)
−0.126±0.956
−0.127±0.957
−0.122±0.950
−0.126±0.964
Fast progressors (MD <−1 dB yr -1 )
394 (9.2%)
276 (9.2%)
62 (9.7%)
56 (8.7%)
Baseline MD proxy, dB (mean ± s.d.)
−6.1±6.1
−6.1±6.1
−6.0±6.1
−6.2±6.2
Table 1: Cohort characteristics of the analysis population (UWHVF, n=4,276 patient-eyes). 0 0 footnotetext: IQR, interquartile range; MD, mean deviation; s.d., standard deviation; UWHVF, University of Washington Humphrey Visual Field dataset.
Figure 1: GLAM model architecture. Longitudinal VF total deviation sequences ( T visits × 54 points) are encoded by a VisualFieldEncoder and processed by a two-layer bidirectional LSTM to produce a 512-dimensional functional representation. A five-feature clinical branch produces a 128-dimensional embedding. An attention-fusion module combines all modality representations with learned soft weights. A dual regression head predicts MD and VFI progression rates with aleatoric uncertainty. The structural branch (dashed) is inactive in this study because UWHVF contains no fundus images and outputs zeros. BiLSTM, bidirectional long short-term memory; MD, mean deviation; MLP, multilayer perceptron; VFI, visual field index; UWHVF, University of Washington Humphrey Visual Field dataset.
Model
MAE
RMSE
R2
AUC (95% CI)
Sens.
Spec.
F1
Naive (global mean)
0.524
0.797
—
0.500 (0.500–0.500)
—
—
—
Ridge (baseline MD)
0.527
0.802
—
0.367 (0.293–0.444)
—
—
—
GLAM (this work)
0.139
0.238
0.927
0.990 (0.982–0.996)
98.2
90.6
0.663
Table 2: Test-set performance of GLAM and clinical baselines ( n=644 patient-eyes). 0 0 footnotetext: MAE and RMSE are MD rate in dB yr -1 . Sensitivity, specificity (%), and F1 for GLAM are reported at the Youden-optimal threshold of −0.58 dB yr -1 . Ridge regression AUC was computed using the negated predicted MD as the discrimination score. AUC, area under the receiver operating characteristic curve; CI, confidence interval (stratified bootstrap, 1,000 resamples); MAE, mean absolute error; MD, mean deviation; RMSE, root mean squared error.
Figure 2: Predicted versus true MD progression rate on the held-out test set ( n=644 patient-eyes). Each point represents one patient-eye. The red dashed line indicates the line of identity. R2=0.927 ; mean absolute error =0.139 dB yr -1 ; signed bias =0.001 dB yr -1 . Point colour encodes prediction error magnitude (darker = larger absolute error). MD, mean deviation.
Figure 3: Receiver operating characteristic curve for fast-progressor detection (MD <−1 dB yr -1 ) on the held-out test set. AUC =0.990 (95% confidence interval by stratified bootstrap: 0.982–0.996). The Youden-optimal operating point (predicted MD threshold =−0.58 dB yr -1 ; sensitivity 98.2%, specificity 90.6%) is marked. AUC, area under the receiver operating characteristic curve; MD, mean deviation.
Figure 4: Aleatoric uncertainty calibration plot. Predicted standard deviation σ=exp(0.5×logvar) plotted against absolute prediction error ∣MDpred−MDtrue∣ for each patient-eye in the held-out test set. The red dashed line denotes ideal calibration ( σ=∣error∣ ). Empirical coverage of nominal 90% prediction intervals is 99.8%, indicating conservative over-coverage; Spearman ρ between predicted standard deviation and absolute error is 0.41 ( P<0.001 ). MD, mean deviation.
Modality
All ( n=644 )
Fast ( n=56 )
Slow ( n=588 )
Structural (placeholder)
0.321
0.329
0.317
Functional (BiLSTM)
0.410
0.396
0.413
Clinical (MLP)
0.271
0.275
0.270
Table 3: Mean modality attention weights by progression subgroup (held-out test set). 0 0 footnotetext: The structural branch carries only zeros and therefore cannot influence predictions; non-zero weights reflect the architectural softmax constraint αs+αf+αc=1 . Fast/slow columns are progressor subgroups. BiLSTM, bidirectional long short-term memory network; MLP, multilayer perceptron.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Test-set MAE by MD-rate quintile. Bars show MAE within each quintile of true MD rate (Q1: ≥0 dB yr -1 , n=129 ; Q2: −0.09 to 0 dB yr -1 , n=129 ; Q3: −0.35 to −0.09 dB yr -1 , n=129 ; Q4: −0.82 to −0.35 dB yr -1 , n=129 ; Q5: <−0.82 dB yr -1 , n=128 ). Error increases monotonically with progression severity from 0.092 dB yr -1 (Q1) to 0.205 dB yr -1 (Q5). MAE, mean absolute error; MD, mean deviation.
Glaucoma is a leading cause of irreversible blindness worldwide, and early detection from fundus images is critical for effective disease management. While deep learning has achieved promising performance in fundus image analysis, most existing methods rely on single time-point images and fail to capture longitudinal structural and vascular changes associated with disease progression. Sequential fundus images acquired during clinical follow-up provide valuable temporal information; however, current sequential models often struggle to detect subtle early progression signals and commonly depend on fixed-length inputs or diagnostic cues from already glaucomatous images, limiting their clinical utility for early prediction. To address these limitations, we propose DiffSight-Former, a framework for glaucoma progression prediction from sequential fundus images. It incorporates a time-variant feature extraction module based on a fundus-specific foundation model to obtain robust anatomical representations. A multi-structure difference modeling module is introduced to quantify progression-related changes in the optic disc/cup region and retinal vasculature. These representations are integrated with temporal interval embeddings and processed by a time-aware Transformer to model disease progression and estimate the probability of future glaucoma onset. Experiments were conducted on two longitudinal datasets, SIGF (405 sequences) and GRAPE (263 sequences). On SIGF, DiffSight-Former achieved an AUC of 91.54% and a sensitivity of 92.16% for progression prediction. On GRAPE, it achieved an average accuracy of 87.48% across three clinical visual-field progression criteria. Compared with existing approaches, DiffSight-Former demonstrates strong performance and robustness across different temporal settings, highlighting its potential for longitudinal glaucoma monitoring and early risk prediction.
Yi Huang, Lei Bi, Jinman Kim
School of Computer Science, The University of Sydney · Institute of Translational Medicine, Shanghai Jiao Tong University
Forecasting visual fields (VFs) is critical for personalized monitoring and treatment planning in glaucoma. This is inherently uncertain due to heterogeneous disease progression and measurement variability, yet most existing methods produce single deterministic predictions that fail to represent this uncertainty. We formulate VF forecasting as a probabilistic prediction problem and the use of conditioned denoising diffusion models to generate distributions of plausible future VFs from longitudinal observations with irregular follow-up intervals. Experiments on two independent VF cohorts show that diffusion-based predictions produce well-calibrated distributions for clinically relevant VF measures. When reduced to a standard point-estimate, the proposed approach achieves state-of-the-art accuracy compared to clinical baselines and prior learning-based methods. Our results highlight the advantages of distributional modeling for VF forecasting and support a shift from point-estimate prediction toward uncertainty-aware, clinically interpretable risk assessment in glaucoma.
Marta Colmenar Herrera, Pablo Márquez Neila, Şerife Seda Kucur Ergünay +2
ARTORG Center for Biomedical Engineering Research, UniBe, Switzerland · PeriVision SA, Epalinges, Switzerland · Department of Ophthalmology, Inselspital, Bern, Switzerland
Background: Artificial intelligence (AI)-based glaucoma detection from colour fundus photographs (CFP) offers scalable screening, but performance may decline on external datasets because of differences in ground-truth definitions, populations, and coexisting conditions such as high myopia (HM). We developed and validated a Vision Transformer-based deep learning (DL) model for glaucoma detection across multi-ethnic cohorts with and without HM. Methods: A ViT-B/16 model with predictive uncertainty estimation was developed using 56,483 CFPs (57.1% with myopia; 14.4% with HM). Glaucoma labels were standardised using clinical, imaging, and perimetry data. The model was validated on 16 independent datasets across three continents, including four datasets with explicit HM labels. Findings: Internal AUROC was 98.7% (95% CI 98.2-99.1%), with sensitivity 94.5% and specificity 97.3%. Across 16 external datasets from eight countries, AUROCs ranged from 86.4% to 99.6%. In HM eyes, internal AUROC was 97.8% (95% CI 96.1-99.2%), with sensitivity 94.8% and specificity 93.7%. External HM AUROCs were 86.5% in the Beijing Eye Study and 93.3%, 91.8%, and 85.5% in hospital-based datasets from Taiwan, Thailand, and South Korea. In an exploratory HM clinical evaluation, the model had higher CFP-only diagnostic accuracy than ophthalmologists and trained graders (92.0% vs 70.0%; p=0.008) and performed comparably to glaucoma specialists using full clinical information. Interpretation: The model showed robust glaucoma detection across myopic and non-myopic multi-ethnic populations and may support AI-assisted screening in settings with high HM prevalence.
Raghavan Lavanya, Yangqin Feng, Ten Cheer Quek +36
Singapore Eye Research Institute, Singapore National Eye Centre, Singapore · Ophthalmology & Visual Sciences Academic Clinical Program (Eye ACP), Duke-NUS Medical School, Singapore · Institute of Advanced Intelligence and Computing, Agency for Science, Technology and Research, Singapore +14