Mobile eye-tracking is crucial for capturing human visual attention in real-world and XR settings, supporting research and human-computer interaction. Yet blinks, pupil-detection errors and lighting changes create missing values that hinder gaze analysis. We present HAGI++, a multi-modal diffusion-based imputation method that, for the first time, leverages integrated head-orientation sensors to exploit the natural correlation between head and eye movements. Using a transformer-based diffusion model, it learns cross-modal dependencies between eye and head data and can additionally incorporate wrist/hand motion when such wearable signals are available. Evaluations on the large-scale Nymeria, Ego-Exo4D and HOT3D datasets show that HAGI++ consistently outperforms traditional interpolation and deep-learning time-series imputation baselines. Statistical analysis confirms that its gaze-velocity distributions closely match real human behaviour, yielding realistic imputations. Even when 100% of gaze data are missing (pure gaze generation), HAGI++ exceeds methods that rely on the visual inputs and the methods rely on full-body motion capture by incorporating wrist motion from commercial wearables. Our approach enables more complete, accurate eye-gaze recordings in real-world contexts, enhancing gaze-based analysis and interaction across many applications. Our code is available at https://git.cai.simtech.uni-stuttgart.de/public-projects/HAGI
Figures & tables
Fig. 1: Missing data is inevitable in mobile eye tracking. HAGI++ is a novel multi-modal diffusion model for gaze data imputation that leverages the close coordination between eye and head movements. The input to our method is gaze data with missing values and time-aligned head movements captured by sensors readily available in mobile eye trackers (left). For XR headsets without built-in eye tracking, HAGI++ can generate gaze data directly from head movements (right). Furthermore, when wrist or hand motion data from commodity wearable devices are available, HAGI++ can exploit eye-hand-head coordination for both tasks. HAGI++ achieves lower mean angular error and produces more realistic gaze trajectories than previous methods.
Fig. 2: Overview of the training pipeline and model architecture of HAGI++.
Nymeria
Data loss ratio
10%
30%
50%
90%
Head direction
23.32
23.43
23.44
23.44
Pose2Gaze [ 24 ]
13.01
13.05
13.05
13.07
Linear
4.96
6.88
9.68
11.54
Nearest
5.29
6.52
8.34
12.61
iTransformer [ 38 ]
8.75
11.10
16.57
24.05
TABLE I: MAE of gaze imputation across different methods and data loss ratios on the Nymeria [ 20 ] test set. The best results are marked in bold, and the second-best are underlined.
Nymeria
Data loss ratio
10%
30%
50%
90%
Head direction
0.139
0.137
0.139
0.146
Linear
0.129
0.078
0.089
0.150
Nearest
0.081
0.073
0.103
0.135
CSDI [ 43 ]
0 .044
0 .042
0.037
0 .030
HAGI
0.042
0.040
0.035
0.017
TABLE II: JS of gaze imputation across different methods and data loss ratios on the Nymeria [ 20 ] test set. The best results are marked in bold, and the second-best are underlined.
Mean angular error (MAE)
Jensen–Shannon divergence (JS)
Dataset
Ego-Exo4D
HOT3D
Ego-Exo4D
HOT3D
Data loss ratio
10%
30%
50%
90%
10%
30%
50%
90%
10%
30%
50%
90%
10%
30%
50%
90%
Head direction
25.82
25.76
25.88
25.82
23.63
23.81
23.70
23.73
0.126
0.125
0.124
0.131
0.148
0.148
0.147
0.145
Linear
3.92
5.54
7.90
9.29
3.87
5.64
7.80
8.98
0.094
0.085
0.073
0.135
0.277
0.102
0.096
0.148
Nearest
4.18
5.18
6.68
10.13
4.14
5.23
6.61
9.86
0.062
0.066
0.081
0.121
0.080
0.081
0.100
0.135
CSDI [ 43 ]
3.78
4.78
6.07
8.91
3.67
4.69
5.85
8.28
0 .042
0.041
0.033
0.026
0 .052
0 .051
0 .042
0.029
TABLE III: Cross-dataset evaluation results: MAE and JS of gaze imputation across different methods and data loss ratios on the Ego-Exo4D [ 21 ] and HOT3D [ 22 ] datasets. The best results are marked in bold, and the second-best are underlined.
Fig. 3: Four examples of gaze imputation results at different missing ratios (10%, 30%, 50%, 90%) using different methods in the cross-dataset evaluation.
Head movements
Nymeria
Ego-Exo4D
HOT3D
Rotation
Translation
10%
30%
50%
90%
10%
30%
50%
90%
10%
30%
50%
90%
CSDI [ 43 ]
✗
✗
4.72
5.90
7.44
10.54
3.78
4.78
6.07
8.91
3.67
4.69
5.85
8.28
HAGI
✗
✓
4.66
5.75
7.29
10.47
3.72
4.63
5.91
8.79
3.60
4.57
5.74
8.19
✓
✗
4.41
5.36
6.66
9.56
3.55
4.35
5.43
8.23
3.42
4.33
5.38
7.79
✓
✓
3.67
4.55
5.77
8.53
3.06
3.80
4.86
7.37
3.01
3.91
4.99
7.34
HAGI++
✗
✓
4.65
5.79
7.34
10.65
3.74
4.63
5.89
8.78
3.61
4.60
5.75
8.38
TABLE IV: The MAE results of the ablation study on head rotation and translation across different data loss ratios in the Nymeria [ 20 ] , Ego-Exo4D [ 21 ] , and HOT3D [ 22 ] datasets. The best results are marked in bold, and the second-best are underlined.
Fig. 4: Head–gaze discrepancy vs. HAGI++ prediction error. Each panel shows one dataset. x-axis: MAE between the head proxy and GT gaze (per sequence). y-axis: MAE between HAGI++ and GT gaze on masked frames. The dashed diagonal is the baseline where the head direction is a proxy of gaze. Colored curves show HAGI++ at four missing ratios; shaded regions are the standard deviation across sequences within each head-proxy bin. HAGI++ errors remain low and largely flat as head-proxy error increases, demonstrating that our predictions do not simply replicate head orientation.
Memory (MB)
Inference Time (s)
CSDI [ 43 ]
3 848
4 .895
HAGI
4380
7.484
HAGI++
2608
3.455
TABLE V: The GPU memory consumption and inference time for three methods tested on one Nvidia Tesla V100 (32GB) GPU with a batch size of 512 during inference. The best results are marked in bold, and the second-best are underlined.
HAGI++ components
Nymeria
Ego-Exo4D
HOT3D
Head information
FiLM fusion
10%
30%
50%
90%
10%
30%
50%
90%
10%
30%
50%
90%
CSDI [ 43 ]
✗
✗
4.72
5.90
7.44
10.54
3.78
4.78
6.07
8.91
3.67
4.69
5.85
8.28
HAGI
✓
✗
3.67
4.55
5.77
8.53
3.06
3.80
4.86
7.37
3.01
3.91
4.99
7.34
HAGI++ (Ours)
✗
✗
4.61
5.74
7.27
10.38
3.70
4.61
5.92
8.78
3.58
4.56
5.72
8.23
✓
✗
3 .55
4 .42
5 .61
8 .24
2 .99
3 .75
4 .79
7 .18
2 .96
3 .86
4 .90
7 .26
✓
✓
3.54
4.40
5.58
8.18
2.98
3.73
4.76
7.08
2.94
3.84
4.86
7.19
TABLE VI: The ablation study results (MAE) on different HAGI++ components across different data loss ratios in the Nymeria [ 20 ] , Ego-Exo4D [ 21 ] , and HOT3D [ 22 ] datasets. The best results are marked in bold, and the second-best are underlined.
Nymeria
Data loss ratio
10%
30%
50%
90%
HAGI
4.01
4.94
6.34
9.17
HAGI++
3 .53
4 .41
5.67
8.17
HAGI++ + Left Wrist
3.52
4.34
5 .52
7 .98
HAGI++ + Right Wrist
3.61
4.48
5.68
8.09
HAGI++ + Two Wrists
3.54
4.34
5.49
7.88
TABLE VII: MAE across different methods and data loss ratios on the Nymeria (Wrist) test set [ 20 ] . The best results are marked in bold, and the second-best are underlined.
Nymeria
MAE
JS
Pose2Gaze [ 24 ]
13.09
0.238
HAGI++
11.65
0 .138
HAGI++ + Left Wrist
11.28
0.153
HAGI++ + Right Wrist
1 0.98
0.156
HAGI++ + Two Wrists
10.79
0.064
TABLE VIII: MAE and JS of gaze generation across different methods on the Nymeria (Wrist) test set [ 20 ] . The best results are marked in bold, and the second-best are underlined.
Fig. 5: Visualisation of gaze generation results from the one-point, two-point, and three-point variants of HAGI++, compared with the state-of-the-art Pose2Gaze [ 24 ] , which requires full-body pose input. Our method achieves higher accuracy using only signals from commodity wearable devices. Additional video results are provided in the supplementary materials.
Wrist movements
MAE
JS
Rotation
Translation
HAGI++
✗
✗
11.65
0.138
HAGI++ + Two Wrists
✗
✓
11.46
0.145
HAGI++ + Two Wrists
✓
✗
1 1.35
0 .129
HAGI++ + Two Wrists
✓
✓
10.79
0.064
TABLE IX: The ablation study results of wrist rotation and translation on the Nymeria [ 20 ] (Wrist) test set. The best results are marked in bold, and the second-best are underlined.
Chuhan Jiao is a PhD student at the University of Stuttgart, Germany. He received his BEng. in Computer Science and Technology from Donghua University, China, in 2020 and MSc. in Computer Science from Aalto University, Finland, in 2022. His research interest lies at the intersection of computer vision and human-computer interaction.
Table 15
Zhiming Hu is a tenure-track Assistant Professor at The Hong Kong University of Science and Technology (Guangzhou), with a joint appointment at The Hong Kong University of Science and Technology. He was a post-doctoral researcher in the University of Stuttgart, Germany from August 2022 to July 2025. He obtained his Ph.D. degree in Computer Software and Theory from Peking University, China in 2022 and received his Bachelor’s degree in Optical Engineering from Beijing Institute of Technology, China in 2017. His research interests include virtual reality, human-computer interaction, eye tracking, and human-centred artificial intelligence.
Table 16
Andreas Bulling is Full Professor of Computer Science at the University of Stuttgart, Germany, where he directs the research group ”Collaborative Artificial Intelligence”. He received his MSc. in Computer Science from the Karlsruhe Institute of Technology, Germany, in 2006 and his PhD in Information Technology and Electrical Engineering from ETH Zurich, Switzerland, in 2010. Before, Andreas Bulling was an Independent Research Group Leader and Senior Researcher at the Max Planck Institute for Informatics as well as Saarland University, Germany. His research interests include computer vision, machine learning, and human-computer interaction.
Table 17
Figure 18
John Doe Use \ begin{IEEEbiographynophoto} and the author name as the argument followed by the biography text.
Table 19
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Ego-Exo4D
HOT3D
10%
30%
50%
90%
10%
30%
50%
90%
iTransformer [ 38 ]
7.89
11.04
17.55
26.36
6.81
9.83
15.80
24.46
DLinear [ 40 ]
12.12
12.14
12.88
14.00
9.83
10.00
10.62
11.52
TimesNet [ 36 ]
21.66
20.27
22.37
25.23
19.83
18.79
20.51
23.23
BRITS [ 37 ]
10.51
12.47
14.96
19.17
9.47
11.36
13.50
17.20
Appendix
TABLE X: Mean angular error (MAE) of iTransformer [ 38 ] , Dlinear [ 40 ] , TimesNet [ 36 ] , BRITS [ 37 ] on the Ego-Exo4D [ 21 ] and HOT3D [ 22 ] datasets.
Fig. 6: HAGI++ Gaze prediction vs. head orientation over time on representative test sequences (90% missing). x-axis: frame index. y-axis: angular error vs. GT (°). Dashed black: head orientation vs. GT. Blue: HAGI++ prediction vs. GT (missing frames). Gray bar: missing frames.
Fig. 7: Visualisation of gaze prediction results from HAGI++ with and without visual input, compared against ground truth across four temporal frames in each sequence. The head-only variant uses eye–head coordination signals alone, while the visual-input variant additionally incorporates video features.
MAE
MSE
HAGI++
11.96
0.0071
HAGI++ + Visual Input
11.32
0.0068
Appendix
TABLE XI: Preliminary study on incorporating egocentric visual features into HAGI++ for gaze estimation on the Nymeria test set.
Department of Information and Telecommunication Engineering Soongsil University Seoul, Republic of Korea · Department of Electronic Engineering Soongsil University Seoul, Republic of Korea