Appearance-based gaze estimation (AGE) has achieved remarkable performance in constrained settings, yet we reveal a significant generalization gap where existing AGE models often fail in practical, unconstrained scenarios, particularly those involving facial wearables and poor lighting conditions. We attribute this failure to two core factors: limited image diversity and inconsistent label fidelity across different datasets, especially along the pitch axis. To address these, we propose a robust AGE framework that enhances generalization without requiring additional human-annotated data. First, we expand the image manifold via an ensemble of augmentation techniques, including synthesis of eyeglasses, masks, and varied lighting. Second, to mitigate the impact of anisotropic inter-dataset label deviation, we reformulate gaze regression as a multi-task learning problem, incorporating multi-view supervised contrastive (SupCon) learning, discretized label classification, and eye-region segmentation as auxiliary objectives. To rigorously validate our approach, we curate new benchmark datasets designed to evaluate gaze robustness under challenging conditions, a dimension largely overlooked by existing evaluation protocols. Our MobileNet-based lightweight model achieves generalization performance competitive with the state-of-the-art (SOTA) UniGaze-H, while utilizing less than 1% of its parameters, enabling high-fidelity, real-time gaze tracking on mobile devices.
Figures & tables
Figure 1 : Comparison of generalization performance between our MobileNet-based model and the SOTA UniGaze [ 31 ] . (a) On our benchmark datasets RealGaze (top) and ZeroGaze (bottom), UniGaze-H (red arrows) manifests significantly higher prediction variance under occlusion. (b) Our lightweight model maintains superior robustness and manifold consistency compared to significantly larger baselines, effectively marginalizing the impact of visual alterations. Detailed analysis is provided in Section 6.2 .
Dataset
Diversity
Image
Label
Subject#
Sample#
Gaze Range † (pitch × yaw)
Quality
Fidelity ∗
DE : EYEDIAP (FT sessions excluded) [ 11 ]
14 [ 36 ] ‡
15K
[5∘,21∘]×[−17∘,16∘]
Low
Medium
DC : GazeCapture [ 18 ]
1474
2M
[−24∘,9∘]×[−21∘,21∘]
Medium
Medium
DM : MPIIFaceGaze [ 47 ]
15
45K
[−24∘,0∘]×[−20∘,20∘]
Medium
Medium
D360 : Gaze360 [ 15 ]
238
100K [ 36 ]
[−62∘,15∘]×[−75∘,72∘]
Low
Low
DX : ETH-XGaze [ 45 ]
80
760K
[−65∘,56∘]×[−86∘,82∘]
High
High
Table 1 : Summary of representative AGE datasets. The datasets selected for our training pipeline are bolded . We define the symbols D∗ for simplicity.
Table 2 : Quantitative and qualitative analysis of inter-dataset label consistency. Perceived consistency rates on pitch and yaw axes reveal a significant discrepancy, confirming the presence of anisotropic inter-dataset label deviation.
Figure 2 : Qualitative comparison of samples from different datasets with near-identical annotations; visually divergent gaze directions, particularly in the pitch dimension, illustrate the inherent unreliability of cross-dataset vertical ground truths.
Figure 3 : Framework of the proposed multi-task learning architecture. Red blocks indicate the streamlined architecture during inference.
Figure 4 : Overview of the automated data augmentation pipeline. During training, we stochastically combine these methods for each sample to expand the training manifold.
Figure 5 : Pipeline for pose-consistent eyeglasses template generation. (a) Original face images, with 30 discrete head poses. (b) GlassesGAN outputs, featuring diverse frame styles. (c) Extracted glasses templates and examples of augmented training samples.
Figure 6 : Results of eye and iris segmentation produced by our MobileNet-based model. In (a) and (b), green regions denote the eye, while red denotes the iris.
Figure 7 : The nine session types in the RealGaze dataset. Sessions vary by illumination (light‑off O , side‑lit S ) and accessories (glasses G , mask M ), except Session a , which uses standard indoor lighting without accessories.
Model (Backbone, Training data)
Model
Overall
Ideal
Side-Lit
Glasses
Masks
size
dX
dY
∥d∥2
dX
dY
∥d∥2
dX
dY
∥d∥2
dX
dY
∥d∥2
dX
dY
∥d∥2
PureGaze (ResNet-50, DX )
31.0M
46.6
71.0
93.5
35.3
73.3
85.7
59.4
54.3
90.3
56.1
79.6
106.0
31.7
76.7
88.4
ETH-XGaze (ResNet-50, DX )
25.6M
42.9
59.7
82.1
24.8
62.7
72.3
33.0
56.9
71.8
73.3
56.6
104.8
31.3
62.0
75.1
UniGaze-B-Joint (ViT-B, 5 datasets)
86.6M
21.1
44.4
52.8
16.1
33.7
40.6
17.4
35.0
41.8
26.9
52.9
63.8
19.7
40.2
48.6
UniGaze-H-Joint (ViT-H, 5 datasets)
632M
18.3
44.1
51.5
13.9
38.1
43.1
14.6
35.5
41.3
25.1
47.9
59.0
16.0
37.9
45.3
Ours (MobileNet-v2, {DX,DN,DC} )
3.8M
22.3
34.9
46.3
16.0
28.8
36.6
18.6
28.5
37.0
25.4
33.4
46.6
20.9
36.0
45.3
Table 3 : Comparative Evaluation on RealGaze (errors in mm). The best results are in bold and the second best results are with underline .
Figure 8 : Distribution of AGE results on ZeroGaze across three views: C lean (blue), G lasses (red), and M asks (green). Our methods maintain concentrated and zero-centered, while baselines yield biased results with long tails, particularly along pitch.
Figure 9 : t-SNE visualization of high-level features on ZeroGaze. Our models successfully learn occlusion-invariant feature manifolds across the three views: Clean (blue), Glasses (red) and Masks (green).
Figure 10 : Performance benchmark across various backbone architectures on RealGaze.
Model
Overall
Ideal
Side-Lit
Glasses
Masks
dX
dY
∥d∥2
dX
dY
∥d∥2
dX
dY
∥d∥2
dX
dY
∥d∥2
dX
dY
∥d∥2
Ours, baseline (MobileNet)
22.3
34.9
46.3
16.0
28.8
36.6
18.6
28.5
37.0
25.4
33.4
46.6
20.9
36.0
45.3
(a) Training Data
DX only
32.9
45.4
61.9
23.5
41.8
52.0
31.0
34.8
52.1
36.5
43.6
63.0
35.3
52.0
69.2
{DX,DC}
26.1
45.1
57.2
18.8
47.4
54.5
23.3
31.5
43.2
31.6
45.3
61.1
25.7
56.2
66.9
{DX,DC,D360}
25.7
43.8
55.6
18.7
36.8
44.6
23.4
31.9
44.1
30.5
45.2
59.8
21.7
48.5
57.4
Table 4 : Ablation studies on RealGaze (errors in mm).
Figure 11 : Ablation of the dataset SupCon loss LDS . t-SNE visualizations of features from DX (blue), DC (red), and DN (green) show that LDS promotes a source-agnostic feature distribution, effectively marginalizing inter-dataset domain gaps.
Figure 12 : (a) Standard normalization pipeline for acquiring NCCS-aligned face images and corresponding 3D gaze vectors g . (b) Morphological variance in skull structures induces ambiguity in the individual head coordinate systems, particularly along the vertical axis (image from [ 21 ] ).
Figure 13 : Divergent HPE results lead to a systemic deviation between the resulting NCCS pitch gaze labels ϕ1 and ϕ2 .
Dataset
Subject #
Sample #
Cam. Param. Fidelity
Raw Label Fidelity
Image Quality & HPE Fidelity
NCCS
Overall Fidelity
DE : EYEDIAP (FT sessions excluded) [ 11 ]
14
15K
Medium
High
Low
Yes
Low
DC : GazeCapture [ 18 ]
1474
2M
Medium
Medium
Medium
Yes
Medium
DM : MPIIFaceGaze [ 47 ]
15
45K
Medium
High
Medium
Yes
Medium
D3 : Gaze360 [ 15 ]
238
172K
High
High
Low
No
Low
DX : ETH-XGaze [ 45 ]
110
760K
High
High
High
Yes
High
DN : GazeGene [ 3 ]
56
1M
N/A
N/A
N/A
No
Medium
Table 5 : Comparative audit of dataset scale and 3D gaze label fidelity. Overall ratings assess the systematic reliability of labels within the NCCS framework across three tiers: High : The label is consistently precise with minimal documented bias; Medium : The label is generally acceptable for training but subject to localized systemic errors (e.g., hardware-limited HPE or distribution skew); Low : The label is considered suspicious, suffering from severe systemic noise originating from multiple sources such as non-standardized coordinate systems or extreme environmental constraints.
Figure 15 : A rigid 2D transformation (incorporating rotation and translation) is applied to the glasses template by anchoring four facial landmarks. Landmark definitions are provided in Fig. 16 (b).
Figure 16 : (a) MediaPipe facial landmarks [ 23 ] ; the region delineated by yellow points is used to simulate mask occlusion. (b) The ocular landmarks. Meganta points guide glasses synthesis ( Fig. 15 ); yellow points define the eye region in the segmentation task, and cyan points specify the eye mask input M for the iris-region generation algorithm ( Algorithm 2 ).
Figure 17 : The eyelid exhibits significant appearance variation correlated with the pitch component of gaze, ϕ .
Figure 18 : Segmentation masks generated by MediaPipe Iris (red) versus our proposed intensity-based method (green), highlighting the latter’s superior alignment.
Figure 19 : RealGaze data collection. A subject participating in session a (no accessories, standard indoor lighting).
Figure 20 : Example ZeroGaze triplets( X,Xg,Xm ), generated via prompts S∗ (clean view), S∗+S1 (glasses view), and S∗+S2 (mask view). Note the high identity stability across diverse wearable prompts.
Figure 21 : Omitting the “ medical ” keyword in S2 induces undesired artifacts ( e.g . , apparel change or non-realistic occlusion) and catastrophic identity shifts ( e.g . , changes in age and gender). Images are shown at raw 512×512 resolution prior to cropping.
Figure 22 : Samples of ( X,Xg ), ( X,Xm ) pairs with top 5% similarity scores. Excessively high similarity scores may imply improper augmentation.
Figure 23 : Scatter plots of ground-truth labels (X-axis) versus model predictions (Y-axis) for a representative RealGaze subject. The persistent linear trend validates our training-free calibration approach. The red line denotes the identity mapping y=x .
Model
Overall
Ideal
Side-Lit
Glasses
Masks
Calibration Setting
dX
dY
∥d∥2
dX
dY
∥d∥2
dX
dY
∥d∥2
dX
dY
∥d∥2
dX
dY
∥d∥2
ETH-XGaze (ResNet-50, 25.6M)
Uncalibrated
42.9
59.7
82.1
24.8
62.7
72.3
33.0
56.9
71.8
73.3
56.6
104.8
31.3
62.0
75.1
1 Point
29.0
28.0
45.5
14.9
21.1
28.6
14.7
19.7
27.2
42.4
30.2
58.1
25.5
33.6
47.2
5 Points
23.3
22.6
36.3
14.3
18.1
25.7
15.1
17.7
25.8
33.7
25.2
46.9
24.7
24.6
38.9
UniGaze-B-Joint (ViT-B, 86.6M)
Table 6 : Results of personalized calibration on RealGaze using different numbers of calibration points (errors in mm).
Model
Model
Uncalibrated
3 Pts
5 Pts
10 Pts
20 Pts
Error (mm)
size
dX
dY
∥d∥2
dX
dY
∥d∥2
dX
dY
∥d∥2
dX
dY
∥d∥2
dX
dY
∥d∥2
Gaze322 (ResNet-18)
11.4M
-
-
101.9
-
-
79.3 ∗
-
-
62.3 ∗
-
-
56.7
-
-
45.4 ∗
ETH-XGaze (ResNet-50)
25.6M
47.4
82.9
104.3
32.1
63.0
76.6
26.6
45.6
57.0
23.3
47.8
56.5
21.3
46.3
54.8
UniGaze-B-Joint
86.6M
36.6
77.7
92.4
27.6
54.3
66.5
18.6
53.3
60.3
18.5
47.5
54.3
18.8
45.7
52.8
Ours (MobileNet)
3.8M
39.8
70.8
88.1
29.8
51.6
63.6
25.6
50.9
59.3
21.1
47.4
54.8
19.6
45.6
52.6
Ours (ViT-B)
86.6M
40.4
73.2
90.7
27.2
56.6
67.6
18.9
49.4
56.2
19.2
47.9
54.7
20.8
46.0
53.2
Table 7 : Results of personalized calibration on MPIIFaceGaze DM using different numbers of calibration points (errors in mm).
Model (Backbone)
Model
→DM
→DE
Size
d
dϕ
dψ
d
dϕ
dψ
(a) Trained on DX only
ETH-XGaze (ResNet-50) [ 45 ]
25.6M
6.94 (7.50 ∗ )
5.15
3.75
8.87 (11.0 ∗ )
5.23
6.40
PureGaze (ResNet-50) [ 6 ]
31M
6.85 (7.08 ∗ )
4.91
3.89
7.40 (7.48 ∗ )
4.67
4.75
CDG (ResNet-50) [ 36 ]
25.6M
6.73 ∗
-
-
7.95 ∗
-
-
FSCI (ResNet-18) [ 20 ]
16.4M
6.93 (5.79 ∗ )
4.64
4.32
7.94 (6.96 ∗ )
6.12
3.88
Table 8 : Cross-domain evaluation results on MPIIFaceGaze DM and EyeDiap DE (angular errors in degrees; dϕ and dψ are the pitch and yaw components respectively).
Model
Overall
Ideal
Side-Lit
Glasses
Masks
dX
dY
∥d∥2
dX
dY
∥d∥2
dX
dY
∥d∥2
dX
dY
∥d∥2
dX
dY
∥d∥2
(a) Variance Loss
Ours (MobileNet, baseline)
22.3
34.9
46.3
16.0
28.8
36.6
18.6
28.5
37.0
25.4
33.4
46.6
20.9
36.0
45.3
Replacing τ=0.5 with variance loss
21.9
38.0
47.9
17.0
33.9
41.2
22.2
28.1
39.6
24.9
42.6
54.3
18.8
40.7
48.6
(b) Pre-training
Ours (ViT-B, MAE-pretrained)
19.1
35.8
44.4
14.9
29.3
36.2
17.7
27.9
35.9
24.4
33.5
46.0
18.8
36.5
45.3
Table 9 : Ablation studies on RealGaze (errors in mm).
National Key Laboratory of Fundamental Science on Synthetic Vision, Sichuan University, Chengdu, China · College of Computer Science, Sichuan University, Chengdu, China
Department of Information and Telecommunication Engineering Soongsil University Seoul, Republic of Korea · Department of Electronic Engineering Soongsil University Seoul, Republic of Korea