Appearance-based gaze estimation (AGE) has achieved remarkable performance in constrained settings, yet we reveal a significant generalization gap where existing AGE models often fail in practical, unconstrained scenarios, particularly those involving facial wearables and poor lighting conditions. We attribute this failure to two core factors: limited image diversity and inconsistent label fidelity across different datasets, especially along the pitch axis. To address these, we propose a robust AGE framework that enhances generalization without requiring additional human-annotated data. First, we expand the image manifold via an ensemble of augmentation techniques, including synthesis of eyeglasses, masks, and varied lighting. Second, to mitigate the impact of anisotropic inter-dataset label deviation, we reformulate gaze regression as a multi-task learning problem, incorporating multi-view supervised contrastive (SupCon) learning, discretized label classification, and eye-region segmentation as auxiliary objectives. To rigorously validate our approach, we curate new benchmark datasets designed to evaluate gaze robustness under challenging conditions, a dimension largely overlooked by existing evaluation protocols. Our MobileNet-based lightweight model achieves generalization performance competitive with the state-of-the-art (SOTA) UniGaze-H, while utilizing less than 1% of its parameters, enabling high-fidelity, real-time gaze tracking on mobile devices.
Figures & tables
Figure 1 : Comparison of generalization performance between our MobileNet-based model and the SOTA UniGaze [ 31 ] . (a) On our benchmark datasets RealGaze (top) and ZeroGaze (bottom), UniGaze-H (red arrows) manifests significantly higher prediction variance under occlusion. (b) Our lightweight model maintains superior robustness and manifold consistency compared to significantly larger baselines, effectively marginalizing the impact of visual alterations. Detailed analysis is provided in Section 6.2 .
Dataset
Diversity
Image
Label
Subject#
Sample#
Gaze Range † (pitch × yaw)
Quality
Fidelity ∗
DE : EYEDIAP (FT sessions excluded) [ 11 ]
14 [ 36 ] ‡
15K
[5∘,21∘]×[−17∘,16∘]
Low
Medium
DC : GazeCapture [ 18 ]
1474
2M
[−24∘,9∘]×[−21∘,21∘]
Medium
Medium
DM : MPIIFaceGaze [ 47 ]
15
45K
[−24∘,0∘]×[−20∘,20∘]
Medium
Medium
D360 : Gaze360 [ 15 ]
238
100K [ 36 ]
[−62∘,15∘]×[−75∘,72∘]
Low
Low
DX : ETH-XGaze [ 45 ]
80
760K
[−65∘,56∘]×[−86∘,82∘]
High
High
Table 1 : Summary of representative AGE datasets. The datasets selected for our training pipeline are bolded . We define the symbols D∗ for simplicity.
Table 2 : Quantitative and qualitative analysis of inter-dataset label consistency. Perceived consistency rates on pitch and yaw axes reveal a significant discrepancy, confirming the presence of anisotropic inter-dataset label deviation.
Figure 2 : Qualitative comparison of samples from different datasets with near-identical annotations; visually divergent gaze directions, particularly in the pitch dimension, illustrate the inherent unreliability of cross-dataset vertical ground truths.
Figure 3 : Framework of the proposed multi-task learning architecture. Red blocks indicate the streamlined architecture during inference.
Figure 4 : Overview of the automated data augmentation pipeline. During training, we stochastically combine these methods for each sample to expand the training manifold.
Figure 5 : Pipeline for pose-consistent eyeglasses template generation. (a) Original face images, with 30 discrete head poses. (b) GlassesGAN outputs, featuring diverse frame styles. (c) Extracted glasses templates and examples of augmented training samples.
Figure 6 : Results of eye and iris segmentation produced by our MobileNet-based model. In (a) and (b), green regions denote the eye, while red denotes the iris.
Figure 7 : The nine session types in the RealGaze dataset. Sessions vary by illumination (light‑off O , side‑lit S ) and accessories (glasses G , mask M ), except Session a , which uses standard indoor lighting without accessories.
Model (Backbone, Training data)
Model
Overall
Ideal
Side-Lit
Glasses
Masks
size
dX
dY
∥d∥2
dX
dY
∥d∥2
dX
dY
∥d∥2
dX
dY
∥d∥2
dX
dY
∥d∥2
PureGaze (ResNet-50, DX )
31.0M
46.6
71.0
93.5
35.3
73.3
85.7
59.4
54.3
90.3
56.1
79.6
106.0
31.7
76.7
88.4
ETH-XGaze (ResNet-50, DX )
25.6M
42.9
59.7
82.1
24.8
62.7
72.3
33.0
56.9
71.8
73.3
56.6
104.8
31.3
62.0
75.1
UniGaze-B-Joint (ViT-B, 5 datasets)
86.6M
21.1
44.4
52.8
16.1
33.7
40.6
17.4
35.0
41.8
26.9
52.9
63.8
19.7
40.2
48.6
UniGaze-H-Joint (ViT-H, 5 datasets)
632M
18.3
44.1
51.5
13.9
38.1
43.1
14.6
35.5
41.3
25.1
47.9
59.0
16.0
37.9
45.3
Ours (MobileNet-v2, {DX,DN,DC} )
3.8M
22.3
34.9
46.3
16.0
28.8
36.6
18.6
28.5
37.0
25.4
33.4
46.6
20.9
36.0
45.3
Table 3 : Comparative Evaluation on RealGaze (errors in mm). The best results are in bold and the second best results are with underline .
Figure 8 : Distribution of AGE results on ZeroGaze across three views: C lean (blue), G lasses (red), and M asks (green). Our methods maintain concentrated and zero-centered, while baselines yield biased results with long tails, particularly along pitch.
Figure 9 : t-SNE visualization of high-level features on ZeroGaze. Our models successfully learn occlusion-invariant feature manifolds across the three views: Clean (blue), Glasses (red) and Masks (green).
Figure 10 : Performance benchmark across various backbone architectures on RealGaze.
Model
Overall
Ideal
Side-Lit
Glasses
Masks
dX
dY
∥d∥2
dX
dY
∥d∥2
dX
dY
∥d∥2
dX
dY
∥d∥2
dX
dY
∥d∥2
Ours, baseline (MobileNet)
22.3
34.9
46.3
16.0
28.8
36.6
18.6
28.5
37.0
25.4
33.4
46.6
20.9
36.0
45.3
(a) Training Data
DX only
32.9
45.4
61.9
23.5
41.8
52.0
31.0
34.8
52.1
36.5
43.6
63.0
35.3
52.0
69.2
{DX,DC}
26.1
45.1
57.2
18.8
47.4
54.5
23.3
31.5
43.2
31.6
45.3
61.1
25.7
56.2
66.9
{DX,DC,D360}
25.7
43.8
55.6
18.7
36.8
44.6
23.4
31.9
44.1
30.5
45.2
59.8
21.7
48.5
57.4
Table 4 : Ablation studies on RealGaze (errors in mm).
Figure 11 : Ablation of the dataset SupCon loss LDS . t-SNE visualizations of features from DX (blue), DC (red), and DN (green) show that LDS promotes a source-agnostic feature distribution, effectively marginalizing inter-dataset domain gaps.
Figure 12 : (a) Standard normalization pipeline for acquiring NCCS-aligned face images and corresponding 3D gaze vectors g . (b) Morphological variance in skull structures induces ambiguity in the individual head coordinate systems, particularly along the vertical axis (image from [ 21 ] ).
Figure 13 : Divergent HPE results lead to a systemic deviation between the resulting NCCS pitch gaze labels ϕ1 and ϕ2 .
Dataset
Subject #
Sample #
Cam. Param. Fidelity
Raw Label Fidelity
Image Quality & HPE Fidelity
NCCS
Overall Fidelity
DE : EYEDIAP (FT sessions excluded) [ 11 ]
14
15K
Medium
High
Low
Yes
Low
DC : GazeCapture [ 18 ]
1474
2M
Medium
Medium
Medium
Yes
Medium
DM : MPIIFaceGaze [ 47 ]
15
45K
Medium
High
Medium
Yes
Medium
D3 : Gaze360 [ 15 ]
238
172K
High
High
Low
No
Low
DX : ETH-XGaze [ 45 ]
110
760K
High
High
High
Yes
High
DN : GazeGene [ 3 ]
56
1M
N/A
N/A
N/A
No
Medium
Table 5 : Comparative audit of dataset scale and 3D gaze label fidelity. Overall ratings assess the systematic reliability of labels within the NCCS framework across three tiers: High : The label is consistently precise with minimal documented bias; Medium : The label is generally acceptable for training but subject to localized systemic errors (e.g., hardware-limited HPE or distribution skew); Low : The label is considered suspicious, suffering from severe systemic noise originating from multiple sources such as non-standardized coordinate systems or extreme environmental constraints.
Figure 15 : A rigid 2D transformation (incorporating rotation and translation) is applied to the glasses template by anchoring four facial landmarks. Landmark definitions are provided in Fig. 16 (b).
Figure 16 : (a) MediaPipe facial landmarks [ 23 ] ; the region delineated by yellow points is used to simulate mask occlusion. (b) The ocular landmarks. Meganta points guide glasses synthesis ( Fig. 15 ); yellow points define the eye region in the segmentation task, and cyan points specify the eye mask input M for the iris-region generation algorithm ( Algorithm 2 ).
Figure 17 : The eyelid exhibits significant appearance variation correlated with the pitch component of gaze, ϕ .
Figure 18 : Segmentation masks generated by MediaPipe Iris (red) versus our proposed intensity-based method (green), highlighting the latter’s superior alignment.
Figure 19 : RealGaze data collection. A subject participating in session a (no accessories, standard indoor lighting).
Figure 20 : Example ZeroGaze triplets( X,Xg,Xm ), generated via prompts S∗ (clean view), S∗+S1 (glasses view), and S∗+S2 (mask view). Note the high identity stability across diverse wearable prompts.
Figure 21 : Omitting the “ medical ” keyword in S2 induces undesired artifacts ( e.g . , apparel change or non-realistic occlusion) and catastrophic identity shifts ( e.g . , changes in age and gender). Images are shown at raw 512×512 resolution prior to cropping.
Figure 22 : Samples of ( X,Xg ), ( X,Xm ) pairs with top 5% similarity scores. Excessively high similarity scores may imply improper augmentation.
Figure 23 : Scatter plots of ground-truth labels (X-axis) versus model predictions (Y-axis) for a representative RealGaze subject. The persistent linear trend validates our training-free calibration approach. The red line denotes the identity mapping y=x .
Model
Overall
Ideal
Side-Lit
Glasses
Masks
Calibration Setting
dX
dY
∥d∥2
dX
dY
∥d∥2
dX
dY
∥d∥2
dX
dY
∥d∥2
dX
dY
∥d∥2
ETH-XGaze (ResNet-50, 25.6M)
Uncalibrated
42.9
59.7
82.1
24.8
62.7
72.3
33.0
56.9
71.8
73.3
56.6
104.8
31.3
62.0
75.1
1 Point
29.0
28.0
45.5
14.9
21.1
28.6
14.7
19.7
27.2
42.4
30.2
58.1
25.5
33.6
47.2
5 Points
23.3
22.6
36.3
14.3
18.1
25.7
15.1
17.7
25.8
33.7
25.2
46.9
24.7
24.6
38.9
UniGaze-B-Joint (ViT-B, 86.6M)
Table 6 : Results of personalized calibration on RealGaze using different numbers of calibration points (errors in mm).
Model
Model
Uncalibrated
3 Pts
5 Pts
10 Pts
20 Pts
Error (mm)
size
dX
dY
∥d∥2
dX
dY
∥d∥2
dX
dY
∥d∥2
dX
dY
∥d∥2
dX
dY
∥d∥2
Gaze322 (ResNet-18)
11.4M
-
-
101.9
-
-
79.3 ∗
-
-
62.3 ∗
-
-
56.7
-
-
45.4 ∗
ETH-XGaze (ResNet-50)
25.6M
47.4
82.9
104.3
32.1
63.0
76.6
26.6
45.6
57.0
23.3
47.8
56.5
21.3
46.3
54.8
UniGaze-B-Joint
86.6M
36.6
77.7
92.4
27.6
54.3
66.5
18.6
53.3
60.3
18.5
47.5
54.3
18.8
45.7
52.8
Ours (MobileNet)
3.8M
39.8
70.8
88.1
29.8
51.6
63.6
25.6
50.9
59.3
21.1
47.4
54.8
19.6
45.6
52.6
Ours (ViT-B)
86.6M
40.4
73.2
90.7
27.2
56.6
67.6
18.9
49.4
56.2
19.2
47.9
54.7
20.8
46.0
53.2
Table 7 : Results of personalized calibration on MPIIFaceGaze DM using different numbers of calibration points (errors in mm).
Model (Backbone)
Model
→DM
→DE
Size
d
dϕ
dψ
d
dϕ
dψ
(a) Trained on DX only
ETH-XGaze (ResNet-50) [ 45 ]
25.6M
6.94 (7.50 ∗ )
5.15
3.75
8.87 (11.0 ∗ )
5.23
6.40
PureGaze (ResNet-50) [ 6 ]
31M
6.85 (7.08 ∗ )
4.91
3.89
7.40 (7.48 ∗ )
4.67
4.75
CDG (ResNet-50) [ 36 ]
25.6M
6.73 ∗
-
-
7.95 ∗
-
-
FSCI (ResNet-18) [ 20 ]
16.4M
6.93 (5.79 ∗ )
4.64
4.32
7.94 (6.96 ∗ )
6.12
3.88
Table 8 : Cross-domain evaluation results on MPIIFaceGaze DM and EyeDiap DE (angular errors in degrees; dϕ and dψ are the pitch and yaw components respectively).
Model
Overall
Ideal
Side-Lit
Glasses
Masks
dX
dY
∥d∥2
dX
dY
∥d∥2
dX
dY
∥d∥2
dX
dY
∥d∥2
dX
dY
∥d∥2
(a) Variance Loss
Ours (MobileNet, baseline)
22.3
34.9
46.3
16.0
28.8
36.6
18.6
28.5
37.0
25.4
33.4
46.6
20.9
36.0
45.3
Replacing τ=0.5 with variance loss
21.9
38.0
47.9
17.0
33.9
41.2
22.2
28.1
39.6
24.9
42.6
54.3
18.8
40.7
48.6
(b) Pre-training
Ours (ViT-B, MAE-pretrained)
19.1
35.8
44.4
14.9
29.3
36.2
17.7
27.9
35.9
24.4
33.5
46.0
18.8
36.5
45.3
Table 9 : Ablation studies on RealGaze (errors in mm).
Gaze target estimation, the task of predicting where a person is looking in a scene, is crucial to understanding human attention and intent. It is a challenging task that combines high-level understanding of global scene semantics and precise spatial reasoning using human appearance (e.g. pose, eye orientation). As a result, human-level performance remains elusive for existing models, limiting their practical application. To this end, we propose PaGE (Practical Gaze Estimator), a gaze estimation model that explicitly models the complex interaction between scene and head features. Using a PaGE model with a large ViT-H+ backbone as the teacher, we further distill student models with lighter backbones on a much larger and more diverse unlabeled dataset. The architectural improvements and novel training recipe allow PaGE to achieve state-of-the-art performance on several gaze estimation tasks, outperforming humans in 7 out of 9 metrics while reducing the human-AI gap by at least 60% in the remaining 2. The distilled student models retain most of the teacher's performance while being lightweight enough for practical deployment on robots and consumer devices. The code and model checkpoints are available at our project page.
Appearance-based gaze estimation always suffers from poor generalization due to limited annotated samples and insufficient dataset diversity. Leading approaches adopt weakly supervised learning to generate large-scale pseudo-labeled data from unconstrained real-world scenarios, aiming to mitigate the domain shifts. In this work, we devise a simple yet effective semi-supervised learning architecture that leverages unlabeled data to enhance domain generalization, thereby reducing reliance on labor-intensive manual annotations. Our key insight is to impose Jacobian regularization to disentangle feature representations into discriminative subspaces dedicated to specific gaze components, such as pitch and yaw angles. We further exploit the intrinsic ordinal ranking within each subspace for contrastive learning, enabling the model to learn robust gaze representations from a small set of labeled samples and an abundance of unlabeled ones. This ultimately yields our Disentangled Subspace Contrastive Learning (DSCL) framework. Extensive experiments on multiple benchmarks verify that the proposed DSCL is plug-and-play, achieving competitive performance using only 20%, 10%, and even 5% of the annotated data under both in-domain and cross-domain evaluation settings. The public code is available at https://github.com/da60266/DSCL.
Qida Tan, Hongyu Yang, Wenchao Du
National Key Laboratory of Fundamental Science on Synthetic Vision, Sichuan University, Chengdu, China · College of Computer Science, Sichuan University, Chengdu, China
Gaze estimation under natural head-eye motion underpins applications from driver monitoring to human-computer interaction. Single-frame methods predict each frame independently, so consecutive outputs fluctuate as jitter. Multi-frame methods reduce this, but they learn motion implicitly inside appearance features, so the gaze trajectory is never an explicit variable. We propose EyeTAG (Eye Trajectory-Aware Gaze Estimation), a causal multi-frame framework built around an explicit first-order gaze prior: at each step it differentiates its own recent predictions and feeds the resulting trajectory back as a compact kinematic token. Because differencing is translation-invariant in gaze space, this token carries subject-invariant motion rather than personal gaze offsets. Face and eye streams supply visual evidence, fused by cross-attention and a causal Transformer decoder. EyeTAG reduces the mean angular error by about 1.0∘ on Gaze360 and performs on par with the strongest baseline on EVE (2.56∘ vs. 2.58∘). Within-model ablations, which keep the encoder and the rest of the architecture fixed and vary only the gaze history, show that the differential formulation, rather than temporal context alone, removes the systematic saccade bias that persists even with an absolute gaze-history prior. Our code is available at https://github.com/peter8366/EyeTAG.
Jungmin Lee, Niamat Ullah, Yoseob Han
Department of Information and Telecommunication Engineering Soongsil University Seoul, Republic of Korea · Department of Electronic Engineering Soongsil University Seoul, Republic of Korea