EyeTAG: Eye Trajectory-Aware Gaze Estimation
Organizations: Department of Information and Telecommunication Engineering Soongsil University Seoul, Republic of Korea · Department of Electronic Engineering Soongsil University Seoul, Republic of Korea
Abstract
Gaze estimation under natural head-eye motion underpins applications from driver monitoring to human-computer interaction. Single-frame methods predict each frame independently, so consecutive outputs fluctuate as jitter. Multi-frame methods reduce this, but they learn motion implicitly inside appearance features, so the gaze trajectory is never an explicit variable. We propose EyeTAG (Eye Trajectory-Aware Gaze Estimation), a causal multi-frame framework built around an explicit first-order gaze prior: at each step it differentiates its own recent predictions and feeds the resulting trajectory back as a compact kinematic token. Because differencing is translation-invariant in gaze space, this token carries subject-invariant motion rather than personal gaze offsets. Face and eye streams supply visual evidence, fused by cross-attention and a causal Transformer decoder. EyeTAG reduces the mean angular error by about 1.0 on Gaze360 and performs on par with the strongest baseline on EVE (2.56 vs. 2.58). Within-model ablations, which keep the encoder and the rest of the architecture fixed and vary only the gaze history, show that the differential formulation, rather than temporal context alone, removes the systematic saccade bias that persists even with an absolute gaze-history prior. Our code is available at https://github.com/peter8366/EyeTAG.
Figures & tables
| Split | (a) Gaze360 Dataset [ Kellnhofer et al.(2019)Kellnhofer, Recasens, Stent, Matusik, and Torralba ] | (b) EVE Dataset [ Park et al.(2020)Park, Aksan, Zhang, and Hilliges ] | ||||
|---|---|---|---|---|---|---|
| # Clips | # Frames | Avg Length (s) | # Clips | # Frames | Avg Length (s) | |
| Train | 321 | 84,902 | 8.8 | 3,015 | 1,286,850 | 13.2 |
| Validation | 85 | 11,318 | 4.4 | 336 | 155,640 | 13.8 |
| Test | 68 | 16,031 | 7.9 | 387 | 169,860 | 14.0 |
| Type | Method | (a) Gaze360 Dataset [ Kellnhofer et al.(2019)Kellnhofer, Recasens, Stent, Matusik, and Torralba ] | (b) EVE Dataset [ Park et al.(2020)Park, Aksan, Zhang, and Hilliges ] | (c) Complexity | ||||
|---|---|---|---|---|---|---|---|---|
| All | Semi-Front | Front | Mean | All | Time (ms) | GFLOPs | ||
| SF | L2CS-Net [ Abdelrahman et al.(2023)Abdelrahman, Hempel, Khalifa, Al-Hamadi, and Dinges ] | 11.10 9.82 | 10.77 8.30 | 9.92 7.97 | 10.60 8.70 | 5.06 6.12 | 5.42 | 16.57 |
| GazeTR-Hybrid [ Cheng and Lu(2022) ] | 10.56 9.24 | 10.30 8.05 | 9.63 7.90 | 10.16 8.40 | 3.63 3.14 | 5.35 | 1.84 | |
| CrossGaze [ Catruna et al.(2024)Catruna, Cosma, and Radoi ] | 10.47 8.45 | 10.29 7.79 | 9.23 7.31 | 10.00 7.85 | 3.30 2.68 | 14.31 | 1.43 | |
| MF | Gaze360 [ Kellnhofer et al.(2019)Kellnhofer, Recasens, Stent, Matusik, and Torralba ] | 10.46 8.92 | 10.18 7.69 | 9.32 7.58 | 9.99 8.06 | 2.73 1.77 | 2.83 | 12.79 |
| STAGE [ Jindal et al.(2024)Jindal, Yadav, and Manduchi ] | 12.08 10.82 | 11.63 8.53 | 9.43 6.72 | 11.05 8.69 | 2.58 2.11 | 10.26 | 39.55 | |
| Type | Method | (a) Fixation Jitter | (b) Saccade Bias | (c) Absolute Mean |
|---|---|---|---|---|
| SF | L2CS-Net [ Abdelrahman et al.(2023)Abdelrahman, Hempel, Khalifa, Al-Hamadi, and Dinges ] | 5.60 0.16 | 0.02 0.17 | 2.81 0.17 |
| GazeTR-Hybrid [ Cheng and Lu(2022) ] | 5.48 0.16 | 0.84 0.18 | 3.16 0.17 | |
| CrossGaze [ Catruna et al.(2024)Catruna, Cosma, and Radoi ] | 5.06 0.14 | 0.23 0.15 | 2.65 0.15 | |
| MF | Gaze360 [ Kellnhofer et al.(2019)Kellnhofer, Recasens, Stent, Matusik, and Torralba ] | 1.71 0.10 | -4.24 0.17 | 2.98 0.14 |
| STAGE [ Jindal et al.(2024)Jindal, Yadav, and Manduchi ] | 3.81 0.11 | -0.89 0.26 | 2.35 0.19 | |
| EyeTAG (Visual-only) | 3.51 0.12 | -0.46 0.16 | 1.99 0.14 |
| Fusion Strategy | All | Semi-Front | Front | Mean | |
|---|---|---|---|---|---|
| (a) | Self-Attn + Concat (face eye) | 9.80 8.14 | 9.60 7.45 | 8.71 7.43 | 9.37 7.67 |
| (b) | Bi. Cross-Attn (face eye) | 9.66 8.16 | 9.49 7.55 | 8.62 7.49 | 9.26 7.73 |
| (Ours) | Uni. Cross-Attn (face eye) | 9.29 7.78 | 9.12 7.23 | 8.08 7.23 | 8.83 7.41 |
| Gaze History | All | Semi-Front | Front | Mean | |
|---|---|---|---|---|---|
| (a) | No Gaze History (visual-only) | 9.72 8.15 | 9.55 7.76 | 8.30 7.53 | 9.19 7.81 |
| (b) | Normal Gaze Prior ( ) | 9.63 7.94 | 9.46 7.38 | 8.22 7.07 | 9.10 7.46 |
| (Ours) | Kinematic Gaze Prior ( ) | 9.29 7.78 | 9.12 7.23 | 8.08 7.23 | 8.83 7.41 |
| 4 | 8 | 16 | 24 | 32 | 48 ∗ | 64 | |
|---|---|---|---|---|---|---|---|
| Mean | 9.87 7.57 | 9.49 7.36 | 9.47 7.41 | 9.35 7.56 | 8.91 7.39 | 8.83 7.41 | 8.82 7.40 |
| Time (ms) | 10.78 | 11.12 | 11.38 | 11.41 | 13.91 | 19.97 | 25.23 |
| GFLOPs | 10.13 | 20.26 | 40.53 | 60.79 | 81.05 | 121.57 | 162.09 |
| Pretraining | All | Semi-Front | Front | Mean | |
|---|---|---|---|---|---|
| (a) | From scratch | 20.41 16.39 | 18.03 15.54 | 16.05 14.50 | 18.16 15.47 |
| (b) | ImageNet [ Deng et al.(2009)Deng, Dong, Socher, Li, Li, and Fei-Fei ] | 11.69 10.22 | 11.34 8.64 | 10.15 8.50 | 11.06 9.12 |
| ( Ours ) | VGGFace2 [ Cao et al.(2018)Cao, Shen, Xie, Parkhi, and Zisserman ] | 9.29 7.78 | 9.12 7.23 | 8.08 7.23 | 8.83 7.41 |
| Eye Backbone | Type | All | Semi-Front | Front | Mean |
|---|---|---|---|---|---|
| EfficientNet-B0 [ Tan and Le(2019) ] | 2D, no temporal | 10.77 10.80 | 10.37 8.83 | 8.26 6.56 | 9.80 8.73 |
| TSM [ Lin et al.(2019)Lin, Gan, and Han ] | 2D + temporal | 9.47 7.89 | 9.30 7.30 | 8.31 7.35 | 9.03 7.51 |
| 3D CNN [ Tran et al.(2015)Tran, Bourdev, Fergus, Torresani, and Paluri ] (scratch) | 3D, temporal | 9.41 7.90 | 9.22 7.27 | 8.17 7.06 | 8.93 7.41 |
| ( Ours ) ResNet-18 (shared) | 2D, no temporal | 9.29 7.78 | 9.12 7.23 | 8.08 7.23 | 8.83 7.41 |
| Input: , window size |
| Output: |
| — Training (per epoch) — |
| 1. for each frame do |
| a. |
| b. Construct gaze history: |
| sample from or with ratio over ep. 1–5 |
| Gaze History Schedule | All | Semi-Front | Front | Mean | |
|---|---|---|---|---|---|
| (a) | Teacher forcing only ( ) | 39.53 32.45 | 38.98 31.81 | 31.64 22.37 | 36.71 28.87 |
| (b) | Fully autoregressive ( ) | 10.13 7.88 | 9.96 7.93 | 8.81 7.54 | 9.63 7.78 |
| ( Ours ) | Scheduled sampling ( ) | 9.63 7.94 | 9.46 7.38 | 8.22 7.07 | 9.10 7.46 |
| All | Semi-Front | Front | Mean | Time (ms) | GFLOPs | |
|---|---|---|---|---|---|---|
| 4 | 10.96 8.28 | 10.80 7.74 | 9.58 7.42 | 10.45 7.81 | 9.87 | 10.13 |
| 8 | 10.17 7.81 | 10.01 7.28 | 9.23 7.06 | 9.80 7.38 | 9.97 | 20.26 |
| 16 | 10.26 7.98 | 10.08 7.42 | 9.15 7.29 | 9.83 7.56 | 10.23 | 40.53 |
| 24 | 9.96 7.97 | 9.80 7.36 | 8.65 7.36 | 9.47 7.56 | 11.45 | 60.79 |
| 32 ∗ | 9.63 7.94 | 9.46 7.38 | 8.22 7.07 | 9.10 7.46 | 13.95 | 81.05 |
| 48 | 9.94 7.94 | 9.65 7.47 | 8.61 7.65 | 9.40 7.69 | 19.98 | 121.57 |
| Type | Method | All | Semi-Front | Front | Mean |
|---|---|---|---|---|---|
| SF | L2CS-Net [ Abdelrahman et al.(2023)Abdelrahman, Hempel, Khalifa, Al-Hamadi, and Dinges ] | 10.07 7.66 | 9.85 6.40 | 9.90 6.37 | 9.94 6.81 |
| GazeTR-Hybrid [ Cheng and Lu(2022) ] | 9.42 7.27 | 9.26 6.50 | 9.46 6.31 | 9.38 6.69 | |
| CrossGaze [ Catruna et al.(2024)Catruna, Cosma, and Radoi ] | 9.22 6.84 | 9.09 6.05 | 8.55 5.90 | 8.95 6.26 | |
| MF | Gaze360 [ Kellnhofer et al.(2019)Kellnhofer, Recasens, Stent, Matusik, and Torralba ] | 9.70 7.02 | 9.47 6.00 | 9.41 5.40 | 9.53 6.14 |
| STAGE [ Jindal et al.(2024)Jindal, Yadav, and Manduchi ] | 8.87 5.84 | 8.87 5.84 | 8.69 5.77 | 8.81 5.82 | |
| EyeTAG ( ) | 8.80 6.41 | 8.68 6.08 | 7.95 6.22 | 8.48 6.23 |