Egocentric gaze prediction enables many downstream applications but remains challenging, as human gaze is inherently stochastic. This stochasticity is constrained by structured temporal dynamics alternating between fixations and saccades, top-down influences from tasks, and bottom-up visual saliency. Based on this observation, we introduce GazeFlow, a framework that directly models gaze as a joint distribution of temporal gaze positions conditioned upon both top-down and bottom-up information. In particular, GazeFlow uses conditional flow matching (CFM): a learned velocity field iteratively transports a Gaussian noise sample into a plausible gaze trajectory drawn from this joint distribution. The velocity field is conditioned on bottom-up visual features extracted by a video encoder and on top-down task information obtained by globally querying these features. On standard datasets, GazeFlow achieves state-of-the-art performance on per-frame metrics, and the generated trajectories align better with human gaze temporal dynamics.
Figures & tables
Figure. 1: Qualitative comparison on EGTEA Gaze+. Example frames where GazeFlow predicts a heatmap tightly localized on the ground-truth gaze positions (red dots) compared to existing methods (AT [ 17 ] , GLC [ 26 ] , EgoM2P [ 31 ] ).
Human gaze mechanism
Method
Paradigm
Prediction
Training objective
Bottom-up
Top-down
Joint
Stochastic
Li et al. [32]
Regression
gt∣V
per-frame MSE
×
∼
×
×
GLC [ 26 ]
Marginal density estimation
qθ(gt∣V)
per-frame KL
✓
×
×
✓
DFG [ 54 ]
qθ(gt∣V)
per-frame KL
✓
×
∼
✓
AT [ 17 ]
Autoregressive
qθ(gt∣gt−1,V≤t)
per-frame BCE
✓
∼
∼
✓
EgoM2P [ 31 ]
qθ(gt∣g<t,V)
masked L2
∼
∼
∼
✓
Table 1: Comparison of egocentric gaze prediction paradigms. V is the input video and V≤t its causal prefix, gt∈R2 is the gaze position at frame t , g1:T the full trajectory, and qθ(⋅∣⋅) denotes a model’s parametric conditional distribution. MSE: mean squared error, KL: Kullback–Leibler divergence, and BCE: binary cross-entropy. ✓ if satisfied, × if not satisfied, and ∼ if architecturally implied but not enforced.
Figure. 2: GazeFlow architecture. (a) Dual-path condition extraction (Section 3.1 ). A LoRA-adapted V-JEPA 2 backbone produces bottom-up visual tokens cv ; spatial–temporal pooling and learnable queries form the top-down task-token bank ctask . (b) Conditional velocity network (Section 3.2 ). Latent gaze tokens encoding xs and s pass through six DiT blocks, each containing four sub-layers; a linear head outputs the velocity field vθ . (c) Flow-matching sampling dynamics. Euler integration of vθ transports Gaussian noise x0 to a clean trajectory x1 ; the cyan sample cloud at s∈{0,0.2,0.4,0.8,1} contracts toward the ground truth.
EGTEA Gaze+
Ego4D
Paradigm
Method
AUC ↑
F1 ↑
P ↑
R ↑
AAE ↓
AUC ↑
F1 ↑
P ↑
R ↑
AAE ↓
Simple Priors
Center Bias
0.906
0.185
0.162
0.215
16.07
0.930
0.222
0.194
0.258
13.92
Random Walk
0.824
0.147
0.090
0.413
16.34
0.763
0.109
0.059
0.662
14.38
Marginal density
GLC [ 26 ]
0.953
0.421
0.348
0.531
10.21
0.947
0.355
0.267
0.530
11.52
Autoregressive
EgoM2P (zero-shot) [ 31 ]
0.921
0.280
0.195
0.498
13.18
0.911
0.253
0.180
0.422
13.86
EgoM2P (LoRA) [ 31 ]
0.921
0.293
0.205
0.515
12.81
0.930
0.297
0.217
0.470
12.69
Table 2: Per-frame accuracy on EGTEA Gaze+ and Ego4D. GazeFlow achieves the best performance; results on more metrics are in Table 9 in Appendix F.1 . Best per column is in bold.
ADE ( ×102 px) ↓
DTW ( ×102 px) ↓
Disp Mean
Disp Med
JSD ↓
Paradigm
Method
Mean
Best
Mean
Best
( H : 12.5 px)
( H : 4.0 px)
Simple Priors
Center Bias
1.35
1.35
1.35
1.35
0.00×
0.00×
0.125
Random Walk
2.07
1.05
1.97
0.92
2.54×
7.43×
0.294
Marginal density
GLC [ 26 ]
1.13
1.01
1.03
0.87
5.77×
11.75×
0.352
Autoregressive
EgoM2P (zero-shot) [ 31 ]
1.47
0.97
1.35
0.83
2.20×
3.05×
0.058
EgoM2P (LoRA) [ 31 ]
1.50
0.96
1.37
0.81
2.69×
3.03×
0.058
Table 3: Alignment with human gaze dynamics on EGTEA Gaze+. Disp Mean and Disp Med are per-step gaze displacements shown as ratios to the human trajectory (subheader H , measured in pixels; closer to 1× is better). Best per column in bold.
Table 6
Variant
AUC ↑
F1 ↑
P ↑
R ↑
AAE ↓
ADE ↓
DTW ↓
Disp Mean
Disp Med
JSD ↓
(H: 12.5 px)
(H: 4.0 px)
Full
0.970
0.493
0.401
0.638
8.81
0.57
0.45
0.91×
1.08×
0.0011
- Bottom-up Visual Feature
0.961
0.431
0.333
0.613
9.22
0.65
0.53
0.89×
1.17×
0.0012
- Top-down Task Guidance
0.965
0.476
0.380
0.638
9.14
0.58
0.46
0.92×
1.09×
0.0013
- Temporal RoPE
0.968
0.477
0.388
0.620
8.94
0.59
0.46
1.21×
1.40×
0.0027
- Joint Distribution Modeling
0.903
0.451
0.394
0.526
9.12
0.84
0.79
0.71×
2.01×
0.0729
Table 6: Ablation study. Full is the default GazeFlow ; each remaining row ablates one architectural component while keeping all others identical, removing either the bottom-up visual feature ( cv , frame-local visual cross-attention), the top-down task guidance ( ctask , task cross-attention), the relative-time RoPE in time-aware self-attention, or the joint-distribution flow-matching head (replaced with a deterministic regression head over the same conditioning). Per-column best in bold .
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Figure. 3: Per-head patterns of the time-aware self-attention aggregated across the EGTEA Gaze+ validation set. Rows index denoiser layers, columns index attention heads. Each cell shows the canonical attention pattern (query frame on the vertical axis, source frame on the horizontal, both normalized to clip-relative position); brighter pixels correspond to higher attention weight after row-wise softmax. The dashed diagonal marks source=query , separating two halves with opposite temporal meaning: mass in the lower triangle ( source<query ) is attention to past frames (look-back), mass in the upper triangle ( source>query ) is attention to future frames (look-ahead), and mass on both sides marks split-attention heads with peaks straddling the current frame. Cell-frame color encodes this role: look-back (blue), look-ahead (red), split-attention (purple), local (grey).
Figure. 4: Per-head time-aware self-attention as a function of temporal offset. Same heads as Fig. 3 , summarized as a 1-D curve by averaging the attention pattern along each diagonal. The horizontal axis is the offset between source and query frame (negative = past, positive = future); the vertical dashed line marks zero offset. Look-ahead heads peak right of zero, look-back heads peak left, and split heads exhibit two peaks straddling zero with a clear valley at the current frame.
Figure. 5: Hand vs. active-object attention mass per (layer, head) on the EGTEA Gaze+ validation set. Each marker is one of the 48 visual-cross-attention heads, plotted by mean attention mass on the hand (x-axis) and active-object (y-axis) regions; markers are colored by layer, error bars are clip-wise standard deviation. Dotted lines mark the global means (hand =4.5% , object =22.6% ). Heads in the lower-right quadrant focus on the hand, those in the upper-left on the active object, and those in the upper-right couple both.
Figure. 6: Layer-wise hand / active-object / background attention mass conditioned on verb. For each of the six most frequent EGTEA verbs we average the per-region attention mass across heads at each layer. Manipulation verbs (“cut”, “put”, “take”, “turn”) keep 25 – 50% of the mass on the active object at every layer, while “open” and “close” route most mass to the background regardless of layer.
Figure. 7: Failure cases. Three rows from two validation clips, each row showing 8 evenly-spaced frames along time. Cyan dots are the K=50 GazeFlow predictions at that frame; the filled red circle is the recorded ground-truth gaze when it lies within the frame, and the hollow red ring with a yellow halo flags frames where the recording is clamped to the image boundary (off-frame). When the wearer saccades to a target outside the camera’s view, the predicted cloud stays on the visible scene while the ground-truth marker drifts to the boundary.
Backbone
Resolution
Layer
AUC ↑
F1 ↑
P ↑
R ↑
AAE ↓
V-JEPA 2 ViT-L
256×256 (native)
last
0.944
0.378
0.292
0.535
11.39
V-JEPA 2 ViT-L
240×320 (matched)
last
0.913
0.296
0.216
0.469
12.87
Qwen3-VL-8B
native
2
0.944
0.370
0.280
0.549
11.37
Qwen3-VL-8B
native
4
0.944
0.370
0.280
0.547
11.37
Qwen3-VL-8B
native
8
0.943
0.369
0.278
0.549
11.34
Qwen3-VL-8B
native
16
0.942
0.368
0.279
0.541
11.24
Appendix
Table 7: Linear probing of frozen backbone features on EGTEA Gaze+ . Top block: V-JEPA 2 ViT-L at its native pre-training resolution and at a 4:3 input that matches the dataset aspect ratio. Bottom block: Qwen3-VL-8B at four LM layers, native resolution. All probes are ridge regressions on the frozen spatial token grid; per-column best in bold.
Dataset
Resolution
fps
Train clips
Val clips
EGTEA Gaze+
640×480
24
8,299
2,022
Ego4D (Aria gaze)
1088×1080
30
12,178
5,202
Appendix
Table 8: Dataset summary.
EGTEA Gaze+
Ego4D
Saliency
Localization
Saliency
Localization
Method
NSS ↑
AUC ↑
CC ↑
KL ↓
SIM ↑
F1 ↑
P ↑
R ↑
AAE ↓
NSS ↑
AUC ↑
CC ↑
KL ↓
SIM ↑
F1 ↑
P ↑
R ↑
AAE ↓
Prior-only reference
Center Bias
0.679
0.906
0.095
6.635
0.110
0.185
0.162
0.215
16.07
0.780
0.930
0.122
5.873
0.136
0.222
0.194
0.258
13.92
Random Walk
0.623
0.824
0.096
3.926
0.101
0.147
0.090
0.413
16.34
0.113
0.763
0.017
4.256
0.072
0.109
0.059
0.662
14.38
Marginal density estimation
Appendix
Table 9: Full saliency and localization comparison on EGTEA Gaze+ and Ego4D.
F1 ↑
Time (s) ↓
K \ S
1
2
5
10
50
1
2
5
10
50
1
0.453
0.465
0.444
0.427
0.405
0.49
0.49
0.51
0.54
0.77
10
0.460
0.481
0.484
0.479
0.468
0.52
0.53
0.57
0.63
1.14
30
0.460
0.484
0.493
0.492
0.486
0.58
0.62
0.73
0.91
2.35
50
0.460
0.485
0.495
0.495
0.491
0.65
0.71
0.89
1.21
3.37
Appendix
Table 10: F1 and inference time over the sample count K and the Euler step count S (EGTEA Gaze+, GazeFlow LoRA). Time is per 64 -frame clip on a single H100. The recommended setting is in bold.
Method
ms / frame ↓
Throughput (fps) ↑
GLC [ 26 ]
12.5
80
AT [ 17 ]
219
4.6
EgoM2P [ 31 ]
506
2.0
GazeFlow ( K=50 , S=50 )
53
19
GazeFlow ( K=30 , S=5 )
11.4
88
Appendix
Table 11: Inference time across methods on a single H100. The ms / frame column gives the time to predict gaze for one video frame, and throughput is its inverse.
ADE ( ×102 px) ↓
DTW ( ×102 px) ↓
K
AUC ↑
F1 ↑
P ↑
R ↑
AAE ↓
Mean
Best
Mean
Best
1
0.866
0.398
0.350
0.462
11.10
1.05
1.05
0.96
0.96
5
0.942
0.457
0.378
0.578
9.61
1.05
0.80
0.95
0.68
10
0.953
0.469
0.395
0.579
9.28
1.05
0.73
0.95
0.60
30
0.961
0.487
0.399
0.623
9.05
1.04
0.64
0.95
0.51
50
0.964
0.491
0.414
0.603
9.01
1.04
0.60
0.95
0.48
Appendix
Table 12: Sample count K sweep on EGTEA Gaze+. ADE and DTW are reported in ×102 px following Table 3 ; “Mean” and “Best” are mean-over- K and best-of- K . At K=1 the two coincide. The K=50 row matches Table 2 and Table 3 .
M
AUC ↑
F1 ↑
P ↑
R ↑
AAE ↓
1
0.968
0.492
0.404
0.630
8.68
2
0.967
0.486
0.392
0.639
8.87
4
0.970
0.493
0.401
0.638
8.81
8
0.968
0.483
0.390
0.635
8.79
16
0.968
0.480
0.381
0.650
8.75
Appendix
Table 13: Capacity sweep over learnable task queries. M is the number of learnable task queries, and every variant keeps the global pool token. Per-column best in bold.
Gaze estimation under natural head-eye motion underpins applications from driver monitoring to human-computer interaction. Single-frame methods predict each frame independently, so consecutive outputs fluctuate as jitter. Multi-frame methods reduce this, but they learn motion implicitly inside appearance features, so the gaze trajectory is never an explicit variable. We propose EyeTAG (Eye Trajectory-Aware Gaze Estimation), a causal multi-frame framework built around an explicit first-order gaze prior: at each step it differentiates its own recent predictions and feeds the resulting trajectory back as a compact kinematic token. Because differencing is translation-invariant in gaze space, this token carries subject-invariant motion rather than personal gaze offsets. Face and eye streams supply visual evidence, fused by cross-attention and a causal Transformer decoder. EyeTAG reduces the mean angular error by about 1.0∘ on Gaze360 and performs on par with the strongest baseline on EVE (2.56∘ vs. 2.58∘). Within-model ablations, which keep the encoder and the rest of the architecture fixed and vary only the gaze history, show that the differential formulation, rather than temporal context alone, removes the systematic saccade bias that persists even with an absolute gaze-history prior. Our code is available at https://github.com/peter8366/EyeTAG.
Jungmin Lee, Niamat Ullah, Yoseob Han
Department of Information and Telecommunication Engineering Soongsil University Seoul, Republic of Korea · Department of Electronic Engineering Soongsil University Seoul, Republic of Korea
We study the problem of human gaze modeling, which aims to generate the gaze patterns a viewer produces while observing a visual stimulus. Gaze is primarily captured through two modalities: continuous eye-tracking trajectories, which describe fine-grained motion dynamics, and discrete scanpaths, which describe high-level fixation structure. Because gaze varies substantially across viewers and trials, we treat this variability as a defining property rather than noise and model gaze as a stochastic generative process. Existing generative gaze models supervise on only one of these two representations in isolation. We hypothesize that trajectories and scanpaths describe gaze at complementary scales and are jointly informative during training, and test this hypothesis through ST-DiffEye, a joint trajectory-scanpath diffusion framework that couples both modalities by concatenating them as an additional raw input channel, requiring no architectural overhead beyond an input and output channel expansion. We further introduce a principled evaluation framework based on the Continuous Ranked Probability Score (CRPS), which generalizes any existing sequence similarity metric into a proper scoring rule that jointly assesses the accuracy and diversity of generated gaze. Experiments on task-driven visual search, covering both target-present and target-absent scenarios, and on free-viewing benchmarks demonstrate state-of-the-art performance. These results, along with detailed ablations, confirm the benefit of joint modeling and the value of distribution-aware evaluation in capturing the intrinsic variability of human gaze. Project webpage: https://st-diffeye.github.io/
We present the first data-driven approach to model temporal gaze-head coordination from large-scale in-the-wild facial videos. To obtain training data for generalizable learning, we propose an automatic pipeline that extracts natural yet diverse gaze and head motions with off-the-shelf appearance-based gaze estimators. To capture the probabilistic correlation and temporal dynamics of gaze-head coordination, we build our model on a generative conditional Variational Autoencoder for plausible yet diverse gaze-conditioned head motion generations. We further apply our framework to gaze-controlled facial video generation, where we enable video generation with natural and realistic head motion correlated to the input gaze - an aspect that has not been emphasized before. Human evaluation and quantitative comparisons demonstrate our method's effectiveness and validate our design choices, with evaluators showing statistically significant preference for our approach over baseline methods.
Xiaohan Liu, Yilin Wen, Yusuke Sugano
Institute of Industrial Science, The University of Tokyo Komaba 4-6-1, Tokyo, Japan