Video gaze prediction is led by gaze-trained models, yet gaze-free priors carry signal those models have not absorbed, if one knows when to trust them. We propose FocusGate, a gated ensemble of gaze-free priors whose members may abstain. A per-frame gate reads three shape statistics of a defocus map and selects the frames on which the estimator is above chance on average, so rejected frames reduce to the base exactly, while midrank normalisation lets an all-zero prior abstain at zero parameters. Gated fusion is significantly positive on film, sports and web video, whereas unconditional fusion is harmful on sports and null on web. Added to four supervised predictors, the NTIRE 2026 champion among them, FocusGate improves all sixteen model-domain cells in shuffled AUC, fifteen significantly, one domain pre-registered and scored once, while adding only 1% to the champion's latency. Alone, it surpasses TASED-Net and UNISAL in shuffled AUC on film with a 16-frame causal mean.
Figures & tables
Figure 1: FocusGate. (a) Eight gaze-free priors (fitted weights wk ) are rank-summed into the base B . The defocus map D enters only through the gate ct , which thresholds three shape statistics of D and never sees the fixations. (b) The gate opens on a bimodal D and closes on a unimodal one, where the prediction reduces to B exactly.
predictor
year
ms/frame
Film (H2)
Sports (UCF)
Web (DHF1K)
DIEM ‡
centre prior (chance calibration)
—
< 0.1
49.9
51.4
50.6
—
flow residual, camera-comp.
—
19.8
63.7
69.9
60.2
—
face detector [ 27 ]
2023
12.1
65.3
52.9
53.9
—
depth, MiDaS [ 21 ]
2022
11.1
61.5
63.2
58.9
—
DINOv2 attention [ 20 , 6 ]
2024
51.5
71.0
74.6
66.9
—
defocus, recipe estimator
ours
15.4
57.6
50.7
54.8
—
Table 1: sAUC (%) of every predictor under one protocol on held-out video (Sec. 3). ms/frame per component on one A100 (batch 1, one forward pass per predicted frame, two for ViSAGE’s expert pair). A “ + FocusGate” row adds 111 ms to its model. Bold with ∗ : paired 95% cluster CI vs. the model’s own row excludes zero and survives a Holm correction over the sixteen cells. ‡ : confirmatory domain, pre-registered and scored once (the ViNet-S cell was added afterwards and is exploratory). These cells measure the ensemble, not the gate, whose pre-specified contrast is null (Sec. 4.2 ).
cue sAUC (%)
Δ vs. base
cue, estimator
rate
alone
on
off
ungated
gated
defocus, classical
0%
45.6
—
45.6
−1.88 †
+0.00
defocus, scratch
38%
52.2
58.5
48.0
−0.94 †
+0.04
defocus, recipe
40%
57.6
62.9
53.9
+0.36 ∗
+0.60 ∗
frozen, sports
38%
50.7
55.7
46.9
−1.40 †
+0.68 ∗
frozen, web
55%
54.8
57.5
51.4
−0.08
+0.57 ∗
Table 2: The gate on seven cue-estimator pairs (H2 test split, reference base centre+face+SR, w=0.5 ). Shaded: the defocus cue that FocusGate gates, with the recipe estimator also transferred frozen. Unshaded: the same rule on other cues (motion: spread ≥0.1 , moving fraction ≤0.35 , train-selected). rate: accept rate. on/off: accepted/rejected frames.
Figure 2: The collapse regime. Standalone sAUC of each cue on the 1,715 test frames by quintile of the spread of its own map (shaded: 95% CI). Defocus maps sit at chance when unimodal and rise when bimodal, while sharpness and motion never approach chance.
Egocentric gaze prediction enables many downstream applications but remains challenging, as human gaze is inherently stochastic. This stochasticity is constrained by structured temporal dynamics alternating between fixations and saccades, top-down influences from tasks, and bottom-up visual saliency. Based on this observation, we introduce GazeFlow, a framework that directly models gaze as a joint distribution of temporal gaze positions conditioned upon both top-down and bottom-up information. In particular, GazeFlow uses conditional flow matching (CFM): a learned velocity field iteratively transports a Gaussian noise sample into a plausible gaze trajectory drawn from this joint distribution. The velocity field is conditioned on bottom-up visual features extracted by a video encoder and on top-down task information obtained by globally querying these features. On standard datasets, GazeFlow achieves state-of-the-art performance on per-frame metrics, and the generated trajectories align better with human gaze temporal dynamics.
Sheng Zhao, Weikai Lin, Yuhao Zhu
Department of Computer Science University of Rochester
Gaze estimation under natural head-eye motion underpins applications from driver monitoring to human-computer interaction. Single-frame methods predict each frame independently, so consecutive outputs fluctuate as jitter. Multi-frame methods reduce this, but they learn motion implicitly inside appearance features, so the gaze trajectory is never an explicit variable. We propose EyeTAG (Eye Trajectory-Aware Gaze Estimation), a causal multi-frame framework built around an explicit first-order gaze prior: at each step it differentiates its own recent predictions and feeds the resulting trajectory back as a compact kinematic token. Because differencing is translation-invariant in gaze space, this token carries subject-invariant motion rather than personal gaze offsets. Face and eye streams supply visual evidence, fused by cross-attention and a causal Transformer decoder. EyeTAG reduces the mean angular error by about 1.0∘ on Gaze360 and performs on par with the strongest baseline on EVE (2.56∘ vs. 2.58∘). Within-model ablations, which keep the encoder and the rest of the architecture fixed and vary only the gaze history, show that the differential formulation, rather than temporal context alone, removes the systematic saccade bias that persists even with an absolute gaze-history prior. Our code is available at https://github.com/peter8366/EyeTAG.
Jungmin Lee, Niamat Ullah, Yoseob Han
Department of Information and Telecommunication Engineering Soongsil University Seoul, Republic of Korea · Department of Electronic Engineering Soongsil University Seoul, Republic of Korea
Deep gaze estimation works well in controlled capture but degrades in unconstrained settings, where systems must reject unreliable predictions. Single-pass uncertainty (e.g., heteroscedastic regression) infers uncertainty from pixels without explicit input-validity cues, while sampling based methods are often too costly for real time use. We propose Factor-Informed Uncertainty Distillation (FIUD), a teacher-student framework that aligns uncertainty with interpretable image-quality failure modes. A gradient-boosting teacher predicts expected gaze error from factors such as illumination, sharpness, eye visibility and symmetry; a neural student distills these signals via curriculum learning and ranking supervision into a lightweight single-pass uncertainty head. Across ETH-XGaze, Gaze360, and MPIIFaceGaze (>300k samples), FIUD improves uncertainty, error rank correlation and selective prediction versus deterministic and sampling-based baselines, with the largest gains in unconstrained settings.
Mohammadreza Jamalifard, Yaxiong Lei, Javier Fumanal Idocin +3