cs.CVOct 6, 2026

Knowing When to Trust a Prior: Reliability-Gated Cue Fusion for Video Gaze Prediction

Authors: Lichen Zhu, Yueqian Lin, Yiheng Wang, Hai "Helen" Li, Yiran Chen

Organizations: Duke University, Durham, NC, USA

Abstract

Video gaze prediction is led by gaze-trained models, yet gaze-free priors carry signal those models have not absorbed, if one knows when to trust them. We propose FocusGate, a gated ensemble of gaze-free priors whose members may abstain. A per-frame gate reads three shape statistics of a defocus map and selects the frames on which the estimator is above chance on average, so rejected frames reduce to the base exactly, while midrank normalisation lets an all-zero prior abstain at zero parameters. Gated fusion is significantly positive on film, sports and web video, whereas unconditional fusion is harmful on sports and null on web. Added to four supervised predictors, the NTIRE 2026 champion among them, FocusGate improves all sixteen model-domain cells in shuffled AUC, fifteen significantly, one domain pre-registered and scored once, while adding only 1% to the champion's latency. Alone, it surpasses TASED-Net and UNISAL in shuffled AUC on film with a 16-frame causal mean.

Figures & tables

Explore similar work

CardsList
  1. GazeFlow: From Human Gaze Behavior to Generative Egocentric Gaze Prediction

    Sep 29, 2026Sheng Zhao, Weikai Lin, Yuhao ZhuGaze Target EstimationGaze Behavior

  2. EyeTAG: Eye Trajectory-Aware Gaze Estimation

    Oct 1, 2026Jungmin Lee, Niamat Ullah, Yoseob HanGaze Target Estimation

  3. Factor-Informed Uncertainty Distillation for Gaze Estimation

    Jul 22, 2026Mohammadreza Jamalifard, Yaxiong Lei, Javier Fumanal Idocin +3Gaze Target EstimationGaze Behavior