cs.ROSep 28, 2026

Gaze Prompts: Temporally Dense Human Attention for Vision-Language-Action Fine-Tuning

Authors: Yihan Zhou, Rui Yan, Mingcong Li, Zheyuan Huang, Xu Yang, Xueyang Guo, Yilin Mo

Organizations: Department of Automation, Tsinghua University · Beijing Key Laboratory of Embodied Intelligence Systems · Institute for Embodied Intelligence and Robotics, Tsinghua University · LingYu Robotics

Abstract

Vision-Language-Action (VLA) fine-tuning pairs images with actions at every step, yet typically provides only a task-level language instruction, leaving moment-to-moment visual relevance implicit. We introduce \emph{eye-tracker-supervised gaze prompting}, which uses gaze recorded during VR teleoperation to provide frame-level visual guidance for VLA fine-tuning. During training, recorded gaze locations are rendered as crosshairs on the robot's head-camera images. At deployment, a lightweight predictor estimates gaze locations from recent images and the instruction, supplying the same type of visual prompt without an eye tracker or changes to the policy architecture. Instantiated with π0π_0, gaze prompting increases mean success from 26.3%26.3\% to 56.0%56.0\% across six real-world bimanual manipulation tasks, with gains also observed when a single policy is trained on all six tasks. We release \textsc{GazeMani}, a dataset of 1,2001{,}200 teleoperated trajectories with synchronized gaze.

Explore similar work

CardsList