cs.ROMar 6, 2026

Glance-Say: Multimodal Human-Robot Collaboration and Intent Recognition via Sticky Glance

Authors: Yuzhi LaiShenghai YuanPeizheng LiBenjamin KieferAndreas Zell

Abstract

Gaze and speech are promising interaction modalities for individuals with motor impairments, yet robust intent recognition in multi-object environments remains challenging due to micro-saccades, semantic ambiguity, and viewpoint changes. This paper presents a multimodal interaction framework for assistive robotic manipulation. We propose a sticky-glance algorithm that stabilizes gaze-based intent by jointly accumulating geometric distance and directional evidence, enabling robust real-time target selection and switching. We further introduce Glance-Say, a gaze-speech interaction paradigm in which gaze specifies objects and speech specifies actions, together with a continuous shared-control scheme that provides high-readiness robot motion and human-in-the-loop feedback. Experiments demonstrate a tracking rate of 0.92 for moving targets, selection accuracy of 0.97 for static targets, and reduced task duration. These results indicate improved robustness, efficiency, and usability over representative interaction paradigms.

Explore similar work

May 28, 2026cs.RO

Gaze2Act: Gaze-Conditioned Vision-Language-Action Policies for Interactive Robot Manipulation

Vision-Language-Action (VLA) models have recently shown strong potential for robot learning by following language instructions. However, in practice, language alone is often insufficient to precisely convey human intent. It is difficult to describe which exact object to interact with among similar candidates, where to act on the object, or how the target may change during execution. To address this limitation, we propose Gaze2Act, a novel VLA framework that leverages human gaze as a dynamic and intuitive intent signal for complex interactive manipulation. Gaze2Act first bridges the ego-exo view gap by mapping first-person gaze into the robot's perspective through cross-view semantic matching, producing both an object mask and a gaze point for coarse-to-fine target specification. These cues are then integrated into the policy through perception-level prompting and action-level conditioning, allowing the robot to attend to relevant regions and execute precise interactions under dynamic intent. In a systematic evaluation across seven task categories and 16 real-robot tasks on a Unitree G1 humanoid, Gaze2Act achieves state-of-the-art performance in both intent accuracy and task success rate. It notably outperforms baselines in object disambiguation, fine-grained interaction, and dynamic intent steering. These results demonstrate that human gaze provides a natural, low-burden, and highly expressive modality for human-in-the-loop VLA control.
Kuangji Zuo, Gen Li, Bofan Lyu +9
Jun 20, 2026cs.RO

Predictive Gaze Is Preserved but Reorganized toward Monitoring during Robot-Mediated Manipulation

Goal-directed eye movements are a fundamental component of visuomotor control, enabling humans to anticipate and guide their actions. Whether this anticipatory and task-driven behavior is preserved when actions are executed through a robot rather than through one's own body remains unclear. Here we address this question by investigating gaze behavior during goal-directed telemanipulation to determine how visuomotor control adapts to altered embodiment. Our findings show that gaze remains strongly aligned with task goals, preserving its predictive role even during robot-mediated manipulation. At the same time, teleoperation systematically redistributes visual attention toward the robotic end-effector and manipulated objects, increasing online monitoring. These findings show that predictive gaze is not lost under altered embodiment, but reorganized in response to changes in sensory feedback and control demands. More broadly, they reveal the flexibility of the human visuomotor system when the natural sensorimotor coupling is disrupted and identify gaze as an informative signal for inferring action intentions in human-robot interaction.
Manuela Uliano, Silvia Fattorini, Marco Controzzi
Jun 11, 2026cs.RO

GIVE: Grounding Human Gestures in Vision-Language-Action Models

Human communication is inherently multimodal, where language is often accompanied by non-verbal cues such as gestures to convey intentions. However, current Vision-Language-Action (VLA) models treat robotic manipulation as a pure text-driven task, overlooking the important role of gestures in Human-Robot Interaction (HRI). This often leads to inaccurate intent grounding and unreliable manipulation when language instructions are ambiguous or underspecified. To address this challenge, we propose GIVE (Gesture Intent via Visual-Semantic Enhancement), an effective approach that enhances pre-trained VLA models with human gesture understanding without architectural modifications. Specifically, GIVE incorporates gesture information through two complementary pathways: a visual pathway that overlays hand skeletons and fingertip rays onto robot observations for explicit object grounding, and a semantic pathway that generates high-level descriptions of human gestures and task instructions for robust intent grounding. By jointly leveraging visual and semantic guidance, GIVE enables VLA policies to better associate gestures with manipulation behaviors and adapt to dynamic interaction intents. In real-world HRI experiments, GIVE substantially outperforms the baseline, improving target object recognition accuracy by 40% and overall task success rate by 80%, while demonstrating strong robustness and generalization to unseen spatial layouts and diverse participants.
Pengfei Liu, Gen Li, Junqiao Fan +4