HINT-Blimp: Human INTent Inference from Multimodal Cues for Robotic Blimps
Authors: Subhadeep Koley, Benjamin Greenberg, Yifei Simon Shao, Juan Aceros, Nadia Figueroa, David Saldaña
Organizations: Lehigh University, Bethlehem, PA 18015, USA · Swarthmore College, Swarthmore, PA 19081, USA · University of Pennsylvania, Philadelphia, PA 19104, USA
In human-robot interaction, traditional interfaces such as joysticks and handheld tablets introduce latency into navigation tasks and require the operator's explicit attention on the device, instead of the robot. We propose a new human-robot interaction framework in which a human communicates intent directly through sparse multimodal signals such as physical pushes and spoken commands. Human intent is represented as a parameterized linear dynamical system (LDS) that encodes the desired goal and motion behavior. The robot estimates this intent (parameters) online using a particle filter, where each particle represents a candidate LDS hypothesis and is reweighted online as new information becomes available. We validate this framework on a robotic blimp, whose inherent compliance and collision tolerance make it well-suited for repeated physical interaction. Experiments with multiple participants across 300 trials show that combining pushes and voice commands identifies the intended goal in 86% of trials within at most five interactions, with most trials resolved in two. The inferred dynamical systems can also produce curved trajectories that avoid obstacles known only to the human.
Figures & tables
Fig. 1 : HINT-Blimp infers human intent from multimodal interactions. Human–robot interaction occurs exclusively through physical pushes and voice commands. In this scenario, the user intends to send the blimp to Goal 2 but cannot push it directly toward the goal because a straight-line trajectory would collide with the obstacle. Instead, the user first pushes the blimp toward Goal 1 and then redirects its motion using the verbal command “Go right!”. Our probabilistic method (right column) maintains beliefs over candidate trajectories to each goal and progressively converges to the correct intent as additional information is provided by the user.
Fig. 2 : Examples for vector fields and robot trajectories for LDS. From left to right, the LDS parameters are (ϑ1,ϑ2,r)=(0.58,103∘,133∘) , (0.90,117∘,70∘) , and (0.47,117∘,70∘) .
Fig. 3 : Setup for Experiment 1 where the user has to send the robot to one of the 8 candidate goals selected randomly at the beginning of the run.
Fig. 5 : Distribution of the number of interactions used in successful trials, pooled across all five participants. Each participant completed 20 trials per modality, giving 100 trials per modality across the five participants
Fig. 6 : Interactions to successful commitment under three prior conditions ( n=40 each). The adversarial prior, despite concentrating 40% of particles on the farthest goal, matches the uniform baseline, while the helpful prior shows a modest reduction.
Embodied foundation models have achieved significant breakthroughs in robotic manipulation, yet they still depend heavily on large-scale robot demonstrations. Although recent works have explored leveraging human data to alleviate this dependency, effectively extracting transferable knowledge remains a significant challenge due to the inherent embodiment gap between human and robot. We argue that the intention underlying human actions can serve as a powerful intermediate representation for bridging this gap. In this paper, we introduce a novel framework that explicitly learns and transfers human intention to facilitate robotic manipulation. Specifically, we model intention through gaze, as it naturally precedes physical actions and serves as an observable proxy for human intent. Our model is first pretrained on a large-scale egocentric human dataset to capture human intention and its synergy with action, followed by finetuning on a small set of robot and human data. During inference, the model adopts a Chain-of-Thought reasoning paradigm, sequentially predicting intention before executing the action. Extensive evaluations in simulation and real-world settings, across long-horizon and fine-grained tasks, and under few-shot and robustness benchmarks, show that our method consistently outperforms strong baselines, generalizes better, and achieves state-of-the-art performance. Project page: https://gazevla.github.io .
Chengyang Li, Kaiyi Xiong, Yuan Xu +3
Shanghai Jiao Tong University · Peking University · ShanghaiTech University +1
Gaze and speech are promising interaction modalities for individuals with motor impairments, yet robust intent recognition in multi-object environments remains challenging due to micro-saccades, semantic ambiguity, and viewpoint changes. This paper presents a multimodal interaction framework for assistive robotic manipulation. We propose a sticky-glance algorithm that stabilizes gaze-based intent by jointly accumulating geometric distance and directional evidence, enabling robust real-time target selection and switching. We further introduce Glance-Say, a gaze-speech interaction paradigm in which gaze specifies objects and speech specifies actions, together with a continuous shared-control scheme that provides high-readiness robot motion and human-in-the-loop feedback. Experiments demonstrate a tracking rate of 0.92 for moving targets, selection accuracy of 0.97 for static targets, and reduced task duration. These results indicate improved robustness, efficiency, and usability over representative interaction paradigms.
Yuzhi Lai, Shenghai Yuan, Peizheng Li +2
University of Tuebingen, Geschwister-Scholl-Platz, 72074 Germany · Nanyang Technological University, 50 Nanyang Avenue, Singapore 639798
Improving the effectiveness of human-robot interaction requires social robots to accurately infer human goals through robust intention understanding. This challenge is particularly critical in multimodal settings, where agents must integrate heterogeneous signals including text, visual cues to form a coherent interpretation of user intent. This paper presents IntentVLM, a novel two-stage video-language framework designed for open-vocabulary human intention recognition. The approach is inspired by forward-inverse modeling in cognitive science by decomposing intention understanding into goal candidate generation followed by structured inference through selection, effectively reducing hallucinations in latent reasoning. Evaluated on the IntentQA and Inst-IT Bench datasets, IntentVLM achieves state-of-the-art results with up to 80% accuracy, notably surpassing the baseline performance by 30% and matches human performance. Our findings demonstrate that this structured reasoning approach enhances open-vocabulary intention understanding without catastrophic forgetting, offering a robust foundation for human-centered robotics.