HINT-Blimp: Human INTent Inference from Multimodal Cues for Robotic Blimps
Authors: Subhadeep Koley, Benjamin Greenberg, Yifei Simon Shao, Juan Aceros, Nadia Figueroa, David Saldaña
Organizations: Lehigh University, Bethlehem, PA 18015, USA · Swarthmore College, Swarthmore, PA 19081, USA · University of Pennsylvania, Philadelphia, PA 19104, USA
In human-robot interaction, traditional interfaces such as joysticks and handheld tablets introduce latency into navigation tasks and require the operator's explicit attention on the device, instead of the robot. We propose a new human-robot interaction framework in which a human communicates intent directly through sparse multimodal signals such as physical pushes and spoken commands. Human intent is represented as a parameterized linear dynamical system (LDS) that encodes the desired goal and motion behavior. The robot estimates this intent (parameters) online using a particle filter, where each particle represents a candidate LDS hypothesis and is reweighted online as new information becomes available. We validate this framework on a robotic blimp, whose inherent compliance and collision tolerance make it well-suited for repeated physical interaction. Experiments with multiple participants across 300 trials show that combining pushes and voice commands identifies the intended goal in 86% of trials within at most five interactions, with most trials resolved in two. The inferred dynamical systems can also produce curved trajectories that avoid obstacles known only to the human.
Figures & tables
Fig. 1 : HINT-Blimp infers human intent from multimodal interactions. Human–robot interaction occurs exclusively through physical pushes and voice commands. In this scenario, the user intends to send the blimp to Goal 2 but cannot push it directly toward the goal because a straight-line trajectory would collide with the obstacle. Instead, the user first pushes the blimp toward Goal 1 and then redirects its motion using the verbal command “Go right!”. Our probabilistic method (right column) maintains beliefs over candidate trajectories to each goal and progressively converges to the correct intent as additional information is provided by the user.
Fig. 2 : Examples for vector fields and robot trajectories for LDS. From left to right, the LDS parameters are (ϑ1,ϑ2,r)=(0.58,103∘,133∘) , (0.90,117∘,70∘) , and (0.47,117∘,70∘) .
Fig. 3 : Setup for Experiment 1 where the user has to send the robot to one of the 8 candidate goals selected randomly at the beginning of the run.
Fig. 5 : Distribution of the number of interactions used in successful trials, pooled across all five participants. Each participant completed 20 trials per modality, giving 100 trials per modality across the five participants
Fig. 6 : Interactions to successful commitment under three prior conditions ( n=40 each). The adversarial prior, despite concentrating 40% of particles on the farthest goal, matches the uniform baseline, while the helpful prior shows a modest reduction.