Human-Robot Interaction

Also known as HRI

Momentum

47 papers in the last four weeks, up 236% on the four weeks before. 0.5% of all new papers.

Jul 13Week of Sep 28

Latest papers 264

Oct 7, 2026cs.RO

COOL: Curiosity-Driven Object Ownership Learning for Personalized Robotic Assistance

Robots are increasingly expected to provide personalized services in everyday environments. To do so, they must ground natural-language commands such as "Where is my backpack?" or "Find my bottle" and execute them by reasoning about object instances, people, locations, and ownership. This is challenging because ownership is rarely labeled explicitly and must be inferred from long-term, behavioral evidence of human-object interactions. To address this, we present COOL, a novel robotic framework for autonomously learning object ownership from everyday observations and maintaining a long-term spatial memory of its environment. To keep its memory current, COOL uses an agent-based curiosity-driven data collection strategy that guides the robot toward the most promising locations to gain information and refresh stale observations. Offline experiments, ablation studies, and real-world evaluations show that COOL can infer ownership relations from real-world interactions and use this knowledge for ownership-conditioned navigation and task execution.
Oct 6, 2026cs.RO

Towards an Extensible Benchmark for Spoken Dialogue with Social Robots

Language models provide a plug-and-play interface between humans and robots, but important challenges remain when speech, dialogue, fast interaction, and collaboration are required. We propose a benchmark for the community to use as a way to explore common spoken dialogue artifacts between robots and humans, including requests for clarification, interruptions, embodied signals (e.g., head nods or facial cues), and time constraints. We also explain our vision to extend the benchmark for other aspects of human-robot interaction that are important to the larger research community. To facilitate the benchmark, we further propose using \textit{Retico}, a real-time communication framework that fulfills important technical requirements to enable robots to have spoken dialogue capabilities.
Oct 6, 2026cs.RO

Feeling Through the Load: Compliant Quadruped Locomotion under Payload Interactions

Quadruped robots are increasingly expected to carry objects while moving through human environments. But what happens when a person interacts directly with the payload rather than with the robot? If the payload is unrestrained, the robot must distinguish intentional external interactions from ordinary payload motion, while still keeping the load balanced and maintaining stable locomotion. How can a quadruped infer and compliantly respond to such interactions using only onboard measurements? In this work, we develop a force-aware locomotion framework that treats payload interactions as commands that shape the motion of the combined robot-payload system. Our approach separates the learning of force-aware locomotion and force estimation on an unrestrained payload. We combine a compliant load-carrying policy with a causal force estimator, trained through estimator-in-the-loop data aggregation and finetuning, to predict interactions from onboard robot measurements. Our simulations and real-world experiments show that the resulting controller can maintain stable payload-carrying locomotion, yield compliantly to external interactions, and use the inferred force to support human-guided changes in the robot's trajectory.
Oct 6, 2026cs.RO

Reactive Task-Oriented Robot-Human Handovers via Generative Hypothesis Selection

When humans hand each other objects, they incorporate both geometric and semantic information into this process. For example, passing a knife with the handle towards the recipient, rather than the blade, is both more ergonomic and safer. Recent state-of-the-art methods for task-oriented robot-human handovers have progressed from modeling object geometry to incorporating object affordances. However, they often forgo predicting the explicit, task-specific hand poses a human selects to utilize an object. Since many objects support multiple interaction modalities, e.g., a claw hammer used to strike or pull nails, this variability must be modeled to achieve robust task-oriented handovers. To tackle this, we propose a novel approach, GENESIS-Handover (GENErative HypotheSIS), which leverages VLM image generation to produce a variety of task-specific hand-object interaction hypotheses. These hypotheses are matched in real time to the observed human hand pose, enabling inference of the most suitable handover configuration. By leveraging VLMs as priors of plausible hand-object interactions, the method produces task-conditioned handover strategies for previously unseen object-task pairs. We evaluate the standalone interaction proposal module before deploying the full system on a mobile manipulator. In a user study with 12 participants across five task-object pairs, 83.3% perceived our method to have better task understanding than the previous state of the art.
Oct 5, 2026cs.HC

MRPilot: Supervising and Intervening LLM-Based Multi-Robot Teams through Mixed Reality

Large language models (LLMs) let users direct heterogeneous multi-robot systems (MRS) through natural language, but make task interpretation, robot assignment, and coordination difficult to inspect and change. Based on a formative study with 12 non-expert users, we developed MRPilot, a mixed reality system organized around four stages of supervision and intervention. MRPilot represents robot-team plans and execution states as structured commitments shared across synchronized situated and overview views. Across four stages, it helps users resolve ambiguous references (Forming), review plans before execution (Reviewing), monitor distributed execution (Following), and make robot-level or team-level changes when problems arise (Repairing). In a within-subjects study with 20 participants in a virtual reality-simulated home, MRPilot reduced workload, increased situational awareness, transparency, trust, and perceived control compared with a conventional LLM-based conversational interface using the same LLM planner and robot capabilities. We provide design implications for multi-scale intervention, adaptive supervision, and calibrated reliance in LLM-based MRS.
Oct 5, 2026cs.RO

Toward Trustworthy Physical AI for Human Interaction

Robots are entering human spaces faster than we can establish when they deserve trust. We propose a framework for trustworthy physical AI that integrates Safety, Behavioral Intelligibility, and Perceptual Alignment across embodiment, control, cognition, and design. Trustworthiness emerges from aligning physical capabilities, observable behavior, and expectations people form during interaction.
Oct 5, 2026cs.RO

Talk, Render, Act: Integrating Social Gesture and Digital Face with Synchronized Speech for Conversational Humanoid Robot

Expressive humanoid interaction requires speech, facial animation, and body gestures to form a coherent response. However, many full-body humanoid robots produce speech and gestures without a visually expressive face, while talking-face animation and robot gesture generation are typically developed separately. We present Talk, Render, Act (TRABot), an agent-based framework comprising specialized agents for motion-atom construction, dialogue generation, motion planning, and facial animation. First, to produce natural and semantically meaningful gestures, we construct Robot-Ready Semantic Motion Atoms by segmenting long-form, G1-retargeted BEAT2 motion into units with natural gesture boundaries, human-verified communicative functions, and feasible trajectories. Second, to preserve semantic order and coordinate body motion with the spoken response, we introduce a Semantic-Conditioned Compositional Planner. Given an ordered semantic function sequence and an estimated response duration, the planner selects approved atoms to realize the longest feasible action sequence while accounting for transitions and neutral recovery. Finally, we deploy a Streaming Face-Speech-Body Integration system on a physical G1 humanoid, combining streaming dialogue audio, audio-driven facial animation, and semantically planned body motion in a unified real-time interaction loop. Quantitative and qualitative experiments demonstrate that TRAbot achieves the best overall performance among all compared conditions in terms of naturalness, expressiveness, and multimodal coherence.
Oct 5, 2026cs.RO

From Social Reasoning to Embodied Interaction: An Agentic Framework for Social Robots

Natural face-to-face human--robot interaction requires a robot to understand an evolving social situation, decide when to engage, and express its intent through coordinated physical behavior. Yet existing approaches rarely close this loop: foundation-model agents provide increasingly capable multimodal reasoning and memory but remain largely disembodied, while expressive virtual agents do not face the physical constraints of real robots, and physical social robots typically address social reasoning and embodied expression only partially. We present ARISE, a unified framework that bridges Agentic Reasoning and Interactive Social Embodiment on the Sophia humanoid robot. ARISE integrates multimodal context understanding, long-term memory, and reactive and proactive interaction to determine when and what to communicate, and translates social intent into robot-native gestures coordinated with speech and mechanical facial expressions through streaming execution. Extensive evaluations on Sophia demonstrate strong perceived interaction quality, expressive and well-coordinated embodied behavior, and substantial latency reductions through streaming execution. These results highlight the importance of jointly reasoning about what to communicate, when to engage, and how to physically express social intent for natural interaction with humanoid robots. Project Page: https://robosocial.github.io/
Oct 4, 2026cs.RO

Social Navigation for Tour-guide Robot

We propose a force-based model for social navigation of a tour-guide robot. Social forces due to various factors like obstacles, user position and heading, have been accounted for in the model. We claim that each one of these forces makes the robot more sociable to the user and we design an experimental setup for evaluation. In the experiment, the user follows an autonomous robot to a destination in a known map, while undertaking a few simple sub-tasks in the middle, which serve as distractions. For each participant, we run several rounds of the experiment, each with different forces and a shortest path, A*-search baseline model. Using per-round subjective indicators, we propose to study the effect of our force model on constructs such as: follow-ability, perceived safety, and perceived intelligence.
Oct 4, 2026cs.RO

CoDance: Learning Reactive and Compliant Human-Humanoid Interaction from Video

Partnered human-humanoid interaction couples locomotion with continuous physical contact. A humanoid needs to coordinate with a person's motion while responding to interaction forces and maintaining stable and natural movement. We present CoDance, a framework for learning reactive and compliant human-humanoid interaction from video. We study partnered dancing as a challenging instantiation, where a humanoid coordinates its footsteps with a moving partner and maintains continuous two-hand contact. Given a single video of two human dancers, CoDance retargets their motions into a robot reference and a moving partner. We introduce a multi-link compliance augmentation that transforms the kinematic demonstration into force-aware training data by adapting the robot reference under structured forces at both hands. Policies trained on this data follow the observed partner while preserving the demonstrated locomotion style and responding compliantly to physical interaction. In simulation, the policies adapt their footsteps to changes in the partner and reproduce approximately 80% of the wrist displacement encoded by the augmented demonstrations. On a physical humanoid, CoDance enables sustained two-hand dancing with a human partner including repeated transitions between forward and backward motions.
Oct 1, 2026cs.RO

Towards Physical Underwater Robotic Assistance for Scuba Diver Movement in Confined Spaces

Scuba divers are taught to control their depth to avoid rapid ascents and descents, which could result in serious injuries such as gas embolisms and barotrauma. However, many underwater tasks necessitate lateral control, maintaining distance between subsea structures such as coral reefs, submerged drilling instrumentation, or unexploded ordnance. In this work, we discuss a first-of-its-kind wearable robotic solution providing thruster-actuated directional guidance to a diver, as distinct from prior propulsive-assistance exoskeletons. We introduce ``Robotic Assisted Diver Movement in Confined Spaces'' (RADMCS), a wearable robot that assists divers in maintaining a fixed distance from subsea structures by leveraging perception techniques in monocular depth estimation and force-feedback from submersible thrusters to provide haptic feedback. Its small and compact form factor creates a foundational platform that could be expanded to include more sophisticated control and navigation behaviors. We present results from Institutional Review Board (IRB) in-water studies with eight human scuba diver participants on threshold sensitivity tests in both a closed-water swimming facility and ocean environments; distance-maintaining experiments in a closed-water facility; and form, fit, and function testing in the ocean. We demonstrate that relatively low thrust values (10 percent of maximum) allow robotic direction of a human's movement using the physical sensation of the robot's guidance.
Oct 1, 2026cs.RO

Real-Time Human-Adaptive Task Allocation for Multi-Human Multi-Robot Supervision

We propose a human-factor-aware method of allocating robot supervision tasks to multiple human operators. In scenarios where multiple operators occasionally teleoperate multiple robots to help the robots overcome difficulties, the allocation of the supervisory control tasks to humans needs to consider the real-time cognitive states of individual operators. However, most existing methods assume fixed supervisory capacity per operator and overlook fluctuations in the human factors such as workload and fatigue. As a result, workload distribution can be unbalanced where some operators become overloaded while the others remain underused. Our method dynamically regulates supervisory capacity and allocates tasks in a way that maintains balanced mental workload, prevents overload, and improves overall team performance. The allocation method uses a greedy strategy that minimizes estimated operator workloads with task prioritization. Robots are assigned to operators by reflecting their current supervisory capacity where the required effort depends on the types of tasks. In the user study, the analysis across predefined time intervals shows that the proposed method consistently achieves higher performance and lower behavioral signs of fatigue compared to a baseline method that does not consider human factors. These results highlight adaptive capacity adjustment as an effective preventive mechanism for sustaining operator performance in long-duration, high-demand settings.
Sep 30, 2026cs.RO

Rethinking Legibility in Social Robot Hallway Navigation: Impact of Intent Representation and Human Distraction

We focus on legible robot motion generation in social navigation settings. Legibility in human-robot interaction (HRI) is often described as the property of robot motion that enables an observer to confidently infer the robot's intent. While mature frameworks exist for generating legible motion in front of static observers, social robot navigation presents a new challenge: the robot must clearly convey its intent while ensuring human safety in dynamic pedestrian environments where human attention is often divided. With the goal of enabling robots to generate legible motion in dynamic and constrained spaces, we investigate how the choice of representation and the level of human attention shape navigation performance and human impressions. Focusing on the ubiquitous and demanding scenario of hallway navigation, we conduct two controlled user studies involving alternative legibility formulations implemented within a shared model predictive control framework. Study 1 (N = 45) investigates the role of intent representation, showing that passing-side legibility, particularly when adaptively updated, leads to smoother human motion and is perceived as more competent and less mentally and physically demanding than destination-based and non-legible baselines. Study 2 (N = 45) examines the effect of pedestrian attention, demonstrating that legible motion allows for smooth human motion even under distraction, even if this is not consistently reflected in subjective ratings. Together, these findings suggest that effective legible motion in social robot navigation benefits from interaction-level intent representations that support coordination, with some effects persisting even when human attention is divided. Code is available at https://github.com/fluentrobotics/Legible_MPPI.
Sep 30, 2026cs.RO

DiFF: Doppler-informed Flow Matching for Human Motion Flow

Perceiving human motion via privacy-preserving 4D millimeter-wave (mmWave) radar is critical for next-generation human-robot interaction (HRI), where point cloud scene flow serves as a foundational motion representation. Yet the extreme sparsity and noise of 4D radar point clouds make non-rigid motion flow estimation severely ill-posed--a challenge that existing rigid-centric methods and prior works fail to adequately address, largely because they neglect the rich Doppler velocity cues inherent in 4D radar. We propose DiFF, a generative framework that marries Doppler-informed motion priors with a Kolmogorov-Arnold Network (KAN)-based conditional flow matching model. At its core, a KAN-attention mechanism enables expressive feature extraction, while a prior-guided generative process harnesses Doppler cues to regularize the ill-posed solution space. Extensive experiments show that DiFF achieves state-of-the-art (SOTA) performance across diverse real-world datasets, reducing 3D endpoint error to the millimeter scale on the mmBody benchmark.
Sep 30, 2026cs.RO

AIfred: Augmented Learning through Functional Robotic Embodiment at the Desk

Desk-based learning and creative activities benefit from handwritten engagement. However, current generative AI tools deliver guidance through a separate screen, creating a gap between where users think and where assistance appears. To address this, in this work we design AIfred, a desk-based robotic arm with a projector mounted at the end-effector that places AI-generated guidance alongside handwritten work. AIfred combines workspace perception, context-aware content generation, and robot-mediated projection to support math assignments, image generation, and drawing tasks. In a user study (n = 36), we compared AIfred against ChatGPT (GPT-5.6 Luna) running on a laptop. Both tools performed comparably while assistance was available during the math assignment (6.7 vs. 7.3/10, p = .41), but AIfred improved short-term learning transfer by 60% once assistance was withdrawn (7.0 vs. 4.4/10, p = .003). In addition, independent art and design professors ranked drawings produced with AIfred better in 33 of 36 cases. Our findings indicate that spatially co-located AI assistance benefits tasks whose guidance shares a spatial frame with the work.
Sep 28, 2026cs.RO

CoBrush: A Hierarchical Planning Framework for Human-Robot Co-Painting

Embodied co-painting requires a robot to repeatedly update a shared physical canvas while human intent evolves over interaction. Existing reference-driven painters or reactive assistants are typically optimized for single-shot rendering or sketch completion, limiting their ability to sustain coherent multi-round collaboration or to construct complex, content-rich scenes over time. We present CoBrush, a hierarchical framework that formulates multi-round co-painting as a coordinated semantic, spatial, and execution process. By separating high-level intent inference from spatial grounding and stroke-level control, the system supports progressive scene development on real acrylic canvases. We evaluate the framework through real human-robot painting sessions, stress tests, and user studies. Compared to single-turn baselines, our approach achieves stronger semantic alignment, more stable spatial progression, and higher perceived plausibility of robot actions. These results demonstrate that structured multi-stage reasoning improves the coherence and robustness of interactive painting and supports the progressive development of content-rich physical artworks.
Sep 28, 2026cs.RO

Bayesian Active Learning for Intent Disambiguation in Interactive Robot Planning

Interactive robot planning requires robots to infer and execute human intentions from natural language instructions that are often ambiguous, incomplete, or underspecified. Although large language models (LLMs) provide a powerful interface for clarification, relying on the generative model to drive an multi-turn conversation can introduce systematic failures. We propose a Bayesian framework that treats clarification as an active learning problem over grounded Signal Temporal Logic (STL) task specifications. Our method uses LLMs to initialize candidate formal specifications and translate informative contrasts into natural-language clarification questions, while Bayesian optimization maintains uncertainty estimation over user intent and selects queries that maximize information gain. After convergence, the inferred STL specification is passed to a formal planner to synthesize a verifiable robot trajectory. Across four simulated and real-world task domains, our approach generally achieves higher task satisfaction and requires fewer clarification rounds than LLM baselines, while helping smaller models close the performance gap against larger reasoning models.
Sep 28, 2026cs.RO

Proprioceptive Force Estimation for Quadruped Locomotion and Human-Robot Interaction

Payload forces must be accommodated during locomotion, while leash forces can specify desired motion. We investigate whether a shared three-dimensional force estimate in newtons, inferred from proprioceptive history under sustained loading, can support both tasks. An estimator and locomotion policy are jointly trained with supervised force and velocity outputs and learned latent context. The estimated force conditions locomotion and additionally generates planar-velocity and yaw-rate commands for leash guidance through an analytical map. In sustained-force simulation sweeps, temporal means of componentwise force root mean square error range from 1.44 to 2.83,N. Compared with a domain-randomized baseline, the framework reduces velocity-tracking and base-orientation error scores by 21.6% and 46.5%, respectively, and increases mean survival from 68.29% to 94.60% in separate sustained-force tests. Unitree Go1 experiments demonstrate stationary vertical and horizontal force estimation, locomotion with an 8.5,kg payload whose weight exceeds the 70,N training force limit, and leash guidance using the same force-estimation interface.
Sep 28, 2026cs.RO

mmHRI: Towards Privacy-Preserving Human-Robot Interaction with Millimeter-Wave Radar

Assistive robots increasingly operate in many human-centered environments and perform various human-robot interaction (HRI) tasks, such as object delivery. However, most existing HRI systems rely on RGB cameras that continuously observe humans to respond to non-verbal commands, such as hand gestures. This raises privacy concerns in privacy- critical environments, such as hospital wards or restaurants, where direct camera observation of humans is restricted. To develop privacy-preserving HRI, we leverage millimeter-wave (mmWave) radar, which can sense human motion through privacy barriers without identifiable imagery. We propose mmHRI, the first multi-modal robot manipulation framework that achieves mmWave radar-guided privacy-preserving HRI. mmHRI introduces two key designs to mitigate the sparsity and temporal inconsistency of radar data in cluttered robot manipulation environments. First, we propose a dual-stream architecture that jointly learns from unfiltered raw radar tensors and radar point clouds to estimate both human actions and 3D poses. To mitigate signal inconsistency, mmHRI further incorporates a memory-based state-space model (MSSM) that retains historical radar features to reduce abrupt changes in pose/action. These estimated human states are then converted into structured textual robot instructions, which control a vision-language-action (VLA) policy for closed-loop robot manipulation and human-aware reactions. Our evaluation covers human action recognition and closed-loop delivery and retrieval. In the privacy-preserving curtain setting, mmHRI achieves 85.09% action-recognition accuracy, outperforming existing radar-based alternatives. Robot trials further demonstrate successful delivery and retrieval under visual occlusion, with stable task performance across unseen subjects, clutter configurations, and environments.
Sep 27, 2026cs.RO

AI-Driven Collaborative Assembly Line Inspection: System Integration and Deployment Challenges

Manual visual inspection on assembly lines is a persistent manufacturing bottleneck: operator fatigue over extended shifts lowers defect-detection rates. This paper presents the design, integration, and field deployment of an AI-assisted collaborative inspection cell at the Silverline kitchen-appliance factory, developed within the AI-PRISM project. The cell couples a Universal Robots UR 10e cobot carrying a machine-vision defect-detection pipeline with a Comau Racer-5 cobot for functional tests, coordinated through ROS 2 Humble on an Ubuntu 22.04 LTS server. Multi-modal data (Basler camera imagery, TIA microphone acoustics, and SPS electrical-safety measurements) are logged locally and visualised in real time with Grafana. We report the practical deployment challenges (close-proximity safety, AI robustness under glare and reflections, ROS 2 namespace collisions across two cobots, and operating-system and dependency issues) together with the engineering solutions adopted, and structure the integration through a four-level Human-Robot Interaction analysis. The deployed cell cuts per-unit quality-check time from 82 s to 61 s (about 25%), raises final-control resource efficiency from 0.75 to 0.88, reduces operator visual-inspection viewing time by 82%, and significantly lowers operator mental demand (p = 0.005, NASA-TLX).
Sep 26, 2026cs.RO

WSM-Aware HRI: An IoT-Enhanced Framework for Early Detection and Norm-Guided Repair of Failures with LLM Guidance

Human-robot interaction (HRI) failures remain a major barrier to deploying robots in real-world environments. Prior work often treats failures as isolated technical faults or focuses on post-hoc recovery behaviors. In practice, many breakdowns arise because humans and robots operate under inconsistent assumptions about the current world state. We propose WSM-Aware HRI, an IoT-enhanced modular framework that unifies diverse HRI breakdowns as World-State Mismatches (WSMs) between a human's instruction-implied assumptions and a robot's grounded world model built from multimodal perception and digital augmentation. A Large Language Model (LLM) is used to make implicit assumptions explicit, map them to a small set of mismatch types, and specify the evidence needed for verification against the robot's world state. WSM-Aware HRI shifts failure handling from execution-time recovery to proactive mismatch detection during intention formation, enabling interventions guided by safety, norm compliance, and multi-user coordination with transparent explanations. We evaluate mismatch identification in ten everyday cases spanning both visual and latent-state mismatches. The system can accurately produce the expected output results, and ablations show that reliable identification depends on appropriate grounding representations and verification-oriented refinement. These results indicate that treating interaction breakdowns as explicit world-state mismatches enables earlier detection of impending failures and offers a principled mechanism for integrating external evidence and social constraints into human-robot interaction.
Sep 24, 2026cs.HC

A Procedure for Classifying Attachments and Affective Social Bonds in Human-Robot Dyads

Human-robot interaction (HRI) claims that people form attachments and social bonds with artificial agents, yet the terms are often applied without the behavioural and physiological criteria that give them content in their source disciplines. Without this empirical grounding, studies deploy widely divergent methods, frequently producing expansive relational claims that far outstrip their underlying evidence. To address this, we propose a standardised four-question procedure, grounded in criteria established in the developmental, ethological, and neuroendocrine literatures, that classifies a given human-robot tie as an attachment, an affective social bond, or no relationship, with intermediate classifications when evidence is incomplete. We specify minimum evidential requirements for each question, and provide candidate HRI study designs, adapted from validated human-human, human-animal, and animal-animal paradigms. We then demonstrate the procedure by applying it to a representative set of published HRI studies, showing how often relational claims outstrip what the reported designs can establish. Finally, we discuss the ethical and regulatory burdens created when artificial agents engage human biobehavioural systems. By replacing the divergent operationalisations with a unified, criterion-based classification, this paper gives HRI practitioners a standardised basis for evaluating, classifying, and comparing human-robot relationships, and sets out the experimental rigour that each classification demands. We therefore call on researchers of human-robot relationships to adopt such rigour, or to consider alternative terminology in their descriptions of these ties.
Sep 24, 2026cs.RO

Robots That Take Initiative: A Framework for Building and Evaluating Proactive Robots

Effective robot assistance beyond narrow roles and repetitive tasks requires robots to be proactive - to decide what needs to be done rather than waiting to be told. While proactivity is increasingly explored, it lacks a unified formulation, and work in the domain is typically evaluated offline against static human models that cannot capture the effect of a robot's actions on the environment and the user's own behavior. We introduce a unified formalism for proactive robot assistance, organize it into three levels, and provide a framework to address the highest level of unprompted proactive assistance. We then show that offline evaluation overstates performance in this setting, and contribute a closed-loop evaluation with a human model that adapts to the robot. Finally, we present a method, GAP, that instantiates our framework, learning from passive observation to anticipate user goals and act. Under closed-loop evaluation, prior state-of-the-art methods collapse, in some cases adding more work than they save, while GAP remains robust and substantially outperforms them.
Sep 22, 2026cs.RO

HINT-Blimp: Human INTent Inference from Multimodal Cues for Robotic Blimps

In human-robot interaction, traditional interfaces such as joysticks and handheld tablets introduce latency into navigation tasks and require the operator's explicit attention on the device, instead of the robot. We propose a new human-robot interaction framework in which a human communicates intent directly through sparse multimodal signals such as physical pushes and spoken commands. Human intent is represented as a parameterized linear dynamical system (LDS) that encodes the desired goal and motion behavior. The robot estimates this intent (parameters) online using a particle filter, where each particle represents a candidate LDS hypothesis and is reweighted online as new information becomes available. We validate this framework on a robotic blimp, whose inherent compliance and collision tolerance make it well-suited for repeated physical interaction. Experiments with multiple participants across 300 trials show that combining pushes and voice commands identifies the intended goal in 86% of trials within at most five interactions, with most trials resolved in two. The inferred dynamical systems can also produce curved trajectories that avoid obstacles known only to the human.
Sep 22, 2026cs.RO

Benchmarking Robots for Everyday Environments: From Lab Experiments to Real-World Operations

This study introduces an interdisciplinary framework for benchmarking robots deployed in public environments, addressing the gap between traditional laboratory metrics and real-world benchmarking requirements. We evaluate three distinct robots across diverse use cases - outdoor park cleaning, pedestrian underpass cleaning, and interactive library assistance - each representing unique challenges in public daily life. Over a three-year benchmarking process (2023-2025) comprising seven benchmarking events, a consensus workshop and six on-site evaluations (two per use case), we utilized realistic indoor and outdoor test environments to assess not only technical performance but also the broader implications of deploying robots in unstructured, human-centric settings. An expert panel, spanning robotics, human-robot interaction, safety, and economics, systematically developed and refined an evaluation concept to analyze the transition from laboratory prototypes to operational systems. Our findings highlight critical factors for successful deployment, including task fulfillment, interaction quality, safety, and economic feasibility. This work provides actionable insights for researchers and practitioners aiming to bridge the gap between robotic innovation and real-world applicability.
Sep 22, 2026cs.RO

Towards Intent-Aware Human-Robot Teaming: A Platform for Search-and-Rescue Operations

We investigate the challenges of enabling effective collaboration between human operators and heterogeneous autonomous agents in complex, dynamic environments by developing an interaction platform that allows study of operator behavior and supports intent inference and decision-making using state-of-the-art frameworks. We demonstrate the extent to which the operator's perception, decisions, and actions could be supported by autonomous systems during search-and-rescue operations with our platform.
Sep 21, 2026cs.RO

Learning to Plan in Human-Robot Collaboration: Multimodal Reinforcement Learning for Adaptive Interaction

Robot assistants for older adults and people with disabilities need to perform collaborative tasks with users effectively. The core component of these systems is an interaction manager whose job is to observe and assess the task and infer the state of the human and their intent for the robot to choose the best course of action. Due to the sparseness of the data in this domain, the policy for such multimodal systems is often crafted by hand; as the complexity of interactions grows, this process is not scalable. This paper proposes a reinforcement learning (RL) approach to automatically generate the multimodal policy of the robot. Our system focuses on a realistic scenario where a robot assists a user in locating objects within a home environment, managing multimodal signals, including language and physical actions, to select the best action. In contrast to traditional dialog systems, our agent is trained with a simulator that uses human data and can deal with multiple modalities. We use a simple high-level reward function that needs no fine-tuning and enforce some preconditions to speed up the training process. A human study evaluating the system in a real-world setting demonstrates promising results, indicating high usability and effective task completion. This RL-based approach offers a scalable and interpretable alternative for designing interaction managers in multimodal human-robot collaborations.
Sep 21, 2026cs.RO

LLM-based Conversational AI Knowledge Assistant for MyBuddy Humanoid Robot

Humanoid robots are increasingly being popular and developed for human-centered applications, yet their ability to provide intelligent conversations and natural interactive knowledge assistance remains constrained by traditional rule-based dialogue systems, pre-defined responses and limited knowledge repositories. Large language models (LLMs) have emerged as a powerful foundation for enabling natural, adaptive, and context-aware Human-Robot Interaction (HRI), which provides a significant opportunity to address such limitations by enabling robots to understand natural speech language, reason over complicated queries, maintain high-quality conversational context, and generate knowledge-rich responses. In this work, we originally present and implement an LLM-based versatile Conversational AI Knowledge Assistant for the Raspberry-Pi-powered 13-Axis MyBuddy humanoid robot, which integrates LLM-driven language understanding and AI reasoning with real-time speech recognition, knowledge retrieval via extensible access of internet engines (e.g., Wikipedia, arXiv), flexible dialogue management, and natural speech synthesis to enable much more intelligent multi-turn continuous conversations and advanced emotional-support Human-Robot Interaction.
Sep 21, 2026cs.RO

MIRA: Real-Time Full-Duplex Human-Robot Interaction for Embodied Companions

% !TEX root = ../main.tex Real-time embodied companion interaction requires a robot to infer user intent from streaming speech, generate timely responses, and execute expressive, interruptible motions. Existing systems typically decouple dialogue orchestration from gesture synthesis, relying on offline motion generation from complete audio. This separation leaves open how a deployed robot can dynamically synchronize response content, prosodic timing, and physical safety under incremental inputs and uncertain turn boundaries. We present MIRA, a unified framework for real-time full-duplex embodied companion interaction. Given streaming user speech, dialogue history, and vocal affect, MIRA predicts both the response text and an explicit embodiment cue that routes the response to the appropriate physical behavior. Discrete social behaviors (\eg listening and greeting) are mapped to validated robot trajectories, while speaking responses are accompanied by streaming, generative co-speech motion. For co-speech motion generation, we propose ROSCO, a prefix-conditioned diffusion model for streaming audio-to-joint motion generation. We further design RHPC, an inference scheme that maintains a sufficiently long temporal context for motion prediction while bounding physical commitment to a short, interruptible prefix. At the interaction level, we design CORTEX, a dual-timescale interaction policy that combines low-latency barge-in preemption and streaming response generation with deliberative turn decisions, backed by a robot-side execution layer that enforces physical safety constraints during execution. MIRA is deployed on an Astribot S1 humanoid robot. Quantitative evaluations demonstrate competitive audio-motion alignment relative to state-of-the-art motion-generation baselines, while real-robot deployment measurements characterize streaming responsiveness and interruption handling.
Sep 21, 2026cs.LG

NAVIR: Neuromorphic Audio-Visual Speech Recognition for Robust Human-Robot Interaction on Edge Hardware

Voice-controlled interaction in industrial settings is hampered by acoustic noise, which severely degrades audio-only speech recognition. Audio-visual speech recognition (AVSR) addresses this by fusing lip-motion cues with the audio stream, but state-of-the-art pipelines rely on three-dimensional convolutions, recurrent units, and attention modules that exceed the budget of typical edge devices. We present NAVIR, an end-to-end AVSR system targeting the BrainChip Akida neuromorphic processor, which natively supports only sequential two-dimensional convolutional inference. The pipeline factorises spatial and temporal encoding into separate AkidaNet-based modules: a per-frame visual encoder, a temporal video encoder, and a spectrogram audio encoder, fused by a lightweight predictor head and decoded by constrained beam search. Models are trained with connectionist temporal classification on noise-augmented audio and then fine-tuned with quantization-aware training. On the GRID benchmark, the quantized audio-visual model reaches 14.0% word error rate (WER) under noise on the unseen-speaker split and 3.3% WER on the overlapped-speaker split, against 22.5% and 11.8% for audio-only baselines, and it attains 98.6% command accuracy at 1.5% WER on a task-specific industrial-command corpus. Operation-count analysis indicates a 13-fold energy advantage of the spiking formulation over its artificial neural network counterpart at 27.6% mean firing rate. On-board measurements show roughly 5-fold lower energy per inference than a Raspberry Pi central processing unit on the lip-reading model, and over 100-fold lower than a laptop graphics processing unit, while sustaining 14.5 inferences per second. To the best of our knowledge, this is the first complete multimodal AVSR pipeline running on neuromorphic hardware of this class.