Organizations: School of Electrical and Electronic Engineering, Nanyang Technological University. · School of Mechanical and Aerospace Engineering, Nanyang Technological University. · The Hong Kong University of Science and Technology (Guangzhou). · School of AI and Robotics, Hunan University.
Assistive robots increasingly operate in many human-centered environments and perform various human-robot interaction (HRI) tasks, such as object delivery. However, most existing HRI systems rely on RGB cameras that continuously observe humans to respond to non-verbal commands, such as hand gestures. This raises privacy concerns in privacy- critical environments, such as hospital wards or restaurants, where direct camera observation of humans is restricted. To develop privacy-preserving HRI, we leverage millimeter-wave (mmWave) radar, which can sense human motion through privacy barriers without identifiable imagery. We propose mmHRI, the first multi-modal robot manipulation framework that achieves mmWave radar-guided privacy-preserving HRI. mmHRI introduces two key designs to mitigate the sparsity and temporal inconsistency of radar data in cluttered robot manipulation environments. First, we propose a dual-stream architecture that jointly learns from unfiltered raw radar tensors and radar point clouds to estimate both human actions and 3D poses. To mitigate signal inconsistency, mmHRI further incorporates a memory-based state-space model (MSSM) that retains historical radar features to reduce abrupt changes in pose/action. These estimated human states are then converted into structured textual robot instructions, which control a vision-language-action (VLA) policy for closed-loop robot manipulation and human-aware reactions. Our evaluation covers human action recognition and closed-loop delivery and retrieval. In the privacy-preserving curtain setting, mmHRI achieves 85.09% action-recognition accuracy, outperforming existing radar-based alternatives. Robot trials further demonstrate successful delivery and retrieval under visual occlusion, with stable task performance across unseen subjects, clutter configurations, and environments.
Figures & tables
Fig. 1: Motivation of privacy-preserving mmHRI using mmWave radar for human sensing. Unlike conventional vision-based HRI that fails under occlusion, mmHRI senses human intent through privacy curtains to coordinate object delivery without directly observing humans. This supports privacy-sensitive applications such as hospitals, restaurants, and unmanned stores.
Fig. 2: Left: Physical robot system setup and the mmWave sensing unit. Middle: Two privacy-preserving scenes, i.e., hospital ward and office. Right: Three gesture/pose based human-robot-interaction decision-making tasks. We show different radar signals that the robot may refer to for different reactions.
Fig. 3: Overview of mmHRI. Radar measurements are first preprocessed into RDT tensors and RPC. The dual-stream radar human perception (DRP) then extracts motion and geometry patterns from both modalities, which are fused to jointly predict action and 3D poses. The human-aware text reasoning (HTR) then converts the estimated human states into structured robot instructions, controlling the downstream VLA policy for different actions.
Fig. 4: Visualization of the complementary radar modalities used by the dual-stream model. The DT and RT maps preserve motion-sensitive Doppler patterns, whereas the RPC retains 3D spatial structure for pose and root estimation. Synchronized RGB images are shown for visual reference.
Methods
HPE
Clear Visibility
Occlusion (Privacy)
Occlusion (Cross-Env)
HAR
HAR
HAR
MPJPE
(cm)
TE
(cm)
ACC
(%)
FPR
(%)
Static-F1
TABLE I: Performance on the radar perception dataset. HPE is evaluated under clear visibility, and HAR is evaluated under clear visibility, privacy occlusion, and cross-environment occlusion.
Methods
Side Decision & Grasping
Object Delivery
Box Retrieval
Collision Avoidance
PA
MA
Success
PA
MA
Success
PA
MA
Success
PA
MA
Success
Clear Visibility
RGB + pi05
30/30
26/30
26/30
30/30
27/30
27/30
27/30
30/30
27/30
30/30
30/30
30/30
mmWave+pi05
30/30
26/30
29/30
26/30
29/30
29/30
30/30
30/30
Occlusion
RGB+pi05
0/30
26/30
0/30
0/30
27/30
0/30
0/30
30/30
0/30
0/30
30/30
0/30
TABLE II: Performance on real-world robot trials. . PA and Success report the perception-only and end-to-end stage success rates. MA reports the success rate over 30 manipulation-only trials with ground-truth human states.
Fig. 5: Qualitative examples of the human states and corresponding robot operations. The columns show an azimuth wave, waiting, approaching, and a radial wave; the camera views show object grasping, tray delivery, and tray retrieval during the interaction.
Subject
Delivery Activation
Retrieve Activation
Collision Avoidance
Seen Subject
28/30
29/30
30/30
Unseen Subject 1
30/30
30/30
30/30
Unseen Subject 2
27/30
26/30
30/30
TABLE III: Generalization to unseen subjects.
Clutter
Delivery Activation
Retrieve Activation
Collision Avoidance
No Additional Clutter
28/30
29/30
30/30
Metal Box
28/30
26/30
30/30
Tall Cup Box
26/30
30/30
30/30
TABLE IV: Generalization to unseen tabletop clutter.
Fig. 6: Real-world robustness evaluation setups. (a) Cross-table-clutter settings with a metal box and an occluding paper box. (b) Cross-subject settings with two unseen subjects.
Institute of Medical Technology, Peking University Health Science Center · National Institute of Health Data Science at Peking University · State Key Laboratory of General Artificial Intelligence, Peking University