Safe and efficient robotic navigation among humans is essential for integrating robots into everyday environments. Most existing approaches focus on simplified 2D crowd navigation and fail to account for the full complexity of human body dynamics beyond root motion. We present HumanHalo, an MPC framework for 3D MAV navigation among humans that combines theoretical safety guarantees with data-driven models for realistic human motion forecasting. Our approach introduces a novel reachability-based safety formulation that constrains only the initial control input for safety while modeling its effects over the entire planning horizon, enabling safe yet efficient navigation. We validate HumanHalo in both simulated experiments using real human trajectories and in the real world, demonstrating its effectiveness across tasks ranging from goal-directed navigation to visual servoing for human tracking. While we apply our method to MAV in this work, it is generic and can be adapted to other platforms. Our results show that the method preserves safety without excessive conservatism, even under imperfect human prediction and in constrained workspaces, while remaining efficient enough for real-time onboard execution.
Figures & tables
Fig. 1 : HumanHalo: A control framework for safe and efficient MAV navigation among humans. Our method combines reachability-based safety constraints with state-of-the-art data-driven human motion forecasting, ensuring recursive safety guarantees without overly conservative behavior.
Fig. 2 : Based on the current MAV state x0 and the tracked 3D positions of human body joints, we first compute both the MAV ’s and the human’s reachable sets over the entire horizon. We then optimize the control inputs uk with the requirement that the initial control input u0 must not lead to a situation where the MAV ’s reachable set becomes a subset of the human’s reachable set at any time. While partial overlap between the two sets is safe, as it can be resolved by future control actions, full containment of the MAV ’s set within the human’s would imply inevitable collision regardless of future control inputs. To run on an MAV , our method is augmented with a real-time perception and forecasting stack (dashed boxes) that provides the current human and MAV state and informs the objective about future human motion.
Fig. 3 : Example evaluation of our safety constraint for three different initial control input choices. While the left choice leads to unsafe behavior, both the middle and the right choices for the initial control input are regarded safe as they leave options to avoid collision later on.
Fig. 4 : Example constructions of the human reachable set for discretization steps Δt=0.025s and Δt=0.500s . The black dots represent the 3D body joints at t=0s .
Parameter
Value
Parameter
Value
Parameter
Value
ρHead
0.2
ρTorso
0.3
ρArm
0.205
ρHand
0.1
vi,max
1.0
ai,max
1.0
λ
1000
RMAV
0.5
τmin
-0.5
τmax
0.25
θmin
$-$$$
θmax
15 °
ϕmin
$-$$$
ϕmax
15 °
b1
1.0
b2
0.1
b3
1.0
b4
0.1
TABLE I : Parameters for reachable sets, box constraints, and MAV model in simulation.
Method
1 Human
2 Humans
Coll. Avoid. ↑ [%]
SR ↑ [%]
TTG ↓ [s]
Solver ↓ [ms]
Coll. Avoid. ↑ [%]
SR ↑ [%]
TTG ↓ [s]
Solver ↓ [ms]
No Constraint
94/92
94/92
3.96/3.96
0.05
89/91
89/91
3.56/3.56
0.07
3D-ORCA [ 18 ]
98/96
84/80
4.98/4.87
0.05
100/93
91/87
4.74/4.74
0.07
CBF-QP [ 15 ]
100/100
100/98
4.43/4.51
0.06
100/98
98/91
4.11/3.99
0.11
MPC-CBF [ 16 ]
100/96
100/96
4.17/4.14
0.57
100/91
100/91
3.94/3.82
1.39
Salazar et al. [ 9 ]
96/90
82/78
4.91/5.03
0.06
98/98
91/89
4.72/4.61
0.08
TABLE II : Simulation evaluation. Each cell reports det/noisy , i.e. under a perfect forecast and under an imperfect one where the executed human motion deviates from the prediction by a seeded per-sequence velocity drift (std 0.6ms−1 ). SR is success rate and TTG time-to-goal. With a perfect forecast all safety-aware methods are comparable, but under prediction error only the reachability variants stay fully collision-free while also reaching the goal faster than the external safe baselines, carrying a formal guarantee (Proposition 1) and, unlike prior baselines, extending to static obstacles ( Table III ).
Method
Human Avoid. ↑
Obstacle Avoid. ↑
Success Rate ↑
Min. Dist. ↑
Solver Time ↓
[%]
[%]
[%]
[m]
[ms]
No Constraint
14
100
14
0.21
0.18
3D-ORCA [ 18 ]
72
100
26
0.75
0.06
CBF-QP [ 15 ]
94
94
50
1.09
0.08
MPC-CBF [ 16 ]
62
96
56
0.77
0.76
Complex HRS
84
100
64
0.97
3.40
TABLE III : Constrained-workspace evaluation: a 2m -wide corridor with a low ceiling below the human’s vertical reach, so a flyover is impossible. Only our method stays simultaneously safe w.r.t. humans and obstacles (both 100% ) while completing the task most often.
Horizon
Coll. Avoid. ↑ [%]
Success Rate ↑ [%]
Time To Goal ↓ [s]
Solver Time ↓ [ms]
20
100
84
7.53
1.09
40
100
98
4.17
3.90
60
100
98
3.52
9.05
TABLE IV : Effect of horizon length in the single-human scene. Safety holds at every horizon. A short horizon yields slow, conservative trajectories, while extending the horizon reduces time-to-goal only up to a point, after which it mainly adds computation. A horizon of 40 balances goal-reaching speed and solver time.
Robust and accurate perception of humans in their 3D scene context is essential for integrating robots into everyday environments. Existing approaches, however, often fail to predict plausible and accurate human motion estimates that are consistent with the surrounding scene, especially in the presence of heavy occlusions or partial visibility. This can limit both safety and efficiency for robotic operations. We introduce HumanFlow, a latent diffusion model that unifies human motion tracking and forecasting, conditioned on the 3D scene context. We show that our human motion model produces smooth and accurate predictions under challenging conditions, including heavy occlusions, and outperforms state-of-the-art methods in tracking accuracy while being significantly more efficient. Furthermore, we show how HumanFlow's latent space can be tightly coupled with control by conditioning a flow-matching-based, approximate MPC policy on these representations. We validate our policy in simulation with real human trajectories for MAV social navigation, demonstrating superior navigation performance and remaining collision-free, even under partial observability of the human.
Safe and socially compliant navigation remains a fundamental challenge for autonomous robots operating in human-populated environments. Beyond collision avoidance, robots must anticipate human motion and respect personal space to ensure human comfort. Model Predictive Control (MPC) offers a robust alternative to classical and data-driven methods, although its effectiveness strongly depends on accurate human motion prediction and efficient computation. This paper introduces SFM-NMPC, a Social Force Model-based Non-linear Model Predictive Control framework that embeds human motion prediction directly within the optimization loop. By incorporating the Social Force Model into the dynamic model of surrounding agents, the controller jointly predicts the trajectories of humans and robots over the prediction horizon, thereby enabling socially-aware planning. A tailored set of social cost functions guides the optimization toward human-compliant behaviors. Despite the increased model complexity, the proposed formulation runs in real time at 20 Hz. Extensive simulated testing in crowded environments demonstrates that SFM-NMPC outperforms state-of-the-art baselines in social compliance metrics while maintaining efficient and smooth navigation. Visual trajectory analysis and an ablation study further highlight the contribution of the embedded SFM dynamics and social cost terms, confirming the effectiveness of the proposed approach for real-world social navigation.
Stefano Trepella, Andrea Ostuni, Mauro Martini +5
Department of Electronics and Telecommunications, Politecnico di Torino, 10129, Torino, Italy. · School of Engineering, Pablo de Olavide University, Crta. Utr-era km 1, Seville, Spain
Safe and efficient robot navigation in crowds requires anticipating pedestrian motion despite uncertain and potentially shifting prediction errors. Existing reactive methods can produce oscillatory behavior, while predictive planners often treat forecasts as exact or rely on restrictive error models. Incorporating conservative uncertainty sets as hard constraints can also render model predictive control (MPC) infeasible. We propose \textit{CoCoNav}, a crowd-navigation framework that combines online conformal calibration with runtime-certified planning. A horizon-specific conformal proportional--integral controller adapts trajectory-error bounds to regulate long-run empirical coverage, enabling the framework to respond to changing prediction errors. A \textit{relax-then-verify} planner preserves solver feasibility by generating nominal trajectories with soft-constrained MPC and separately certifying them, together with contingency maneuvers, against the calibrated bounds before execution. Simulations and quadruped experiments show that CoCoNav achieves a favorable balance among collision avoidance, task success, and navigation efficiency relative to the evaluated baselines.
Cheng Guo, Mingzhe Ni, Zheng Liang +5
Department of Computer Science, The University of Manchester, Manchester, UK. · Human-Robot Interfaces and Interaction Laboratory, Italian Institute of Technology, Genoa, Italy. · Genisom AI, Shanghai, China. +3