WAND: Learning Robust Navigation under Complex Wind Disturbances and Dense Obstacles for Quadrotors
Authors: Zhonghan Tang, Chenhui Li, Shuai Liang, Zhongrui You, Jianan Li, Bin Zhao, Zhigang Wang, Xuelong Li
Organizations: Department of Electronic Engineering and Information Science, University of Science and Technology of China, Hefei 230027, China · Shanghai Artificial Intelligence Laboratory, Shanghai 200232, China · Northwestern Polytechnical University, Xi’an 710072, China · Institute of Artificial Intelligence (TeleAI), China Telecom, Beijing 100033, China
Robust navigation in cluttered environments remains a fundamental challenge for quadrotors, particularly when strong wind disturbances arise, which perturb vehicle dynamics, limit control authority, and substantially increase collision risk. Existing learning-based navigation policies typically rely on obstacle perception and proprioceptive observations, requiring the policy to infer time-varying disturbance effects implicitly and thereby limiting robustness under partial observability. This paper proposes WAND (Wind-Aware Navigation with Disturbance Estimation), a reinforcement learning framework for navigation under time-varying wind disturbances in dense obstacle fields. Specifically, WAND estimates wind-induced disturbance acceleration from historical proprioceptive states using a Temporal Convolutional Network (TCN). This estimation is integrated into the policy via a zero-initialized residual module, \emph{WindAdapter}, while simultaneously providing feedforward compensation for low-level control. The dual use of the estimate couples disturbance-conditioned navigation with feedforward disturbance rejection. Across 12 wind-disturbed simulation settings, WAND improved the observed success rate by 8.3 percentage points on average relative to feedforward compensation alone. Controlled opposite-crosswind experiments further showed wind-direction-dependent trajectory adaptation. In indoor fan-induced flight tests, WAND succeeded in 18 of 20 trials, demonstrating the feasibility of real-time onboard navigation.
Figures & tables
Fig. 1 : Real-world deployment of WAND. (a) A quadrotor navigates a cluttered indoor scene under fan-induced airflow; the inset shows an anemometer reading of approximately 7.2m/s . (b) The corresponding visualization in RViz.
Fig. 2 : Overview of WAND framework. Ray-cast obstacle perception is first encoded into compact features, while a TCN-based wind disturbance estimator predicts the wind-induced disturbance acceleration from a temporal history of proprioceptive states. The estimated disturbance is encoded by WindAdapter and fused with the state feature to form the input to the policy network. The policy outputs a desired acceleration command, which is further combined with the feedforward disturbance compensation term for execution.
Fig. 3 : Analytical wind-velocity profiles used in simulation.
Fig. 4 : Reach-goal rates during PPO training: (a) navigation-only comparison and (b) WAND ablations. Curves show the mean over five independent random seeds, and shaded regions indicate 95% confidence intervals across seeds.
Fig. 5 : Navigation success rates under constant, turbulent, gust, and mixed wind. Markers show success rates over 50 trials; error bars indicate 95% Wilson confidence intervals. Markers are slightly offset horizontally for clarity.
Method
Constant
Turbulent
Gust
Mixed
EKF
0.3204
0.4620
0.4203
0.4474
LSTM
0.1166
0.2405
0.2059
0.2525
TCN
0.0777
0.1876
0.1159
0.1970
TABLE I : RMSE( m/s2 ) IN DIFFERENT WIND TYPES
Fig. 6 : Fixed-scene trajectories under opposite 7.5m/s crosswinds. The translucent volume and arrows indicate the prescribed wind region and direction. Each panel overlays 20 trials per method; WAND and Base+Comp are shown in blue and orange, while red crosses and yellow squares mark collisions and timeouts.
Wind
Method
S/C/T (%)
dmin (m)
+Wx
Base+Comp
60/20/20
0.086±0.089
WAND
100/0/0
0.376±0.022
−Wx
Base+Comp
0/100/0
–
WAND
100/0/0
0.449±0.107
TABLE II : FIXED-SCENE RESULTS UNDER OPPOSITE CROSSWINDS.
Fig. 7 : Real-world trajectories under (a) weak and strong fixed-fan wind and (b) strong handheld-fan wind.
Method
Scene 1
Scene 2
Total
EGO-Planner
4/10
6/10
10/20
NavRL
5/10
6/10
11/20
Base+Comp
7/10
6/10
13/20
WAND
10/10
8/10
18/20
TABLE III : REAL-WORLD NAVIGATION RESULTS (SUCCESSES/TRIALS)
Small multirotor aircraft are increasingly tasked with operations in the atmospheric boundary layer, where turbulent winds comparable to the vehicle's airspeed degrade trajectory tracking and can defeat conventional feedback control. This work illustrates a two-stage learning pipeline that first estimates the local wind from onboard kinematics and dynamics and then exploits that estimate inside a reinforcement learning (RL) flight controller. The wind estimator, an attention-augmented gated recurrent network trained on thousands of simulated flights through von Karman turbulence with power-law shear and veer, recovers the horizontal wind vector with a per-flight root-mean-square error of 0.40 m/s and a direction error of 3.2 degrees on unseen wind regimes, an accuracy near the floor imposed by unresolved turbulence, and generalizes to vertical ascent profiles with a skill score of 0.861 over a constant-wind reference. A proximal policy optimization controller receiving the frozen estimator's output reduces horizontal trajectory tracking error by 48% relative to a wind-blind proportional-derivative baseline across mean winds of 4 m/s to 12 m/s, winning on 100% of evaluation episodes. A three-way ablation decomposes this improvement into a kinematic component, available without wind information, and a wind-perception component; the perception share rises with wind speed, from small in light winds toward roughly half the total benefit in strong winds, consistent with the quadratic scaling of aerodynamic drag. The controller degrades gracefully on out-of-distribution winds of 13 m/s to 15 m/s, where the baseline fails catastrophically.
Safe quadrotor navigation in cluttered and dynamic environments depends not only on instantaneous geometric perception, but more critically on anticipating collision risks induced by relative motion. Conventional modular pipelines frequently suffer from perception latency, while end-to-end learning methods relying on implicit scalar rewards often struggle to extract reliable spatio-temporal features without physics-grounded supervision. To address this, we propose an anticipatory risk-guided reinforcement learning framework. Leveraging privileged simulator states, we construct a directionally aligned future collision risk map based on the Closest Point of Approach (CPA). Through an asymmetric actor-critic architecture, the network is trained to self-predict this structured risk, which explicitly guides the visual policy during deployment. A lightweight spatio-temporal encoder extracts motion cues directly from onboard depth sequences, bypassing explicit object tracking or optical flow estimation. Extensive simulated and real-world experiments demonstrate that our method effectively improves safety margins and flight efficiency in dense dynamic clutters compared to existing baselines. Furthermore, the learned policy achieves robust zero-shot Sim-to-Real transfer on a physical quadrotor, relying purely on abstracted spatio-temporal depth sequences and its self-predicted risk priors, validating the effectiveness of our approach and its robust generalization from simulation to reality.
Yuchao Mei, Guohao Zhang, Luxia Ai +2
National Key Laboratory of Science and Technology on Multi-spectral Information Processing, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology, Wuhan 430074, China
Deep Reinforcement Learning (DRL) for quadrotor flight control typically relies on Domain Randomization (DR) for sim-to-real transfer, resulting in overly conservative policies that struggle with dynamic disturbances. To overcome this, we propose a novel adaptive control architecture that actively perceives and reacts to instantaneous perturbations. First, we train an optimal outer-loop policy, then replace its reliance on ground-truth disturbance data with a Residual Dynamics Predictor (RDP). The RDP estimates the external forces and moments acting on the aircraft in flight online using only the history of states and control actions. For seamless hardware transfer, we introduce a data-efficient linear calibration bridge and an online thrust correction mechanism that align the simulated latent space with reality using mere seconds of flight data. Real-world validations on a Crazyflie micro-quadrotor demonstrate that our adaptive controller significantly outperforms baselines, maintaining precise trajectory tracking under severe uncertainties including mass variations, asymmetric payloads, and dynamic slung loads