Battery-Aware Reinforcement Learning for Aggressive Quadrotor Flight
Authors: Alejandro Sanchez Roncero, Olov Andersson, Petter Ogren
Organizations: Department of Robotics, Perception and Learning, School of Electrical Engineering and Computer Science, Royal Institute of Technology (KTH), SE-100 44 Stockholm, Sweden
Agile flight tasks such as drone racing and pursuit-evasion require strong acceleration and precise turns, but the available thrust changes as the battery discharges and voltage drops under load. Conservative command limits make this variation easier to tolerate, at the cost of unused performance. We investigate how learned controllers can use that additional thrust while retaining the flight controller's voltage compensation and rate control. Our training simulator couples an identified load-transient battery model to rotor dynamics and firmware saturation. The feedforward policy receives filtered voltage during both training and deployment. Controlled ablations distinguish the benefit of a larger thrust-command range from that of voltage information. On a 38 g Crazyflie Brushless, the resulting policy reduces circle tracking error by 49% relative to stock-authority RL at 3.84 m/s, while preserving easy-task precision. Mean 20-lap race time decreases from 106.22 s to 95.24 s. Compared with a voltage-blind policy with the same increased authority, hardware error and race time are lower by 15.3% and 4.5%, respectively. In simulation, replacing the policy's voltage input with a recording from a different battery condition worsens hard-circle tracking, with a smaller, voltage-dependent effect in racing. Together, these results show where a simple voltage input complements existing actuator compensation in aggressive learned flight.
Figures & tables
Fig. 1: Hardware flights with Stock (default thrust-command limit, no voltage input) and High-V (higher limit with voltage input). (a,b) Tracking at 3.36 and 3.84 m/s, with root-mean-square error (RMSE) averaged over repeated flights. (c) Racing over 20 laps. The overlays show measured paths on calibrated camera images and virtual gates in racing.
Parameter
Circle
Racing
Actor hidden layers, ELU
256,256,128
256,256,128
Symmetric critic hidden layers
256,256,128
256,256,128
Privileged critic: observation
256,128
256,128
Privileged critic: battery
64,64
64,64
Privileged critic: value
256,128,1
256,128,1
No voltage / scalar voltage
42/43 inputs
43/44 inputs
TABLE I: Learning and command settings
Quantity
Nominal
Randomization
Mass m [g]
38
±5%
diagJ [ 10−5 kg m 2 ]
3.3,3.6,5.9
±20%
Rotor inertia Jr [kg m 2 ]
5×10−8
Fixed
Motor gain multiplier
1
±3/2%
Thrust multiplier
1
±5/3%
Motor lag τm [ms]
50
±15/5%
TABLE II: Physical model and training randomization
Metric
Stock
High
High-V
Four-lap race [s]
21.30 ± 0.10
20.07 ± 0.16
19.26 ± 0.27
Flying lap, short [s]
5.33 ± 0.03
5.04 ± 0.04
4.81 ± 0.08
20-lap race [s]
106.22 ± 0.26
99.76 ± 0.84
95.24 ± 0.44
Flying lap, long [s]
5.31 ± 0.01
4.99 ± 0.04
4.76 ± 0.02
Mean speed, long [m/s]
3.74
3.89
4.10
Peak speed, long [m/s]
6.08 ± 0.11
6.74 ± 0.46
6.95 ± 0.05
TABLE III: Hardware racing performance
Fig. 9: Mean paired percentage reduction in error or lap time with a privileged rather than symmetric critic, where positive values indicate improvement. Parentheses identify alternative voltage filters, and colors are capped at ±25% .