Micro Neural Policies for Safe Real-Time Robotic Control
Authors: Hongpeng Cao, Riccardo Curcio, Daniele Ottaviano, Marco Caccamo
Organizations: School of Engineering and Design, Technical University of Munich, Boltzmannstraße 15, 85748 Garching b. München, Germany · Department of Computer Science, Sapienza University of Rome, Via Salaria 113, 00198 Rome, Italy
In this paper, we investigate the synthesis of Micro Neural Policies (MNP) to enable safe and robust real-time robotic control on computationally constrained embedded devices. We demonstrate that integrating Evolution Strategy (ES) and Statistical Model Checking (SMC)-based verification for policy search can drastically reduce neural network size without compromising safety and robustness. We conduct a large-scale training and evaluation of MNP on Cartpole and Quadrotor control tasks, varying control frequencies and network architectures. After validating these policies in simulation, we evaluate their deployability through zero-shot transfer to physical systems. Our experiments show that MNP can successfully achieve safe sim-to-real transfer without sacrificing control performance. We then show that the policies' memory footprint, ranging from 0.5 to 7.5 kB, allows deployment on microcontrollers, where they achieve real-time inference latency with under 25 ns of jitter while leaving the chip idle for over 97% of the time for additional workloads. This makes them a highly practical solution for severely resource-constrained robotic systems.
Figures & tables
System
Variable
Constraint
Variable
Constraint
Cartpole
xt
∣xt∣<0.35m
θt
∣θt∣<0.8rad
x˙t
∣x˙t∣<2.0m/s
θ˙t
∣θ˙t∣<4.0rad/s
at
∣at∣<10N
Quadrotor
xt,yt
∣⋅∣<4.0m
ϕt,θt
∣⋅∣<1.5rad
zt
∣zt∣<2.5m
ψt
∣ψt∣<πrad
x˙t,y˙t,z˙t
∣⋅∣<5.0m/s
pt,qt,rt
∣⋅∣<2.0rad/s
TABLE I : State and action constraints for the Cartpole and Quadrotor systems.
Fig. 1 : Real-world experimental platforms: (a) Quanser Cartpole equipped with a Raspberry Pi 4B. (b) ANT-X quadrotor equipped with a Pixhawk flight controller (STM32F427).
Freq.
Metric
DRL
DRL-Res
ES
ES-Res
50 Hz
Verifiability
0/5
3/5
5/5
5/5
Return Lower Bound
7.46±4.39
251.97±223.71
461.57±11.64
457.35±5.26
Relative Enlarged Region ( % )
—
0.00±0.00
96.71±3.63
96.62±3.39
100 Hz
Verifiability
1/5
2/5
5/5
5/5
Return Lower Bound
85.52±167.63
188.05±207.56
442.40±7.46
430.52±8.33
Relative Enlarged Region ( % )
0.00±0.00
0.00±0.00
95.50±3.01
97.79±2.57
TABLE II : Safety and performance comparison in Cartpole simulation at the 50Hz and 100Hz control frequencies.
Freq.
Metric
DRL
DRL-Res
ES
ES-Res
50 Hz
Verifiability
0/5
0/5
3/5
5/5
Return Lower Bound
68.71±147.02
12.61±15.24
319.60±165.26
455.69±14.33
Relative Enlarged Region ( % )
—
—
81.64±8.93
94.13±2.87
100 Hz
Verifiability
0/5
0/5
4/5
5/5
Return Lower Bound
2.91±5.23
12.60±16.91
300.68±137.79
424.09±32.30
Relative Enlarged Region ( % )
—
—
85.80±15.76
87.46±6.08
TABLE III : Safety and performance comparison in Quadrotor simulation at the 50Hz and 100Hz control frequencies.
Fig. 2 : Learning capability evaluation in Cartpole Simulation.
Fig. 3 : Learning capability evaluation in Quadrotor Simulation.
Fig. 4 : Real-world Cartpole evaluation.
Fig. 5 : Quadrotor trajectory tracking in simulation and real-world experiments.
System
Architecture
Latency
Flash
RAM
Energy/inference
Power @50 Hz
Power @100 Hz
ES
ES-Res
ES
ES-Res
ES
ES-Res
ES
ES-Res
ES
ES-Res
ES
ES-Res
Cartpole
2×2
10.26 ± 0.00 µs
10.10 ± 0.00 µs
536 B
712 B
40 B
56 B
0.25 ± 0.07 µJ
0.29 ± 0.07 µJ
0.69 mW
0.69 mW
0.71 mW
0.70 mW
4×4
19.72 ± 0.00 µs
21.13 ± 0.00 µs
696 B
852 B
40 B
56 B
0.45 ± 0.05 µJ
0.40 ± 0.08 µJ
0.70 mW
0.70 mW
0.72 mW
0.71 mW
8×8
39.50 ± 0.02 µs
38.41 ± 0.00 µs
1,016 B
1,172 B
64 B
80 B
0.50 ± 0.07 µJ
0.55 ± 0.08 µJ
0.70 mW
0.70 mW
0.73 mW
0.74 mW
16×16
81.34 ± 0.00 µs
82.97 ± 0.00 µs
2,040 B
2,200 B
128 B
144 B
0.98 ± 0.08 µJ
0.87 ± 0.08 µJ
0.72 mW
0.72 mW
0.78 mW
0.77 mW
32×32
219.27 ± 0.00 µs
206.64 ± 0.01 µs
5,624 B
5,780 B
256 B
272 B
2.88 ± 0.09 µJ
2.83 ± 0.10 µJ
0.81 mW
0.81 mW
0.95 mW
0.98 mW
TABLE IV : On-device inference latency, memory and energy cost for Cartpole and Quadrotor policies.
Department of Computer Science, Sapienza University of Rome, via Salaria 113, 00198 Rome, Italy · School of Engineering and Design, Technical University of Munich, Boltzmannstraße 15, 85748 Garching b. München, Germany