Swarm robotics presents a robust and cost-effective paradigm for advanced automation in complex, dynamic environments, such as those encountered in search and rescue or environmental monitoring. A fundamental challenge for this field is the data-driven design of decentralized controllers capable of generating emergent collective behaviors. This paper proposes a novel, AI-driven hybrid methodology for the automatic synthesis of swarm robotic controllers for autonomous visual navigation. This approach synergistically combines multi-agent reinforcement learning with neuro-evolutionary strategies, specifically leveraging implementations of the cross-entropy method and the covariance matrix adaptation evolution strategy to optimize a pre-trained individual navigation policy. The underlying deep architecture is engineered for low-cost, resource-constrained platforms, utilizing a compact neural network that relies exclusively on monocular camera imagery. This vision-based design emphasizes computational and energy efficiency, a critical requirement for practical swarm deployments. Experiments, performed in a high-fidelity physics simulator, demonstrate that the resulting controllers enable robust and scalable collective exploration of diverse indoor environments. The controller trained using our cross-entropy method achieves superior exploration coverage, visiting 36.20% more regions compared to the covariance matrix adaptation evolution strategy. Critically, our best vision-based policy achieves exploration performance statistically comparable to traditional methods relying on more expensive distance sensors, while delivering a significant 31.40% average reduction in energy consumption. These findings validate an effective and economically viable autonomous control system, establishing a path for deploying highly efficient collective intelligence in real-world engineering applications.
Figures & tables
Figure 1: Microscopic model architecture. This figure details the layer structure of the convolutional variational autoencoder (bottom), which handles perception by compressing high-dimensional images, and the lightweight classification neural network (top), which serves as the controller for action selection.
Figure 2: Relationship between the microscopic and macroscopic models of the proposed swarm robotic system. The parameters of the collective navigation policy are updated based on the individual rewards assigned to agents for actions they execute in response to their partial observations of the environment.
Component
Definition
Implementation
Agents
E
Homogeneous swarm of J robots operating under a fully decentralized, map-less paradigm.
States
S
Unobservable global variables ( e.g. , time steps T per episode and collision termination threshold C ).
Observations
O
Local partial images: 64x64 RGB camera feed per agent, processed into a 32-element latent vector by the VAE G .
Table 1: Systematic summary of the Dec-POMDP formulation for our swarm exploration task.
Figure 3: Flowchart of the proposed CEM algorithm.
Figure 4: Flowchart of the proposed CMA-ES algorithm.
Figure 5: Division of the simulated environment into 22 rectangular sectors. This map shows the partitioned indoor corridor (in green) used for quantitative evaluation of spatial coverage. The corridor has a navigable surface approximately 20 meters long and 2 meters wide.
Figure 6: Evolution of the mean accumulated reward during 2000 generations of CEM training. The curves represent the filtered improvement progress for hyperparameter configurations (1) (in blue), (2) (in orange), and (3) (in green).
Hyperparameter
Description
Configuration
(1)
(2)
(3)
J
Agents in the system
4
4
4
T
Time steps per episode
500-4500
500-4500
500-4500
C
Collision termination threshold
0
0
0
κ
Total generations
2000
2000
2000
ρ
Population size
32
32
32
Table 2: Hyperparameter configurations for the CEM algorithm sensitivity analysis.
Figure 7: Number of tests achieving specific numbers of visited sectors. The diagrams compare the ability of our three CEM hyperparameter configurations to explore the indoor corridor across 100 validation trials.
Configuration
Sectors
Reward
(1)
16.93
3573.56
(2)
20.38
4673.10
(3)
14.55
2483.32
Table 3: Quantitative performance comparison of CEM configurations across 100 validation trials.
Figure 8: Evolution of the mean accumulated reward during 2000 generations of CMA-ES training. The curves represent the filtered improvement progress for hyperparameter configurations (4) (in blue), (5) (in orange), and (6) (in green).
Hyperparameter
Description
Configuration
(4)
(5)
(6)
J
Agents in the system
4
4
4
T
Time steps per episode
500-4500
500-4500
500-4500
C
Collision termination threshold
0
0
0
κ
Total generations
2000
2000
2000
ρ
Population size
32
32
32
Table 4: Hyperparameter configurations for the CMA-ES algorithm sensitivity analysis.
Figure 9: Number of tests achieving specific numbers of visited sectors. The diagrams compare the ability of our three CMA-ES hyperparameter configurations to explore the indoor corridor across 100 validation trials.
Configuration
Sectors
Reward
(4)
17.75
1195.83
(5)
19.44
4587.92
(6)
16.86
890.18
Table 5: Quantitative performance comparison of CMA-ES configurations across 100 validation trials.
Figure 10: Number of tests achieving specific numbers of visited sectors. This bar chart compares the collective exploration performance of the LiDAR baseline (in blue), the proposed vision-based controller trained with our CEM algorithm (in orange), and the proposed vision-based controller trained with our CMA-ES algorithm (in green) over 100 multi-agent reinforcement learning episodes.
Method
Sectors
Reward
Collisions (%)
Distance (m)
Consumption (Ws)
LiDAR
20.91
5827.51
4.68
162.05
3341.99
CEM
20.38
4673.10
2.69
166.39
2316.93
CMA-ES
19.44
4587.92
2.41
157.58
1987.14
Table 6: Average test results for sectors explored, reward points, percentage of time with inter-robot collisions, traveled distance, and energy consumption.
Figure 11: Navigational trajectory frequency heatmaps. This figure illustrates the number of times four robots traversed different positions in the simulated environment over 100 episodes, comparing the behaviors resulting from a) the LiDAR baseline, b) the policy obtained through our proposed cross-entropy method, and c) the policy generated using our proposed covariance matrix adaptation evolution strategy.
Evaluation Rule
Method
Sectors
Reward
With collision recovery
CEM
20.38
4673.10
With collision recovery
CMA-ES
19.44
4587.92
Without collision recovery
CEM
11.48
752.35
Without collision recovery
CMA-ES
13.61
1427.84
Table 7: Average test results for sectors explored and reward points obtained by the vision-based controllers with and without the ability to recover and continue navigation after a collision.
Comparison
Difference
p-Value
Confidence Interval (95%)
Significant
CEM vs LiDAR
-0.53
0.2008
[-1.2575, 0.1975]
No
CMA-ES vs LiDAR
-1.47
0.0
[-2.1975, -0.7425]
Yes
CMA-ES vs CEM
-0.94
0.0072
[-1.6675, -0.2125]
Yes
Table 8: Results of the Tukey HSD test for comparing means of sectors explored by the three collective autonomous navigation controllers evaluated.
Identifier
Name
Sectors
(0)
Original corridor
22
(1)
Corridor with corner
26
(2)
Corridor with intersection
30
(3)
Corridor with passage
33
Table 9: Synthetic environments used to evaluate the robustness, flexibility, and scalability of the developed robotic swarm.
Figure 12: Navigational trajectory frequency heatmaps with 5% noise. This figure illustrates the number of times four robots traversed different positions in the simulated environment (0) over 100 episodes with 5% Gaussian noise added to sensory inputs, comparing behaviors from a) the LiDAR baseline, b) the proposed CEM-trained policy, and c) the proposed CMA-ES-trained policy.
Figure 13: Navigational trajectory frequency heatmaps with 10% noise. This figure illustrates the number of times four robots traversed different positions in the simulated environment (0) over 100 episodes with 10% Gaussian noise added to sensory inputs, comparing behaviors from a) the LiDAR baseline, b) the proposed CEM-trained policy, and c) the proposed CMA-ES-trained policy.
Figure 14: Navigational trajectory frequency heatmaps with 15% noise. This figure illustrates the number of times four robots traversed different positions in the simulated environment (0) over 100 episodes with 15% Gaussian noise added to sensory inputs, comparing behaviors from a) the LiDAR baseline, b) the proposed CEM-trained policy, and c) the proposed CMA-ES-trained policy.
Noise
Method
Sectors
Reward
Collisions (%)
Distance (m)
Consumption (Ws)
5%
LiDAR
20.76
5305.83
4.10
154.82
3367.04
5%
CEM
20.61
4221.22
2.42
165.15
2862.62
5%
CMA-ES
18.65
3142.97
1.99
136.91
2113.32
10%
LiDAR
19.76
4174.42
5.26
133.09
3614.75
10%
CEM
20.66
3723.24
3.01
159.79
2367.03
10%
CMA-ES
15.58
1631.54
3.09
109.89
1979.73
Table 10: Average robustness test results across different noise levels for sectors explored, reward points, percentage of time with inter-robot collisions, traveled distance, and energy consumption.
Figure 15: New synthetic environments for flexibility analysis. This figure illustrates the division of the three novel simulated environments into discrete sectors (in green) with their corresponding indices for quantitative evaluation. The coordinate axes correspond to three distinct interior spaces: a) a corridor with a corner, b) a corridor with an intersection, and c) a corridor with a narrow passage.
Figure 16: Navigational trajectory frequency heatmaps in environment (1). This figure illustrates the number of times four robots traversed different positions in the simulated corridor with a corner over 100 episodes. This flexibility assessment compares behaviors from: a) the LiDAR baseline, b) the proposed CEM-trained policy, and c) the proposed CMA-ES-trained policy.
Figure 17: Navigational trajectory frequency heatmaps in environment (2). This figure illustrates the number of times four robots traversed different positions in the simulated corridor with an intersection over 100 episodes. This flexibility assessment compares behaviors from: a) the LiDAR baseline, b) the proposed CEM-trained policy, and c) the proposed CMA-ES-trained policy.
Figure 18: Navigational trajectory frequency heatmaps in environment (3). This figure illustrates the number of times four robots traversed different positions in the simulated corridor with a passage over 100 episodes. This flexibility assessment compares behaviors from: a) the LiDAR baseline, b) the proposed CEM-trained policy, and c) the proposed CMA-ES-trained policy.
Corridor
Method
Sectors
Reward
Collisions (%)
Distance (m)
Consumption (Ws)
(1)
LiDAR
24.83
7177.43
1.92
193.27
3252.71
(1)
CEM
21.54
4493.92
1.64
156.60
2248.05
(1)
CMA-ES
18.42
2335.72
2.20
118.79
2087.76
(2)
LiDAR
20.41
5509.19
3.76
154.49
2887.14
(2)
CEM
21.61
5177.19
1.76
165.69
1781.02
(2)
CMA-ES
18.49
1842.17
2.55
113.78
1809.04
Table 11: Average flexibility test results across different environments for sectors explored, reward points, percentage of time with inter-robot collisions, traveled distance, and energy consumption.
Robots
Method
Sectors
Reward
Collisions (%)
Distance (m)
Consumption (Ws)
2
LiDAR
28.64
4198.51
0.37
111.96
1803.83
2
CEM
18.24
2976.72
0.13
91.25
1144.42
2
CMA-ES
9.33
337.84
0.45
38.14
843.78
8
LiDAR
32.71
15540.27
6.81
419.95
7348.64
8
CEM
25.09
12147.43
4.35
373.93
4827.89
8
CMA-ES
16.71
2124.36
11.29
177.44
3654.71
Table 12: Average scalability test results across different swarm sizes for sectors explored, reward points, percentage of time with inter-robot collisions, traveled distance, and energy consumption.
Figure 19: Navigational trajectory frequency heatmaps with 2 robots. This figure illustrates the number of times two robots traversed different positions in the simulated environment (3) over 100 episodes. This scalability assessment compares behaviors from: a) the LiDAR baseline, b) the proposed CEM-trained policy, and c) the proposed CMA-ES-trained policy.
Figure 20: Navigational trajectory frequency heatmaps with 8 robots. This figure illustrates the number of times eight robots traversed different positions in the simulated environment (3) over 100 episodes. This scalability assessment compares behaviors from: a) the LiDAR baseline, b) the proposed CEM-trained policy, and c) the proposed CMA-ES-trained policy.
Figure 21: Navigational trajectory frequency heatmaps with 16 robots. This figure illustrates the number of times sixteen robots traversed different positions in the simulated environment (3) over 100 episodes. This scalability assessment compares behaviors from: a) the LiDAR baseline, b) the proposed CEM-trained policy, and c) the proposed CMA-ES-trained policy.
Noise
Comparison
Difference
p-Value
Confidence Interval (95%)
Significant
5%
CEM vs LiDAR
-0.15
0.8726
[-0.8604, 0.5604]
No
5%
CMA-ES vs LiDAR
-2.11
0.0
[-2.8204, -1.3996]
Yes
5%
CMA-ES vs CEM
-1.96
0.0
[-2.6704, -1.2496]
Yes
10%
CEM vs LiDAR
0.9
0.0329
[0.0581, 1.7419]
Yes
10%
CMA-ES vs LiDAR
-4.18
0.0
[-5.0219, -3.3381]
Yes
10%
CMA-ES vs CEM
-5.08
0.0
[-5.9219, -4.2381]
Yes
Table 13: Results of the Tukey HSD test for comparing means of sectors explored by the three collective autonomous navigation controllers evaluated under different Gaussian noise levels.
Corridor
Comparison
Difference
p-Value
Confidence Interval (95%)
Significant
(1)
CEM vs LiDAR
-3.29
0.0
[-4.4961, -2.0839]
Yes
(1)
CMA-ES vs LiDAR
-6.41
0.0
[-7.6161, -5.2039]
Yes
(1)
CMA-ES vs CEM
-3.12
0.0
[-4.3261, -1.9139]
Yes
(2)
CEM vs LiDAR
1.2
0.2998
[-0.704, 3.104]
No
(2)
CMA-ES vs LiDAR
-1.92
0.0476
[-3.824, -0.016]
Yes
(2)
CMA-ES vs CEM
-3.12
0.0004
[-5.024, -1.216]
Yes
Table 14: Results of the Tukey HSD test for comparing means of sectors explored by the three collective autonomous navigation controllers evaluated in novel, unknown environments.
Robots
Comparison
Difference
p-Value
Confidence Interval (95%)
Significant
2
CEM vs LiDAR
-10.4
0.0
[-11.8718, -8.9282]
Yes
2
CMA-ES vs LiDAR
-19.31
0.0
[-20.7818, -17.8382]
Yes
2
CMA-ES vs CEM
-8.91
0.0
[-10.3818, -7.4382]
Yes
8
CEM vs LiDAR
-7.62
0.0
[-9.2031, -6.0369]
Yes
8
CMA-ES vs LiDAR
-16.0
0.0
[-17.5831, -14.4169]
Yes
8
CMA-ES vs CEM
-8.38
0.0
[-9.9631, -6.7969]
Yes
Table 15: Results of the Tukey HSD test for comparing means of sectors explored by the three collective autonomous navigation controllers evaluated across different swarm sizes.
Laboratory of Intelligent Systems, Ecole Polytechnique Federale de Lausanne (EPFL), CH1015 Lausanne, Switzerland · Department of Electronic and Computer Engineering, Hong Kong University of Science and Technology, Hong Kong