Learning to Orchestrate Evolutionary Search: Progression-Aware Deep Reinforcement Learning for Dynamic DE-CMA-ES Coordination in Optimization and Structural Model Updating
Authors: Lechen Li, Rongye Shi, Wanhuan Zhou
Organizations: State Key Laboratory of Internet of Things for Smart City, University of Macau, Macau 519000, China · College of Water Conservancy and Hydropower Engineering, Hohai University, Nanjing 210098, China · School of Artificial Intelligence, Beihang University, Beijing 100191, China
Solving high-dimensional structural model updating problems requires an algorithm capable of navigating complex, non-convex landscapes with correlated parameters. Existing hybrid evolutionary algorithms typically rely on static architectures or fixed switching rules, resulting in disjointed search phases. To address this, this study proposes a Deep Reinforcement Learning-governed dynamic DE-CMAES Orchestration (DRL-DCO) algorithm, in which a Deep Deterministic Policy Gradient (DDPG)-based actor-critic agent continuously governs the evolutionary process as a single, unified system rather than a mechanical concatenation of algorithms. Guided by a progression-aware state representation and a diversity-informed reward, the agent fluidly reallocates computational resources between the differencevector-based exploration of Differential Evolution (DE) and the covariance-guided exploitation of CMA-ES, while jointly regulating population size, elite preservation, and a restart mechanism to escape local optima. This allows DRL-DCO to autonomously transition between exploration-dominant, exploitation-dominant, and mixed-strategy regimes across generations. Beyond the training phase, the trained actor can operate in a supervision-free inference mode, where the internalized policy autonomously orchestrates DE and CMA-ES control from observed search states through forward inference alone, without critic evaluation or weight updates, enabling faster deployment while retaining full effectiveness. Validated on high-dimensional single-objective optimization benchmarks and the IASC-ASCE structural health monitoring benchmark, DRL-DCO achieves superior convergence accuracy and robustness compared to state-of-the-art adaptive and hybrid evolutionary algorithms, as well as single-operator DRL-governed baselines.
Figures & tables
Figure 1: The flowchart of the proposed DRL-DCO algorithm.
Variable
F(t)
CR(t)
λ(t)
LRcov(t)
Range
[0.01, 2.00]
[0.01, 1.00]
[0.75, 1.00]
[0.50, 1.50]
Variable
Δσ(t)
pres(t)
reli(t)
β(t)
Range
[ − 0.20, 0.20]
[0.00, 0.30]
[0.00, 0.30]
[0.10, 0.90] †
Table 1: Prescribed lower and upper bounds for the action variables in the DRL-DCO algorithm.
Input: FEMU or general single-objective objective function f(θ) , search bounds, population size, maximum episodes Nepi , maximum generations per episode Ngen .
Output: The optimal solution vector θ∗ .
1: Initialize DE operators, CMA-ES parameters, actor π(s;θπ) , critic Q(s,a;θQ) , and replay memory.
2: for episode e=1,…,Nepi do
3: Initialize or inherit population χ(0) according to the warm-up/training strategy.
4: Construct initial state s(0) from χ(0) .
Table 2: Pseudocode of the proposed DRL-DCO algorithm.
Function
Standard search domain
Global optimum
Ackley
[−32.768,32.768]Nd
f(x∗)=0 at x∗=0
Griewank
[−600,600]Nd
f(x∗)=0 at x∗=0
Rastrigin
[−5.12,5.12]Nd
f(x∗)=0 at x∗=0
Rosenbrock
[−30,30]Nd
f(x∗)=0 at x∗=1
Schwefel
[−500,500]Nd
f(x∗)=0 at x∗=420.9687
Table 3: Standard search domains, global optima, and optimal points of the five benchmark optimization functions adopted for algorithm validation.
Figure 2: Convergence histories of best fitness and effective population distribution of the five baseline algorithms and the DRL-DCO algorithm (agent-training phase) on the Rastrigin and Rosenbrock functions at dimensionality Nd=50 . (a) Convergence history on Rastrigin; (b) effective population size distribution on Rastrigin; (c) convergence history on Rosenbrock; (d) effective population size distribution on Rosenbrock.
Figure 3: Evolution of the elite ratio and orchestration ratio in the DRL-DCO algorithm (agent-training phase) on the Rastrigin and Rosenbrock functions at dimensionality Nd=50 . (a) Elite ratio evolution on Rosenbrock; (b) orchestration ratio evolution on Rosenbrock; (c) elite ratio evolution on Rastrigin; (d) orchestration ratio evolution on Rastrigin.
Function
Algorithm
Avg total generations
Avg total function evaluations
Approx. time (s)
Ackley
JADE
40,000 (failure)
4,000,000 (failure)
890–990
LSHADE
573 (118)
56,988 (11,756)
14–18
LSHADE-SPACMA
386 (72)
38,592 (7,143)
10–13
DRL-CMA-ES
18,194 (1,023)
1,268,258 (99,143)
230–300
DRL-DE
15,956 (1,312)
1,139,602 (105,672)
310–390
DRL-DCO
5,923 (556)
514,622 (47,443)
85–130
Table 4: Convergence performance of the five baseline algorithms and the developed DRL-DCO algorithm (agent-training phase) on the five benchmark optimization functions at dimensionality Nd=50 . All results are averages over 10 independent runs, with standard deviations in parentheses. “Failure” indicates that none of the 10 independent runs achieved the convergence threshold ε=1×10−8 within the maximum budget of 40,000 total generations, in which case no standard deviation is reported.
Function
Avg total generations
Avg total function evaluations
Approx. time (s)
Ackley
5,176 (202)
406,656 (16,195)
32–35
Griewank
556 (63)
42,571 (3,014)
4–6
Rastrigin
5,510 (217)
386,286 (17,276)
33–36
Rosenbrock
14,331 (498)
1,003,834 (45,277)
90–96
Schwefel
8,989 (379)
677,712 (31,168)
52–57
Table 5: Convergence performance of the developed DRL-DCO algorithm (agent-inference phase) on the five benchmark optimization functions at dimensionality Nd=50 , obtained from one round of the agent-training phase (40 episodes). All results are averages over 10 independent runs, with standard deviations in parentheses.
Figure 4: Convergence performance of the developed DRL-DCO algorithm during the agent-inference phase on the Rastrigin function at dimensionality Nd=50 . (a) Convergence history; (b) effective population size evolution; (c) elite ratio evolution; (d) orchestration ratio evolution.
Figure 5: Convergence performance of the developed DRL-DCO algorithm during the agent-inference phase on the Rosenbrock function at dimensionality Nd=50 . (a) Convergence history; (b) effective population size evolution; (c) elite ratio evolution; (d) orchestration ratio evolution.
Figure 6: (a) Physical IASC–ASCE benchmark structure; (b) analytical FEM of the benchmark structure; and (c) simplified shear-type FEM used for stiffness parameter identification.
Figure 7: Convergence histories of the DRL-DCO algorithm for the FEMU on the IASC-ASCE benchmark structure, specifically regarding the best fitness, effective population size, elite ratio, and orchestration ratio, with (a) – (d) for the agent-training phase and (e) – (h) for the agent-inference phase.
Figure 8: Identified stiffness values of the IASC-ASCE benchmark structure in its baseline condition by the DRL-DCO algorithm, with (a) agent-training phase results and (b) agent-inference phase results; values above each bar are the corresponding averages and standard deviations (in parentheses), computed across 10 independent runs.
Total generations
Total evaluations
Ave. time (s)
Agent-training phase
261 (63)
17,642 (3,828)
12–16
Agent-inference phase
338 (41)
25,176 (2,463)
14–17
Table 6: Convergence performance of the DRL-DCO algorithm for the FEMU of the IASC-ASCE benchmark structure’s baseline condition. All results are averages over 10 independent runs, with standard deviations in parentheses.
Condition index
Damage description
Damage Scenario 1
All braces at the 1 st floor are broken.
Damage Scenario 2
All braces at both the 1 st and 3 rd floors are broken.
Damage Scenario 3
One brace at the 1 st floor is broken.
Damage Scenario 4
Two braces are broken, at the 1 st and the 3 rd floors.
Damage Scenario 5
One brace at the 1 st floor has a one-third reduction in stiffness.
Table 7: Considered damage scenarios of the benchmark structure in Case Study 2.
Stiffness
Scenario 1
Scenario 2
Scenario 3
Scenario 4
Scenario 5
k1(1)
45.24%
45.24%
0
0
0
k2(1)
71.03%
71.03%
17.76%
17.76%
5.92%
k3(1)
64.96%
64.96%
9.87%
9.87%
2.88%
k1(2)
0
0
0
0
0
k2(2)
0
0
0
0
0
k3(2)
0
0
0
0
0
Table 8: Equivalent stiffness percentage loss of each DOF for the damage scenarios in Case Study 2 [48].
Figure 9: Identified stiffness values of the IASC-ASCE benchmark structure for the considered five damaged conditions by the DRL-DCO algorithm in its agent-inference phase, averaged across 10 independent runs, with (a) – (e) representing Damage Scenarios 1–5, respectively.
This paper proposes RCMAES, a novel variant of the Covariance Matrix Adaptation Evolution Strategy (CMA-ES) for CEC benchmark optimization. RCMAES integrates a dimension-dependent nonlinear population-size reduction strategy with an adaptive restart mechanism within a pure CMA-ES framework. RCMAES is evaluated on three benchmark suites (CEC2017, CEC2020, and CEC2022) and compared with state-of-the-art DE algorithms as well as its closely related counterpart, BIPOP-aCMAES. Experimental results show that RCMAES achieves competitive and robust performance across all benchmarks.
Khoirul Faiq Muzakka, Sören Möller, Martin Finsterbusch
Institute of Energy Materials and Devices (IMD-2), Forschungszentrum Jülich GmbH, Germany
While deep Reinforcement Learning (deep-RL) has been increasingly applied to parameter control in evolutionary algorithms, rigorous theoretical analysis of parameter control remains largely restricted to single-parameter settings, owing to the difficulty of deriving effective, interpretable multi-parameter policies amenable to formal study. We demonstrate how deep-RL can be leveraged to overcome this barrier, using the (1+(λ,λ))-genetic algorithm optimizing OneMax, one of the few problems where a super-constant speedup of dynamic control has been formally proven, as a representative case study. We first show that standard approaches struggle to converge in this multi-parameter setting, and introduce algorithm-agnostic enhancements targeting action-space decomposition, reward shifting, and long-horizon discounting. With these in place, we compare common deep-RL methods and find that Double Deep Q-Networks uniquely avoid the policy collapse observed in Proximal Policy Optimization, yielding trajectories suitable for downstream analysis. Crucially, we move beyond the ``black-box'' nature of neural networks by distilling the learned behaviors into a transparent, symbolic control policy. This resulting policy does not only offer interpretability for future theoretical analysis but also yields exceptional performance, consistently outperforming existing baselines across a wide range of problem sizes.
Tai Nguyen, Phong Le, Carola Doerr +1
University of St Andrews, St Andrews, United Kingdom · Sorbonne Universit´e, CNRS, LIP6, Paris, France
In the last few decades, Markov chain Monte Carlo (MCMC) methods have been widely applied to Bayesian updating of structural dynamic models in the field of structural health monitoring. Recently, several MCMC algorithms have been developed that incorporate neural networks to enhance their performance for specific Bayesian model updating problems. However, a common challenge with these approaches lies in the fact that the embedded neural networks often necessitate retraining when faced with new tasks, a process that is time-consuming and significantly undermines the competitiveness of these methods. This paper introduces a newly developed adaptive meta-learning stochastic gradient Hamiltonian Monte Carlo (AM-SGHMC) algorithm. The idea behind AM-SGHMC is to optimize the sampling strategy by training adaptive neural networks, and due to the adaptive design of the network inputs and outputs, the trained sampler can be directly applied to various Bayesian updating problems of the same type of structure without further training, thereby achieving meta-learning. Additionally, practical issues for the feasibility of the AM-SGHMC algorithm for structural dynamic model updating are addressed, and two examples involving Bayesian updating of multi-story building models with different model fidelity are used to demonstrate the effectiveness and generalization ability of the proposed method.
Xianghao Meng, James L. Beck, Yong Huang +1
Key Lab of Smart Prevention and Mitigation of Civil Engineering Disasters of the Ministry of Industry and Information Technology, Harbin Institute of Technology, Harbin, China · Key Lab of Structures Dynamic Behavior and Control of the Ministry of Education, Harbin Institute of Technology, Harbin, China · Division of Engineering and Applied Science, California Institute of Technology, CA, USA