Learning to Orchestrate Evolutionary Search: Progression-Aware Deep Reinforcement Learning for Dynamic DE-CMA-ES Coordination in Optimization and Structural Model Updating
Authors: Lechen Li, Rongye Shi, Wanhuan Zhou
Organizations: State Key Laboratory of Internet of Things for Smart City, University of Macau, Macau 519000, China · College of Water Conservancy and Hydropower Engineering, Hohai University, Nanjing 210098, China · School of Artificial Intelligence, Beihang University, Beijing 100191, China
Solving high-dimensional structural model updating problems requires an algorithm capable of navigating complex, non-convex landscapes with correlated parameters. Existing hybrid evolutionary algorithms typically rely on static architectures or fixed switching rules, resulting in disjointed search phases. To address this, this study proposes a Deep Reinforcement Learning-governed dynamic DE-CMAES Orchestration (DRL-DCO) algorithm, in which a Deep Deterministic Policy Gradient (DDPG)-based actor-critic agent continuously governs the evolutionary process as a single, unified system rather than a mechanical concatenation of algorithms. Guided by a progression-aware state representation and a diversity-informed reward, the agent fluidly reallocates computational resources between the differencevector-based exploration of Differential Evolution (DE) and the covariance-guided exploitation of CMA-ES, while jointly regulating population size, elite preservation, and a restart mechanism to escape local optima. This allows DRL-DCO to autonomously transition between exploration-dominant, exploitation-dominant, and mixed-strategy regimes across generations. Beyond the training phase, the trained actor can operate in a supervision-free inference mode, where the internalized policy autonomously orchestrates DE and CMA-ES control from observed search states through forward inference alone, without critic evaluation or weight updates, enabling faster deployment while retaining full effectiveness. Validated on high-dimensional single-objective optimization benchmarks and the IASC-ASCE structural health monitoring benchmark, DRL-DCO achieves superior convergence accuracy and robustness compared to state-of-the-art adaptive and hybrid evolutionary algorithms, as well as single-operator DRL-governed baselines.
Figures & tables
Figure 1: The flowchart of the proposed DRL-DCO algorithm.
Variable
F(t)
CR(t)
λ(t)
LRcov(t)
Range
[0.01, 2.00]
[0.01, 1.00]
[0.75, 1.00]
[0.50, 1.50]
Variable
Δσ(t)
pres(t)
reli(t)
β(t)
Range
[ − 0.20, 0.20]
[0.00, 0.30]
[0.00, 0.30]
[0.10, 0.90] †
Table 1: Prescribed lower and upper bounds for the action variables in the DRL-DCO algorithm.
Input: FEMU or general single-objective objective function f(θ) , search bounds, population size, maximum episodes Nepi , maximum generations per episode Ngen .
Output: The optimal solution vector θ∗ .
1: Initialize DE operators, CMA-ES parameters, actor π(s;θπ) , critic Q(s,a;θQ) , and replay memory.
2: for episode e=1,…,Nepi do
3: Initialize or inherit population χ(0) according to the warm-up/training strategy.
4: Construct initial state s(0) from χ(0) .
Table 2: Pseudocode of the proposed DRL-DCO algorithm.
Function
Standard search domain
Global optimum
Ackley
[−32.768,32.768]Nd
f(x∗)=0 at x∗=0
Griewank
[−600,600]Nd
f(x∗)=0 at x∗=0
Rastrigin
[−5.12,5.12]Nd
f(x∗)=0 at x∗=0
Rosenbrock
[−30,30]Nd
f(x∗)=0 at x∗=1
Schwefel
[−500,500]Nd
f(x∗)=0 at x∗=420.9687
Table 3: Standard search domains, global optima, and optimal points of the five benchmark optimization functions adopted for algorithm validation.
Figure 2: Convergence histories of best fitness and effective population distribution of the five baseline algorithms and the DRL-DCO algorithm (agent-training phase) on the Rastrigin and Rosenbrock functions at dimensionality Nd=50 . (a) Convergence history on Rastrigin; (b) effective population size distribution on Rastrigin; (c) convergence history on Rosenbrock; (d) effective population size distribution on Rosenbrock.
Figure 3: Evolution of the elite ratio and orchestration ratio in the DRL-DCO algorithm (agent-training phase) on the Rastrigin and Rosenbrock functions at dimensionality Nd=50 . (a) Elite ratio evolution on Rosenbrock; (b) orchestration ratio evolution on Rosenbrock; (c) elite ratio evolution on Rastrigin; (d) orchestration ratio evolution on Rastrigin.
Function
Algorithm
Avg total generations
Avg total function evaluations
Approx. time (s)
Ackley
JADE
40,000 (failure)
4,000,000 (failure)
890–990
LSHADE
573 (118)
56,988 (11,756)
14–18
LSHADE-SPACMA
386 (72)
38,592 (7,143)
10–13
DRL-CMA-ES
18,194 (1,023)
1,268,258 (99,143)
230–300
DRL-DE
15,956 (1,312)
1,139,602 (105,672)
310–390
DRL-DCO
5,923 (556)
514,622 (47,443)
85–130
Table 4: Convergence performance of the five baseline algorithms and the developed DRL-DCO algorithm (agent-training phase) on the five benchmark optimization functions at dimensionality Nd=50 . All results are averages over 10 independent runs, with standard deviations in parentheses. “Failure” indicates that none of the 10 independent runs achieved the convergence threshold ε=1×10−8 within the maximum budget of 40,000 total generations, in which case no standard deviation is reported.
Function
Avg total generations
Avg total function evaluations
Approx. time (s)
Ackley
5,176 (202)
406,656 (16,195)
32–35
Griewank
556 (63)
42,571 (3,014)
4–6
Rastrigin
5,510 (217)
386,286 (17,276)
33–36
Rosenbrock
14,331 (498)
1,003,834 (45,277)
90–96
Schwefel
8,989 (379)
677,712 (31,168)
52–57
Table 5: Convergence performance of the developed DRL-DCO algorithm (agent-inference phase) on the five benchmark optimization functions at dimensionality Nd=50 , obtained from one round of the agent-training phase (40 episodes). All results are averages over 10 independent runs, with standard deviations in parentheses.
Figure 4: Convergence performance of the developed DRL-DCO algorithm during the agent-inference phase on the Rastrigin function at dimensionality Nd=50 . (a) Convergence history; (b) effective population size evolution; (c) elite ratio evolution; (d) orchestration ratio evolution.
Figure 5: Convergence performance of the developed DRL-DCO algorithm during the agent-inference phase on the Rosenbrock function at dimensionality Nd=50 . (a) Convergence history; (b) effective population size evolution; (c) elite ratio evolution; (d) orchestration ratio evolution.
Figure 6: (a) Physical IASC–ASCE benchmark structure; (b) analytical FEM of the benchmark structure; and (c) simplified shear-type FEM used for stiffness parameter identification.
Figure 7: Convergence histories of the DRL-DCO algorithm for the FEMU on the IASC-ASCE benchmark structure, specifically regarding the best fitness, effective population size, elite ratio, and orchestration ratio, with (a) – (d) for the agent-training phase and (e) – (h) for the agent-inference phase.
Figure 8: Identified stiffness values of the IASC-ASCE benchmark structure in its baseline condition by the DRL-DCO algorithm, with (a) agent-training phase results and (b) agent-inference phase results; values above each bar are the corresponding averages and standard deviations (in parentheses), computed across 10 independent runs.
Total generations
Total evaluations
Ave. time (s)
Agent-training phase
261 (63)
17,642 (3,828)
12–16
Agent-inference phase
338 (41)
25,176 (2,463)
14–17
Table 6: Convergence performance of the DRL-DCO algorithm for the FEMU of the IASC-ASCE benchmark structure’s baseline condition. All results are averages over 10 independent runs, with standard deviations in parentheses.
Condition index
Damage description
Damage Scenario 1
All braces at the 1 st floor are broken.
Damage Scenario 2
All braces at both the 1 st and 3 rd floors are broken.
Damage Scenario 3
One brace at the 1 st floor is broken.
Damage Scenario 4
Two braces are broken, at the 1 st and the 3 rd floors.
Damage Scenario 5
One brace at the 1 st floor has a one-third reduction in stiffness.
Table 7: Considered damage scenarios of the benchmark structure in Case Study 2.
Stiffness
Scenario 1
Scenario 2
Scenario 3
Scenario 4
Scenario 5
k1(1)
45.24%
45.24%
0
0
0
k2(1)
71.03%
71.03%
17.76%
17.76%
5.92%
k3(1)
64.96%
64.96%
9.87%
9.87%
2.88%
k1(2)
0
0
0
0
0
k2(2)
0
0
0
0
0
k3(2)
0
0
0
0
0
Table 8: Equivalent stiffness percentage loss of each DOF for the damage scenarios in Case Study 2 [48].
Figure 9: Identified stiffness values of the IASC-ASCE benchmark structure for the considered five damaged conditions by the DRL-DCO algorithm in its agent-inference phase, averaged across 10 independent runs, with (a) – (e) representing Damage Scenarios 1–5, respectively.
Key Lab of Smart Prevention and Mitigation of Civil Engineering Disasters of the Ministry of Industry and Information Technology, Harbin Institute of Technology, Harbin, China · Key Lab of Structures Dynamic Behavior and Control of the Ministry of Education, Harbin Institute of Technology, Harbin, China · Division of Engineering and Applied Science, California Institute of Technology, CA, USA