cs.NEOct 8, 2026

Learning to Orchestrate Evolutionary Search: Progression-Aware Deep Reinforcement Learning for Dynamic DE-CMA-ES Coordination in Optimization and Structural Model Updating

Authors: Lechen Li, Rongye Shi, Wanhuan Zhou

Organizations: State Key Laboratory of Internet of Things for Smart City, University of Macau, Macau 519000, China · College of Water Conservancy and Hydropower Engineering, Hohai University, Nanjing 210098, China · School of Artificial Intelligence, Beihang University, Beijing 100191, China

Abstract

Solving high-dimensional structural model updating problems requires an algorithm capable of navigating complex, non-convex landscapes with correlated parameters. Existing hybrid evolutionary algorithms typically rely on static architectures or fixed switching rules, resulting in disjointed search phases. To address this, this study proposes a Deep Reinforcement Learning-governed dynamic DE-CMAES Orchestration (DRL-DCO) algorithm, in which a Deep Deterministic Policy Gradient (DDPG)-based actor-critic agent continuously governs the evolutionary process as a single, unified system rather than a mechanical concatenation of algorithms. Guided by a progression-aware state representation and a diversity-informed reward, the agent fluidly reallocates computational resources between the differencevector-based exploration of Differential Evolution (DE) and the covariance-guided exploitation of CMA-ES, while jointly regulating population size, elite preservation, and a restart mechanism to escape local optima. This allows DRL-DCO to autonomously transition between exploration-dominant, exploitation-dominant, and mixed-strategy regimes across generations. Beyond the training phase, the trained actor can operate in a supervision-free inference mode, where the internalized policy autonomously orchestrates DE and CMA-ES control from observed search states through forward inference alone, without critic evaluation or weight updates, enabling faster deployment while retaining full effectiveness. Validated on high-dimensional single-objective optimization benchmarks and the IASC-ASCE structural health monitoring benchmark, DRL-DCO achieves superior convergence accuracy and robustness compared to state-of-the-art adaptive and hybrid evolutionary algorithms, as well as single-operator DRL-governed baselines.

Figures & tables

Explore similar work

CardsList
  1. RCMAES: A Robust CMA-ES Variant for CEC2026 Competition

    Apr 29, 2026Khoirul Faiq Muzakka, Sören Möller, Martin FinsterbuschEvolutionary OptimizationBlack-Box Optimization

  2. Discovering Interpretable Multi-Parameter Control Policies for Evolutionary Algorithms Using Deep Reinforcement Learning

    Jun 8, 2026Tai Nguyen, Phong Le, Carola Doerr +1Evolutionary OptimizationDeep Q-Learning

  3. Adaptive Meta-Learning Stochastic Gradient Hamiltonian Monte Carlo Simulation for Bayesian Updating of Structural Dynamic Models

    Apr 28, 2026Xianghao Meng, James L. Beck, Yong Huang +1Markov Chain Monte CarloMeta-Learning