Constraint handling methods typically intervene before the optimizer acts, by modifying the objective or the gradient. Yet momentum, adaptive scaling, and structured preconditioning can substantially reshape that signal before it becomes a parameter update. We formulate optimizer relative constrained learning, where constraint compatibility is assessed on the post optimizer update. Building on this view, we introduce ORACLE, which evaluates the native optimizer's realized step through a joint endpoint linearization of heterogeneous constraint families, constructs the resulting alignment in the optimizer's own geometry, bounds its authority, and commits it only after validation. We evaluate ORACLE across eight Partial Differential Equation benchmarks and four optimizers spanning Euclidean, diagonal adaptive, and structured preconditioned geometries, where it improves or matches native optimizer in 94% of configurations. Cross model analysis shows the same behavior in 92% of configurations, while matched comparisons show improvements over alternative constraint-handling methods acting at the objective, gradient, and post-optimizer levels.
We propose Dynamic Constraint Learning (DCL), a data-driven framework for constrained optimization when constraint functions are unknown and cannot be queried during optimization. At each iteration, the method learns a local surrogate from nearby data and solves a subproblem within a data-supported trust region. Compared with offline global constraint learning, the approach uses local surrogates that adapt to the data distribution during optimization and can achieve solution quality comparable to that of global models while using simpler local models and smaller optimization subproblems. We demonstrate the performance of DCL on a synthetic test problem and two case studies from the literature.
Ezgi Oztekin, Figen Oztoprak, S. Ilker Birbil
Department of Mathematics, Gebze Technical University · Department of Industrial Engineering, Gebze Technical University · Amsterdam Business School, University of Amsterdam
Modern LLM optimizers such as Muon often produce weight matrices with higher effective rank than Adam, yet further spectral control has delivered only modest gains. We identify a tension behind this result: concentrated spectra can suppress gradient directions in coupled weight matrices and slow optimization, while constraints maintained throughout training can limit task-specific adaptation and raise the attainable loss floor. We introduce ORCA (Orthogonal Regularization, Cooled After), a minimal optimizer intervention that applies strong but temporary soft orthogonality regularization early in training, then removes it. This allows the weights to benefit from a broader spectrum early on and adapt freely afterward. Across LLaMA, Qwen3, and fine-grained mixture-of-experts models ranging from 130M to 8B parameters, ORCA achieves lower final validation loss than Muon. Its loss reduction relative to Muon matches or exceeds Muon's reduction relative to Adam. Ablations support the early-shaping, later-release design. Further, ORCA requires no architectural changes and adds minimal overhead.
Yuanshi Liu, Boyuan Jiang, Liang Hou +4
School of Intelligence Science and Technology, Peking University · Kling Team
Constraint handling is central to constrained black-box optimization (BBO), where objective improvement and feasibility restoration often provide conflicting search signals. Existing ε-relaxation methods are simple and effective, but their relaxation schedules are usually fixed or manually designed for a limited range of problems. To address this limitation, this letter proposes MeCO, a meta-learning-assisted optimizer that learns an adaptive ε-relaxation policy for constrained BBO. MeCO couples a SHADE optimizer with a Double Deep Q-Network controller. At each optimization step, the controller observes compact population and constraint features and selects a scalar action, which is decoded into a relaxation vector for the candidate comparison rule. The policy is trained across constrained BBO instances and then deployed on held-out problems without problem-specific tuning. Experiments on the CEC2017 constrained benchmark, 16 UAV path-planning tasks and eight real-world engineering problems provide evidence that MeCO transfers across held-out benchmark functions, higher dimensions, and an application-domain setting. Ablation and behavior analyses further clarify the roles of constraint-related state features, action scaling, reward shaping, and meta-training.
Sijie Ma, Zeyuan Ma, Yue-Jiao Gong +1
Hong Kong Polytechnic University · South China Normal University · South China University of Technology