EvE: An Alternate Optimizer to Adam
Organizations: Department of Computer Science The University of Alabama Tuscaloosa, AL 35487, USA
Abstract
Adam and its variants dominate neural network training, but a single run only reveals whether a configuration works well after most of its budget is spent, a poor fit for hyperparameter or architecture search, where configurations must be ranked cheaply and pruned early. We introduce EvE (Evolutionary Explorer), a steady-state, population-of-four differential evolution (DE) optimizer with a targeted Adam fallback: each iteration proposes one candidate via DE, running a short burst of gradient descent only if the DE step fails to improve on the incumbent. Selection is greedy, so on a deterministic objective the best-so-far value is provably monotone non-increasing, and since gradients are used only as a targeted rescue, per-iteration cost stays within a constant factor of a single Adam step regardless of dimension. Under a fixed, evaluation-cost-matched budget, EvE wins or ties Adam on 76% of 70 (problem, dimension) cells across seven scalable benchmarks up to one million variables. On three real neural-network tasks (an MLP on MNIST, and LoRA fine-tuning of a 1.5B-parameter language model on two datasets) EvE finishes the same charged budget 1.7-3.9x faster, at a modest cost in final quality (about one accuracy point on MNIST, 9-11% higher relative test loss on the two fine-tuning tasks; on GSM8K, Adam is about 5 accuracy points more accurate, and fine-tuning lowers accuracy below the base model for both). Inside successive halving on UCI Adult, EvE completes hyperparameter and architecture searches 3.1-3.5x faster, ranking configurations about as consistently with Adam as Adam does with itself across seeds (Kendall's tau 0.66-0.69). EvE is not a total replacement for Adam as a final-stage trainer, but a fast, gradient-aware proxy for the search-heavy, budget-constrained regime one level up.
Figures & tables
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Task | Methods with same rate | EvE rate (test / val.) | EvE result (test / val.) | Largest baseline change |
|---|---|---|---|---|
| MLP/MNIST | of | / | / | points |
| LoRA/Alpaca | of | / | / | |
| LoRA/GSM8K | of | / | / |
| Task | EvE (s) | Baselines (s) | Speedup |
|---|---|---|---|
| MLP/MNIST ( ) | – | – | |
| LoRA/Alpaca ( ) | – | – | |
| LoRA/GSM8K ( ) | – | – | |
| UCI Adult, hyperparameter search | – | – | |
| UCI Adult, architecture search | – | – |
| Setting | Value |
|---|---|
| Evaluation-cost budget | (Section 5.1 accounting) |
| Problem dimensions | |
| Seeds | (21 independent runs per cell) |
| EvE population size | (fixed, Section 3.1 ) |
| Population initialization | Latin Hypercube Sampling (see below), not proximity (Section 3.2 ) |
| DE weight | , jitter (Section 3.3 ) |
| Problem | EvE better | Indistinguishable | Adam better | Where Adam wins |
|---|---|---|---|---|
| Ackley | 7 | 3 | 0 | – |
| Griewank | 10 | 0 | 0 | – |
| Rastrigin | 7 | 3 | 0 | – |
| Rosenbrock | 0 | 10 | 0 | – |
| Schwefel | 1 | 2 | 7 | |
| Sphere | 0 | 0 | 10 | all |
| Layer | Shape | Activation | Parameters |
|---|---|---|---|
| Input | – | – | |
| Hidden 1 | ReLU | (weight+bias) | |
| Hidden 2 | ReLU | ||
| Output | – | ||
| Total |
| Setting | Value |
|---|---|
| Evaluation-cost budget | (Section 5.1 accounting) |
| Learning-rate grid | |
| Seeds | (5 independent runs per cell) |
| Minibatch size | |
| Weight decay | (decoupled for AdamW; L2 for SGD+momentum and EvE’s Adam fallback) |
| SGD momentum |
| Method | Time (s) | Best Val Acc |
|---|---|---|
| EvE | 21.8 0.3 | 0.8544 0.0018 |
| SGD+mom. | 73.6 1.3 | 0.8549 0.0020 |
| NAdam | 75.1 0.5 | 0.8591 0.0018 |
| RAdam | 75.5 0.9 | 0.8584 0.0009 |
| AdaBelief | 75.6 0.8 | 0.8579 0.0010 |
| Adam | 75.7 0.9 | 0.8582 0.0019 |
| Method | Time (s) | Best Val Acc |
|---|---|---|
| EvE | 21.5 0.9 | 0.8544 0.0010 |
| AdaBelief | 66.9 1.2 | 0.8587 0.0017 |
| Lookahead | 68.2 1.1 | 0.8588 0.0006 |
| SGD+mom. | 69.6 0.9 | 0.8533 0.0014 |
| AMSGrad | 69.8 2.1 | 0.8588 0.0017 |
| RAdam | 70.5 1.7 | 0.8585 0.0012 |
| Comparison | Hyperparameter search | Architecture search |
|---|---|---|
| EvE vs. Adam | ||
| EvE vs. AdamW | ||
| Adam vs. Adam (different training seeds) | ||
| Adam vs. AdamW |
| Setting | Value |
|---|---|
| Evaluation-cost budget | (Section 5.1 accounting) |
| Learning-rate grid | |
| Seeds | (5 independent runs per cell) |
| Minibatch size | |
| Max sequence length | tokens |
| LoRA rank / alpha | / |
| Checkpoint | Strict ( "#### N" ) | Flexible | Flexible interval | Training time (s) |
|---|---|---|---|---|
| Base model (no fine-tuning) | – | |||
| EvE, / evaluations | / | |||
| EvE, evaluations | ||||
| Adam, evaluations | ||||
| Adam, evaluations | ||||
| Adam, evaluations |