Adam and its variants dominate neural network training, but a single run only reveals whether a configuration works well after most of its budget is spent, a poor fit for hyperparameter or architecture search, where configurations must be ranked cheaply and pruned early. We introduce EvE (Evolutionary Explorer), a steady-state, population-of-four differential evolution (DE) optimizer with a targeted Adam fallback: each iteration proposes one candidate via DE, running a short burst of gradient descent only if the DE step fails to improve on the incumbent. Selection is greedy, so on a deterministic objective the best-so-far value is provably monotone non-increasing, and since gradients are used only as a targeted rescue, per-iteration cost stays within a constant factor of a single Adam step regardless of dimension. Under a fixed, evaluation-cost-matched budget, EvE wins or ties Adam on 76% of 70 (problem, dimension) cells across seven scalable benchmarks up to one million variables. On three real neural-network tasks (an MLP on MNIST, and LoRA fine-tuning of a 1.5B-parameter language model on two datasets) EvE finishes the same charged budget 1.7-3.9x faster, at a modest cost in final quality (about one accuracy point on MNIST, 9-11% higher relative test loss on the two fine-tuning tasks; on GSM8K, Adam is about 5 accuracy points more accurate, and fine-tuning lowers accuracy below the base model for both). Inside successive halving on UCI Adult, EvE completes hyperparameter and architecture searches 3.1-3.5x faster, ranking configurations about as consistently with Adam as Adam does with itself across seeds (Kendall's tau 0.66-0.69). EvE is not a total replacement for Adam as a final-stage trainer, but a fast, gradient-aware proxy for the search-heavy, budget-constrained regime one level up.
Figures & tables
Table 1: EvE vs. Adam on the scalable synthetic benchmarks, at fixed evaluation-cost budget (Section 5 ). Lower is better. Min/Med/Max over 21 seeds per (problem, n ) pair. A shaded, bold median is the significantly lower (better) side (one-sided Wilcoxon signed-rank, Holm-Bonferroni corrected, family-wise α=0.05 ); shaded, bold, italic on both sides marks no significant difference.
Figure 1: EvE vs. eight baselines on MLP/MNIST, under the same evaluation-cost budget (Section 5.1 ), at each method’s own best learning rate over 5 seeds. (a) training runs vs. budget and vs. wall-clock time; see Figure 2 for panels (b)–(d).
Figure 2: (cont.) (b) learning-rate sensitivity. (c) wall-clock cost of the shared budget, sorted. (d) final test accuracy reached.
Figure 3: SHA sweep wall-clock time vs. best validation accuracy found, UCI Adult, mean ± std over 5 seeds. EvE reaches comparable accuracy in roughly a third of the wall-clock time of every baseline in both search spaces.
Figure 4: EvE vs. eight baselines fine-tuning Qwen2.5-1.5B-Instruct with LoRA on Alpaca, under the same evaluation-cost budget (Section 5.1 ), at each method’s own best learning rate over 5 seeds. (a)/(b) training runs vs. budget and vs. wall-clock time. (c) learning-rate sensitivity. (d) wall-clock cost of the shared budget, sorted. (e) final test loss reached (lower is better).
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Measured peak GPU memory, EvE vs. Adam, on MLP/MNIST (Section 5.3 ), 200 real training iterations per method, each in its own CUDA context. The measured ratio ( 2.95× peak allocated) sits below the parameter-storage-only theoretical ratio ( 3.25× ) derived above.
Task
Methods with same rate
EvE rate (test / val.)
EvE result (test / val.)
Largest baseline change
MLP/MNIST
5 of 9
10−3 / 10−3
97.03% / 97.03%
0.12 points
LoRA/Alpaca
7 of 9
3×10−4 / 3×10−4
1.240 / 1.240
0.002
LoRA/GSM8K
7 of 9
10−2 / 10−4
0.640 / 0.654
0.0004
Appendix
Table 2: Effect of choosing each method’s learning rate on validation instead of test data, from the same saved runs.
Task
EvE (s)
Baselines (s)
Speedup
MLP/MNIST ( B=3,000 )
49.3
182.3 – 192.3
3.7 – 3.9×
LoRA/Alpaca ( B=300 )
35.0±0.6
58.8 – 70.7
1.7 – 2.0×
LoRA/GSM8K ( B=900 )
100.4±4.7
193.8 – 198.7
1.9 – 2.0×
UCI Adult, hyperparameter search
21.8±0.3
73.6 – 76.8
3.4 – 3.5×
UCI Adult, architecture search
21.5±0.9
66.9 – 75.8
3.1 – 3.5×
Appendix
Table 3: Wall-clock training time in seconds under the shared budget, EvE against the eight baselines (range).
Setting
Value
Evaluation-cost budget B
2,000 (Section 5.1 accounting)
Problem dimensions n
{2,10,20,50,100,200,500,1,000,105,106}
Seeds
{0,…,20} (21 independent runs per cell)
EvE population size npop
4 (fixed, Section 3.1 )
Population initialization
Latin Hypercube Sampling (see below), not proximity (Section 3.2 )
DE weight F
0.5 , jitter γ=10−4 (Section 3.3 )
Appendix
Table 4: Shared parameters for the synthetic scalable-benchmark sweep (Table 1 ).
Problem
EvE better
Indistinguishable
Adam better
Where Adam wins
Ackley
7
3
0
–
Griewank
10
0
0
–
Rastrigin
7
3
0
–
Rosenbrock
0
10
0
–
Schwefel
1
2
7
n≥50
Sphere
0
0
10
all n
Appendix
Table 5: Per-problem summary of EvE vs. Adam on the synthetic benchmarks: over the ten dimensions swept, the number of (problem, n ) cells where EvE is significantly better, where the two are statistically indistinguishable, and where Adam is significantly better (one-sided Wilcoxon signed-rank, paired by seed, Holm-Bonferroni corrected, family-wise α=0.05 ; 21 seeds per cell). The full min/median/max at every cell is Table 1 in the main text.
Layer
Shape
Activation
Parameters
Input
784
–
–
Hidden 1
784→1,024
ReLU
802,816+1,024 (weight+bias)
Hidden 2
1,024→1,024
ReLU
1,048,576+1,024
Output
1,024→10
–
10,240+10
Total
1,863,690
Appendix
Table 6: MLP architecture. fan_in is the preceding layer’s width; biases inherit their weight tensor’s fan-in.
Setting
Value
Evaluation-cost budget B
3,000 (Section 5.1 accounting)
Learning-rate grid
{10−4,3×10−4,10−3,3×10−3,10−2}
Seeds
{0,1,2,3,4} (5 independent runs per cell)
Minibatch size
256
Weight decay
5×10−4 (decoupled for AdamW; L2 for SGD+momentum and EvE’s Adam fallback)
SGD momentum
0.9
Appendix
Table 7: Shared MLP/MNIST training protocol.
Method
Time (s)
Best Val Acc
EvE
21.8 ± 0.3
0.8544 ± 0.0018
SGD+mom.
73.6 ± 1.3
0.8549 ± 0.0020
NAdam
75.1 ± 0.5
0.8591 ± 0.0018
RAdam
75.5 ± 0.9
0.8584 ± 0.0009
AdaBelief
75.6 ± 0.8
0.8579 ± 0.0010
Adam
75.7 ± 0.9
0.8582 ± 0.0019
Appendix
Table 8: UCI Adult, learning-rate/weight-decay SHA sweep, mean ± std over 5 seeds, sorted by wall-clock time.
Method
Time (s)
Best Val Acc
EvE
21.5 ± 0.9
0.8544 ± 0.0010
AdaBelief
66.9 ± 1.2
0.8587 ± 0.0017
Lookahead
68.2 ± 1.1
0.8588 ± 0.0006
SGD+mom.
69.6 ± 0.9
0.8533 ± 0.0014
AMSGrad
69.8 ± 2.1
0.8588 ± 0.0017
RAdam
70.5 ± 1.7
0.8585 ± 0.0012
Appendix
Table 9: UCI Adult, architecture (width/depth) SHA sweep, mean ± std over 5 seeds, sorted by wall-clock time.
Comparison
Hyperparameter search
Architecture search
EvE vs. Adam
0.663±0.093
0.691±0.032
EvE vs. AdamW
0.677±0.063
0.709±0.041
Adam vs. Adam (different training seeds)
0.694±0.061
0.670±0.029
Adam vs. AdamW
0.714±0.078
0.848±0.018
Appendix
Table 10: Rank agreement (Kendall’s τ ) at the first SHA rung over the same 48 configurations, mean ± std over 5 seeds.
Setting
Value
Evaluation-cost budget B
300 (Section 5.1 accounting)
Learning-rate grid
{10−4,3×10−4,10−3,3×10−3,10−2}
Seeds
{0,1,2,3,4} (5 independent runs per cell)
Minibatch size
2
Max sequence length
512 tokens
LoRA rank / alpha
32 / 64
Appendix
Table 11: Shared LoRA/Alpaca training protocol ( B=300 ; GSM8K uses the same protocol except B=900 , Appendix F ).
Figure 6: EvE vs. eight baselines fine-tuning Qwen2.5-1.5B-Instruct with LoRA directly on GSM8K, under a shared evaluation-cost budget of 900 ( 3× Figure 4 ’s Alpaca budget), at each method’s own best learning rate over 5 seeds. (a)/(b) training runs vs. budget and vs. wall-clock time. (c) learning-rate sensitivity. (d) wall-clock cost of the shared budget, sorted. (e) final test loss reached (lower is better).
Checkpoint
Strict ( "#### N" )
Flexible
Flexible interval
Training time (s)
Base model (no fine-tuning)
27.9%
73.5%
[71.1,75.9]
–
EvE, 900 / 30,000 evaluations
54.4%
54.4%
[51.7,57.1]
105 / 3,396
EvE, 10,000 evaluations
53.7%
53.7%
[51.0,56.4]
1,104
Adam, 900 evaluations
59.4%
59.4%
[56.7,62.0]
196
Adam, 10,000 evaluations
58.2%
58.2%
[55.5,60.9]
2,093
Adam, 30,000 evaluations
52.5%
52.6%
[49.9,55.3]
6,284
Appendix
Table 12: GSM8K downstream exact-match accuracy on the full 1,319 -question test set, one training run per row, with the wall-clock training time of that run in seconds. The interval is the 95% Wilson interval on flexible accuracy. EvE’s 30,000 -evaluation adapter is identical to its 900 -evaluation adapter, so its accuracy is shown once, with both training times.
We introduce AdamX, a first-order optimizer that incorporates cosine similarity as an adaptive mechanism for controlling update magnitudes. The proposed method is scalable, model-agnostic, and straightforward to integrate into existing training pipelines. We further introduce a variance rectification scheme that promotes smoother optimization during the early stages of training. Overall, we provide empirical evidence that AdamX achieves competitive convergence rates across a range of benchmark datasets and architectures. Performance is evaluated in terms of the number of epochs required to reach predefined performance thresholds under a fixed hyperparameter budget. Code and Experiments available at: https://github.com/FranciscoCaldas/adamX.
Francisco Caldas, Ruben Belo, Cláudia Soares
NOVA School of Science of Technology Universidade Nova de Lisboa, Caparica, Portugal
Would experience designing faster GPU kernels also help close in on a long-standing open mathematical conjecture? Large Language Models (LLMs) integrated into evolutionary search have recently produced state-of-the-art solutions on optimization tasks, including open mathematical conjectures, GPU kernel design, scientific law discovery, and combinatorial puzzles. To achieve this, prior work applied search scaffolds to one target task at a time, so every new problem is approached from scratch and the experience accumulated during search is discarded once the model finishes its attempt. This leaves the capability of iteratively evolving a solution (e.g., knowing which part to mutate and how, deciding when to backtrack) entirely in the scaffold rather than in the model itself. Whether the model itself could acquire this capability and reuse it across different tasks has been largely unexamined. To address this, we introduce Evolution Fine-Tuning (EFT), a mid-training paradigm that teaches LLMs to evolve solutions across tasks by converting evolutionary search trajectories into supervision. We construct Finch Collection, a 156K-trajectory dataset spanning 10 domains and 371 optimization tasks, and fine-tune open-source LLMs from 2B to 9B parameters. Empirically, EFT confers cross-task generalization: across 22 held-out tasks, our models surpass their base counterparts by 10.22% on average. Furthermore, when paired with test-time RL, our model matches state-of-the-art performance on two circle-packing tasks and outperforms its base-model counterpart on the Erdős minimum-overlap problem. EFT thus serves as a "practice phase" for general-purpose discovery agents that do not solve new problems from scratch.
Young-Jun Lee, Seungone Kim, Minki Kang +5
University of Minnesota · Carnegie Mellon University · KAIST +3
Interpreting optimizers as gradient-flow discretizations has motivated applying higher-order Runge-Kutta (RK) integrators to neural networks. We build a representative Adam variant (Bogacki-Shampine 3(2) RK pair, FSAL reuse, local-error step control) and evaluate it under a strict compute-matched protocol giving every method the same gradient-evaluation budget - an accounting this literature rarely enforces. Under it the RK variant loses to plain Adam on training loss in both minibatch and full-batch (RK's best-case) training. Instrumenting it shows the "adaptivity" is illusory: normalized error stays far below tolerance, the step size pins at its growth cap from step one (98-100 percent of steps), and no rtol x hmax x h0 setting makes it act; tolerances spanning 100x give bit-identical trajectories. The method is exactly fixed-step Adam with an averaged gradient at 3-4x cost. Repairing it (true reject branch; error on the applied map) reverses the full-batch result - about 40x lower training loss than tuned Adam - and a fixed-step control isolates adaptivity (an emergent warmup-and-growth schedule) as the mechanism. But the gain is fragile to the initial step size and does not reach test accuracy. A pre-registered follow-up rules out the obvious explanations: deeper minimization does not overfit, and an explicit temperature knob only hurts - leaving a trajectory effect, the controller selecting a minimum generalizing 1.3-3.4 points below first-order descent at equal depth. An n=10 study confirms one secondary effect: gradient averaging is a genuine implicit regularizer, beating lr-matched Adam and AdamW on 10/10 seeds - yet RMSprop and NAdam match or beat it at a third the per-step cost. Higher-order adaptive integration buys deeper deterministic minimization and a small regularization effect, but nothing a cheaper, well-tuned first-order baseline does not already provide.