Early Memory Selection for Balanced Adam
Organizations: Universitat Politècnica de València Valencia, Spain · Universitat Jaume I Castelló de la Plana, Spain
Abstract
We propose a method for choosing the shared memory parameter in Adam from a short pilot training. The selected remains fixed during the subsequent full training. A local model of Adam's normalized direction balances sampling variability against the delay introduced by averaging past gradients. This balance gives a cubic memory rule, whose two coefficients are estimated from gradient probes at a few pilot checkpoints. The estimator uses the numerator and denominator jointly, preserving their covariance. With a 200-update pilot and sixteen probe gradients at each of four checkpoints, a seed-matched retrospective evaluation on eleven vision and language workloads reduces mean relative validation gap by 40.7% and worst-quarter mean gap by 44.3% against the grid representative of shared . The mean gap is also 32.3% lower than that of the best constant chosen across all eleven workloads.
Figures & tables
| 1 | Run pilot updates with shared . At and , collect probe gradients with parameters fixed. |
| 2 | Compute pooled and group moment estimates, normalize them using ( 13 ), and store their images under . |
| 3 | Estimate variability using ( 14 ); fit ( 17 ) and compute change using ( 18 ). |
| 4 | Apply ( 21 ), restrict memory to the candidate range, convert to , and round to . |
| 5 | Restart from the pilot’s initialization and use this fixed shared with the prescribed full-training recipe. |
| Network | Dataset | Oracle | Selected | Relative gap (%) |
|---|---|---|---|---|
| EfficientNet-B0 | Cars | 0.82217 | 0.96838 | 1.0124 |
| Llama60M | C4 | 0.94377 | 0.90000 | 0.2582 |
| Llama60M | SlimPajama | 0.94377 | 0.90000 | 0.1175 |
| NanoGPT | OpenWebText | 0.94377 | 0.90000 | 0.1157 |
| NanoGPT | WikiText-103 | 0.94377 | 0.90000 | 0.1613 |
| ResNet50 | Food-101 | 0.99684 | 0.96838 | 1.0613 |
| Mean | Maximum | CVaR 25 | |
|---|---|---|---|
| Fixed 0.90000 | 1.4292 | 3.8770 | 3.3543 |
| Fixed 0.94377 | 0.6236 | 2.4853 | 2.0409 |
| Best fixed 0.96838 | 0.5461 | 2.1862 | 1.4304 |
| Tracking rule | 0.3695 | 1.3339 | 1.1359 |
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Network | Dataset | Steps | Batch | Accum. | LR max | Decay | Clip |
|---|---|---|---|---|---|---|---|
| EfficientNet-B0 | Cars | 12800 | 64 | 1 | 0.0008 | 5e-05 | 1 |
| Llama60M | C4 | 10000 | 64 | 8 | 0.001 | 1e-05 | 0 |
| Llama60M | SlimPajama | 10000 | 64 | 8 | 0.001 | 1e-05 | 0 |
| NanoGPT | OpenWebText | 10000 | 8 | 128 | 0.0006 | 0.01 | 1 |
| NanoGPT | WikiText-103 | 10000 | 8 | 128 | 0.0006 | 0.01 | 1 |
| ResNet50 | Food-101 | 29550 | 128 | 1 | 0.0003 | 0.01 | 0 |
| Asset | License information or provider record |
|---|---|
| PyTorch / torchvision | BSD-style / BSD-3-Clause: PyTorch license and torchvision license . |
| NumPy / Transformers | BSD-3-Clause / Apache-2.0: NumPy license and Transformers license . |
| T5-small checkpoint | Apache-2.0, as stated in the provider model card . The other evaluated models start from random weights. |
| C4 / WikiText-103 | C4 card: ODC-By . WikiText card: CC-BY-SA-3.0 and GFDL . |
| SlimPajama-6B | The subset card points to the original SlimPajama record for source and license information. This original provider record supplies the access provenance used here. |
| OpenWebText / BookCorpus | Provider records: OpenWebText and BookCorpus . Rights in underlying web text and books remain with their respective sources. |
| Decision update | Mean gap | Maximum gap | CVaR 25 |
|---|---|---|---|
| 100 | 0.6345 | 2.1862 | 1.5271 |
| 150 | 0.3695 | 1.3339 | 1.1359 |
| 200 | 0.3695 | 1.3339 | 1.1359 |
| 250 | 0.3695 | 1.3339 | 1.1359 |
| Effective batches | Groups | Mean gap | Maximum gap | CVaR 25 |
|---|---|---|---|---|
| 2 | 2 | 0.9185 | 2.5913 | 2.4600 |
| 4 | 4 | 0.6829 | 2.4853 | 2.0409 |
| 8 | 4 | 0.6829 | 2.4853 | 2.0409 |
| 16 | 4 | 0.3695 | 1.3339 | 1.1359 |
| Fixed shared or method | Mean gap | Maximum gap | CVaR 25 |
|---|---|---|---|
| 0.68377 | 4.8152 | 11.6651 | 9.9106 |
| 0.82217 | 2.0501 | 6.4732 | 4.5653 |
| 0.90000 | 1.4292 | 3.8770 | 3.3543 |
| 0.94377 | 0.6236 | 2.4853 | 2.0409 |
| 0.96838 | 0.5461 | 2.1862 | 1.4304 |
| 0.98222 | 1.8128 | 7.2826 | 4.5072 |
| Gap benefit (pp) | Loss improvement (%) | ||||||
|---|---|---|---|---|---|---|---|
| Network | Dataset | Rule gap | .94377 | .96838 | .94377 | .96838 | |
| EfficientNet-B0 | Cars | 0.96838 | 1.0124 | -0.2853 | +0.0000 | -0.2833 | +0.0000 |
| Llama60M | C4 | 0.90000 | 0.2582 | -0.2582 | -0.0700 | -0.2582 | -0.0699 |
| Llama60M | SlimPajama | 0.90000 | 0.1175 | -0.1175 | -0.0924 | -0.1175 | -0.0924 |
| NanoGPT | OpenWebText | 0.90000 | 0.1157 | -0.1157 | -0.1137 | -0.1157 | -0.1137 |
| NanoGPT | WikiText | 0.90000 | 0.1613 | -0.1613 | -0.0800 | -0.1613 | -0.0800 |
| Excluded task | Mean-gap reduction (%) | Best fixed on ten | ||||
|---|---|---|---|---|---|---|
| Network | Dataset | Rule mean | vs .94377 | vs .96838 | Reduction (%) | |
| EfficientNet-B0 | Cars | 0.3052 | +50.23 | +38.91 | 0.96838 | +38.91 |
| Llama60M | C4 | 0.3806 | +44.51 | +34.60 | 0.96838 | +34.60 |
| Llama60M | SlimPajama | 0.3947 | +42.46 | +34.03 | 0.96838 | +34.03 |
| NanoGPT | OpenWebText | 0.3948 | +42.44 | +34.25 | 0.96838 | +34.25 |
| NanoGPT | WikiText | 0.3903 | +43.10 | +34.14 | 0.96838 | +34.14 |
| Network | Dataset | Continuous | Grid | |||
|---|---|---|---|---|---|---|
| EfficientNet-B0 | Cars | 27.494 | 0.963628 | 0.96838 | ||
| Llama60M | C4 | 8.823 | 0.886654 | 0.90000 | ||
| Llama60M | SlimPajama | 8.361 | 0.880399 | 0.90000 | ||
| NanoGPT | OpenWebText | 7.678 | 0.869751 | 0.90000 | ||
| NanoGPT | WikiText-103 | 8.871 | 0.887272 | 0.90000 | ||
| ResNet50 | Food-101 | 26.458 | 0.962205 | 0.96838 |
| Network | Dataset | Continuous | Nearest boundary | distance | Ratio change (%) |
|---|---|---|---|---|---|
| EfficientNet-B0 | Cars | 0.963628 | 0.956075 | 0.007553 | -43.224 |
| Llama60M | C4 | 0.886654 | 0.861085 | 0.025569 | -45.679 |
| Llama60M | SlimPajama | 0.880399 | 0.861085 | 0.019314 | -36.180 |
| NanoGPT | OpenWebText | 0.869751 | 0.861085 | 0.008666 | -17.571 |
| NanoGPT | WikiText-103 | 0.887272 | 0.861085 | 0.026187 | -46.563 |
| ResNet50 | Food-101 | 0.962205 | 0.956075 | 0.006130 | -36.294 |
| Network | Dataset | Adjacent | ratio (%) | gap (pp) | loss (%) | |
|---|---|---|---|---|---|---|
| EfficientNet-B0 | Cars | 0.96838 | 0.94377 | -43.22 | -0.2853 | -0.2825 |
| Llama60M | C4 | 0.90000 | 0.82217 | -45.68 | +0.6477 | +0.6460 |
| Llama60M | SlimPajama | 0.90000 | 0.82217 | -36.18 | +0.7058 | +0.7050 |
| NanoGPT | OpenWebText | 0.90000 | 0.82217 | -17.57 | +0.4311 | +0.4306 |
| NanoGPT | WikiText | 0.90000 | 0.82217 | -46.56 | +0.2984 | +0.2979 |
| ResNet50 | Food-101 | 0.96838 | 0.94377 | -36.29 | +1.4240 | +1.4091 |
| Network | Dataset | Pooling contribution (%) | Finite-lag contribution (%) |
|---|---|---|---|
| EfficientNet-B0 | Cars | 3.12 | 100.00 |
| Llama60M | C4 | 4.45 | 54.57 |
| Llama60M | SlimPajama | 4.63 | 91.79 |
| NanoGPT | OpenWebText | 3.17 | 100.00 |
| NanoGPT | WikiText-103 | 3.34 | 100.00 |
| ResNet50 | Food-101 | 3.09 | 100.00 |
| Coefficient diagnostic | Mean gap | Maximum gap | CVaR 25 |
|---|---|---|---|
| Constant-target calculation | 0.5461 | 2.1862 | 1.4304 |
| Endpoint contribution only | 0.5682 | 2.1862 | 1.5271 |
| Intercept contribution only | 17.9403 | 72.4969 | 51.6408 |
| Full measured rule | 0.3695 | 1.3339 | 1.1359 |
| Mean relative gap (%) | Direct improvement (%) | |||||
|---|---|---|---|---|---|---|
| Summary | Rule | Fixed .94377 | Best fixed | Fixed | Seed 1 | Common seeds |
| Minimum recorded | 0.3695 | 0.6236 | 0.5461 | 0.96838 | +0.2464 | -0.1049 |
| Last recorded | 0.4537 | 0.6242 | 0.5707 | 0.96838 | +0.1629 | -0.1093 |
| Mean of last three | 0.4244 | 0.4767 | 0.4767 | 0.94377 | +0.0492 | -0.1586 |
| Network | Dataset | Common seeds | Improvement (%) | Rule |
|---|---|---|---|---|
| EfficientNet-B0 | Cars | 1,2,3 | -2.1064 | 0.96838 |
| Llama60M | C4 | 1,2,3 | -0.2015 | 0.90000 |
| Llama60M | SlimPajama | 1,2,3 | -0.2985 | 0.90000 |
| NanoGPT | OpenWebText | 1,2,3 | -0.2046 | 0.90000 |
| NanoGPT | WikiText-103 | 1,2,3 | -0.1235 | 0.90000 |
| ResNet50 | Food-101 | 1 | 1.3895 | 0.96838 |