The sharpness-aware minimization (SAM) algorithm and its variants, including gap guided SAM (GSAM), have been successful at improving the generalization capability of deep neural network models by finding flat local minima of the empirical loss in training. Meanwhile, it has been shown theoretically and practically that increasing the batch size or decaying the learning rate avoids sharp local minima of the empirical loss. In this paper, we consider the GSAM algorithm with increasing batch sizes or decaying learning rates, such as cosine annealing or linear learning rate, and theoretically show its convergence. Moreover, we numerically compare SAM (GSAM) with and without an increasing batch size and conclude that using an increasing batch size { achieves a lower worst-case ℓ∞ adaptive sharpness} than compared with using a constant batch size and learning rate.
Figures & tables
Algorithm
Gradient
Leaning Rate
Perturbation
Convergence Analysis
(1) SAM
Mini-batch b
ηT=Θ(T1/21)
ρT=Θ(T1/41)
E[∥∇fS∗∥]=O(T1/41+bT1/41)
(2) SSAM
Noise
ηt=Θ(t1/21)
ρt=Θ(t1/21)
E[∥∇fS∗∥]=O(T1/4logT)
(3) GSAM
Noise
ηt=Θ(t1/21)
ρt=Θ(t1/21)
E[∥∇f^S,ρtSAM∗∥]=O(T1/4logT)
(4) m -SAM
Noise
ηT=O(T1/21)
ρ
E[∥∇fS∗∥]=O(T1/21+ρ2)
(5) VaSSO
Noise
ηT=Θ(T1/21)
ρT=Θ(T1/21)
E[∥∇f^S,ρSAM∗∥]=O(T1/41)
(6) FSAM
Noise
ηT=Θ(T1/21)
ρt=Θ(t1/21)
E[∥∇fS∗∥]=O(T1/4logT)
Table 1: Convergence of SAM and its variants to minimize f^S,ρSAM(x)=fS(x)+ρ∥∇fS(x)∥ over the number of steps T . “Noise” in the Gradient column means that algorithm uses noisy observation, i.e., g(x)=∇f(x)+(Noise) , of the full gradient ∇f(x) , while “Mini-batch” in the Gradient column means that algorithm uses a mini-batch gradient ∇fB(x)=b1∑i∈[b]∇fξi(x) with a batch size b . Here, we let E[∥∇f^S,ρSAM∗∥]:=mint∈[T]E[∥∇f^S,ρSAM(xt)∥] , where (xt)t=0T is the sequence generated by Algorithm. Results (1)–(6) were presented in (1) ( Andriushchenko and Flammarion, 2022 , Theorem 2) , (2) ( Mi et al., 2022 , Theorem 2) , (3) ( Zhuang et al., 2022 , Theorem 5.1) , (4) ( Si and Yun, 2023 , Theorem 4.6) , (5) ( Li and Giannakis, 2023 , Corollary 1) , and (6) ( Li et al., 2024 , Theorem 2) . † See Appendix C for the derivation of the rates and the conditions on ϵ , ρ , and b .
Dataset / Model
Metric
SGD
SAM
GSAM
SGD + B
SAM + B
GSAM + B
SGD + C
SAM + C
GSAM + C
CIFAR-10
Test Error
7.17
6.34
6.21
6.39
6.03
6.06
6.63
6.08
6.2
/ ResNet-18
Sharpness
106.46
51.82
48.39
8.73
4.27
4.08
141.37
88.53
89.3
CIFAR-10
Test Error
7.04
6.13
5.98
5.53
5.10
5.05
6.67
5.87
5.92
/ WRN-28-10
Sharpness
1124.45
546.93
543.27
48.25
38.79
40.28
1467.67
771.79
794.45
CIFAR-100
Test Error
26.61
26.39
26.61
25.58
25.10
25.18
26.63
25.87
26.12
/ ResNet-18
Sharpness
154.27
46.23
47.55
1.33
0.94
0.90
155.88
72.70
71.86
Table 2: Mean value of the test error (Test Error) and worst-case ℓ∞ adaptive sharpness (Sharpness) on the CIFAR-10/100 and Tiny-ImageNet datasets. “(algorithm) + B” refers to “(algorithm) + increasing_batch” and “(algorithm) + C” refers to “(algorithm) + Cosine”.
Dataset / Model
Metric
SGD
SAM
GSAM
SGD + B
SAM + B
GSAM + B
SGD + C
SAM + C
GSAM + C
ImageNet
Test Error
32.23
30.62
30.63
29.19
29.06
29.23
30.58
29.96
29.72
/ ResNet-50
Sharpness
492.12
268.60
297.82
55.31
41.83
38.13
1504.53
853.04
670.51
Table 3: Mean value of the test error (Test Error) and worst-case ℓ∞ adaptive sharpness (Sharpness) on the ImageNet dataset. “(algorithm) + B” refers to “(algorithm) + increasing_batch” and “(algorithm) + C” refers to “(algorithm) + Cosine”.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 1: (Left) Loss function value in training and (Right) accuracy score in testing for the algorithms versus the number of epochs in training ResNet-18 on the CIFAR10 dataset. The learning rate of each algorithm was fixed at 0.1. In SGD/SAM/GSAM, the batch size was fixed at 128. In SGD/SAM/GSAM + increasing_batch, the batch size was set at 16 for the first 40 epochs and then it was doubled every 40 epochs afterwards, i.e., to 32 for epochs 41-80, 64 for epochs 81-120, etc.
Figure 2: (Left) Loss function value in training and (Right) accuracy score in testing for the algorithms versus the number of epochs in training ResNet-18 on the CIFAR-10 dataset. The batch size of each algorithm was fixed at 128. In SGD/SAM/GSAM, the constant learning rate was fixed at 0.1. In SGD/SAM/GSAM + Cosine, the maximum learning rate was 0.1 and the minimum learning rate was 0.001.
Figure 3: (Left) Loss function value in training and (Right) accuracy score in testing for the algorithms versus the number of epochs in training Wide-ResNet-28-10 on the CIFAR-100 dataset. The learning rate of each algorithm was fixed at 0.1. In SGD/SAM/GSAM, the batch size was fixed at 128. In SGD/SAM/GSAM + increasing_batch, the batch size was set at 16 for the first 40 epochs and then it was doubled every 40 epochs afterwards, i.e., to 32 for epochs 41-80, 64 for epochs 81-120, etc.
Figure 4: (Left) Loss function value in training and (Right) accuracy score in testing for the algorithms versus the number of epochs in training Wide-ResNet-28-10 on the CIFAR-100 dataset. The batch size of each algorithm was fixed at 128. In SGD/SAM/GSAM, the constant learning rate was fixed at 0.1. In SGD/SAM/GSAM + Cosine, the maximum learning rate was 0.1 and the minimum learning rate was 0.001.
Figure 5: (Left) Loss function value in training and (Right) accuracy score in testing for the algorithms versus the number of epochs in training Wide-ResNet-28-10 on the Tiny-ImageNet dataset. The learning rate of each algorithm was fixed at 0.1. In SGD/SAM/GSAM, the batch size was fixed at 128. In SGD/SAM/GSAM + increasing_batch, the batch size was set at 16 for the first 25 epochs and then it was doubled every 25 epochs afterwards, i.e., to 32 for epochs 26-50, 64 for epochs 51-75, etc.
Figure 6: (Left) Loss function value in training and (Right) accuracy score in testing for the algorithms versus the number of epochs in training Wide-ResNet-28-10 on the Tiny-ImageNet dataset. The batch size of each algorithm was fixed at 128. In SGD/SAM/GSAM, the constant learning rate was fixed at 0.1. In SGD/SAM/GSAM + Cosine, the maximum learning rate was 0.1 and the minimum learning rate was 0.001.
Figure 7: (Left) Loss function value in training and (Right) accuracy score in testing for the batch sizes versus the number of steps in training ResNet-18 on the CIFAR-100 dataset. The learning rate for each batch size was fixed at 0.1. This is a comparison between the case of a varying batch size [16, 32, 64, 128, 256] (iteration: 242,120) and the case of a fixed batch size of 41 (iteration: 243,800).
batch_increasing
16
32
64
128
256
Test Error
25.58
27.14
27.28
27.11
26.62
27.06
Sharpness
1.34
0.93
4.28
46.22
154.27
329.86
Appendix
Table 4: Mean test error (Test Error) and worst-case ℓ∞ adaptive sharpness (Sharpness) when training ResNet-18 from scratch on CIFAR-100 with different batch sizes.
Figure 8: Test errors versus perturbation radius ρ when training ResNet-18 on CIFAR-10 using SAM with a constant batch size, 16 or 256, and an increasing batch size [16,32,64,128,256] .
Figure 9: Test errors versus perturbation radius ρ when training ResNet-50 on CIFAR-100 using SAM with constant batch size, 16 or 256, and an increasing batch size [16,32,64,128,256] .
Sharpness-Aware Minimization (SAM) has established itself as a powerful and widely adopted optimizer for training machine learning models. By explicitly minimizing the sharpness of the loss landscape, SAM often improves generalization while delivering strong empirical performance. However, SAM and its variants, like most training algorithms, are sensitive to the choice of learning rate, which is typically selected through extensive hyperparameter tuning or predefined schedulers. In this work, motivated by recent advances on the effectiveness of stochastic Polyak step sizes for Stochastic Gradient Descent (SGD), we derive Polyak schedulers tailored to SAM-style updates, yielding novel adaptive algorithms in both deterministic and stochastic settings. In the smooth setting, we prove linear convergence for strongly convex objectives and an O(1/T) convergence rate for convex objectives in the deterministic case. In the stochastic setting, we establish analogous convergence guarantees up to a neighborhood of the optimum. Numerical experiments demonstrate that the proposed Polyak schedulers achieve performance comparable to or better than carefully tuned SAM baselines, while substantially reducing the need for learning-rate tuning.
Dimitris Oikonomou, Nicolas Loizou
Mathematical Institute for Data Science (MINDS), Johns Hopkins University, Baltimore, MD, USA · Department of Computer Science, Johns Hopkins University, Baltimore, MD, USA · Department of Applied Mathematics and Statistics, Johns Hopkins University, Baltimore, MD, USA
Sharpness-Aware Minimization (SAM) improves generalization by seeking parameters whose loss is robust to local adversarial perturbations, but the quantitative mechanism underlying its implicit bias toward flat minima remains unclear. In particular, the perturbation radius ρ is typically treated as an isolated tuning parameter, despite defining the neighborhood in which SAM measures sharpness. We analyze mini-batch SAM near an interpolating minimum through linear stability. Under local linearization and gradient-noise alignment assumptions, we prove that every linearly stable minimum satisfies λmax≤3bΓ/(2ρη2), where λmax is the largest Hessian eigenvalue, b is the batch size, η is the learning rate, and Γ bounds the gradient norm. The bound quantitatively characterizes SAM's implicit flatness bias: holding the other quantities fixed, a smaller batch size, a larger learning rate, or a larger radius restricts linearly stable SAM to flatter minima. It also exposes a necessary trade-off: ρ should be large enough to promote flatness, yet remain local enough to preserve the approximation and stable training. We validate this prediction in a controlled study of 900 models on CIFAR-100 with ResNet-18 and VGG-19, where increasing ρ is consistently associated with a smaller largest Hessian eigenvalue across batch-size and learning-rate settings. Finally, we instantiate the analysis in Taylor-Locality Controlled SAM (TLC-SAM), which adjusts ρ using the observed Taylor-approximation error and further reduces the top Hessian eigenvalue relative to fixed-radius SAM. Our results provide quantitative hyperparameter bounds and a stability--locality perspective for analyzing and designing SAM variants.
Sharpness-Aware Minimization (SAM) improves generalization by minimizing the worst-case loss within a fixed parameter-space radius neighborhood. SAM and its variants mainly rely on a first-order linearized surrogate, while flat minima are inherently a second-order (curvature) notion.We revisit this mismatch and propose Loss-Equated SAM (LE-SAM), which inverts the traditional SAM mechanism that fixed perturbation radius with a fixed loss-space budget,effectively removing gradient-norm-dominated learning signals and shifting optimization toward curvature-dominated terms. Extensive experiments across diverse benchmarks and tasks demonstrate the strong generalization ability of LESAM that consistently outperforms SAM and even its variants, achieving the state-of-the-art performance.
Jinping Wang, Qinhan Liu, Zhiwu Xie +1
Wenzhou-Kean University · International Frontier Interdisciplinary Research Institute, Wenzhou-Kean University