LionMuon: Alternating Spectral and Sign Descent for Efficient Training
Authors: Arman Bolatov, Artem Riabinin, Nikita Kornilov, Andrey Veprikov, Samuel Horváth, Martin Takáč, Aleksandr Beznosikov
Organizations: Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) · Basic Research of Artificial Intelligence Laboratory (BRAIn Lab) · Applied Artificial Intelligence Institute · Innopolis University
Pretraining a language model takes enormous compute, and the right optimizer can save a good part of it. Muon's spectral step gives a stronger direction than a sign step, but it is expensive. Every step runs Newton-Schulz iterations on the full matrix and, in distributed training, an extra all-reduce. Sign steps, as in Lion and Signum, are cheap and stay local to each device. We propose LionMuon, which takes one Muon step every P iterations and Lion steps in between, with a single dual-EMA momentum buffer shared by both. Muon's compute and communication are paid once per P steps, and the optimizer state is half of AdamW's. A single-EMA variant, SignMuon, already improves on Muon. We prove complexity bounds under heavy-tailed noise in which the period sets an interpolation between Muon's and Lion's smoothness and noise constants, and which say when LionMuon is faster than both. On 124M and 355M models trained on FineWeb, LionMuon with P=2 and P=5 reaches a lower loss than Muon, AdamW, Lion and Signum at the same number of tokens. Under 4-GPU data-parallel training it reaches Muon's final loss with a third less wall-clock on PCIe, and it beats the communication-efficient Muon variants Dion and MuonBP on loss at no more exposed communication, while keeping the exact gradient. Code: https://github.com/brain-lab-research/lion-muon
Figures & tables
Figure 1: Sign steps are cheap, Muon steps are strong but expensive, and our methods alternate between the two.
Figure 2: Loss against training FLOPs at 124M on FineWeb (left) and WikiText-103 (right). Every optimizer is tuned, and the error bars are over three seeds. LionMuon with P∈{2,5} is lowest on both datasets and uses fewer FLOPs than Muon .
Figure 3
Exposed
Step time / Muon
Optimizer
Newton–Schulz
MB/step
State
PCIe
NVLink
AdamW
no
0
2∣W∣
0.75
0.85
Lion / Signum
no
0
∣W∣
–
–
Muon
every step
170
∣W∣
1.00
1.00
LionMuon P=2
every 2nd step
85
∣W∣
0.88
0.92
LionMuon P=5
every 5th step
34
∣W∣
0.81
0.89
Table 1: What one step costs on top of forward and backward, at 124M on four GPUs. Exposed communication is the optimizer’s own all-reduce, which cannot start before the backward pass ends. Step times are relative to Muon , measured by running all methods back to back in a random order within each round and taking the ratio inside the round, so that any background load on the node falls on every method alike (30 rounds on PCIe, 40 on NVLink, 95% intervals within ±0.01 and ±0.08 ). Muon ’s median step took 1.56 s on PCIe and 0.24 s on NVLink. Dion ’s bytes replace the gradient all-reduce. State is per 2D parameter W .
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
α
ρnuc
ρ1
L2
L∞
L∞/(α2L2)
ρ1/(αρnuc)
FineWeb
67
13.6
894
0.79
857
0.24
0.98
WikiText-103
75
10.9
835
1.46
1312
0.16
1.02
Appendix
Table 2: Constants of the bound, measured during 124M training with P=2 . At each step we take the median over the 2D parameters, and the table gives the mean over the run (100 records per run). L2 and L∞ are single-batch estimates and therefore noisier. The last two columns are the ratios the bound depends on.
Figure 5: The same constants over the course of training, on 124M runs with P=2 , with the mean over the run in each legend. The gradient ratio α falls over the first 2000 steps and then holds. The last panel is the update ratio ∥U∥2/∥U∥∞ , which stays far below the worst case mn that the dense approximation would otherwise have to assume.
Optimizer
Momentum
Period
Signum [ Bernstein et al., 2018 ]
β1=β2
P=∞
Lion [ Chen et al., 2023 ]
β1=β2 (dual-EMA)
P=∞
Muon [ Jordan et al., 2024 ]
β1=β2
P=1
SignMuon (this work)
β1=β2
any P
LionMuon (this work)
β1=β2 (dual-EMA)
any P
Appendix
Table 3: Special cases of LionMuon . Both conditions, on β1,β2 and on P , must hold.
Method
ηM / ηL
Loss
MB per step
Steps to Muon ’s loss
AdamW
10−3
3.407
496
137k
Muon
10−3
3.418
667
147k
Muon
3×10−4
3.467
667
never
Muon
3×10−3
3.756 at 96k, stopped
667
never
LionMuon P=1
10−3
3.391
667
133k
SignMuon P=2
3×10−3 / 10−4
3.395
581
136k
Appendix
Table 4: 124M on four GPUs with data parallelism, 150,000 steps, one seed per run. Loss is the best validation loss of the run, bytes are what one step sends (Table 1 ), and the last column is the first evaluation at or below Muon ’s own final loss (evaluations every 1,000 steps).
Steps
FLOPs
Bytes
Time, PCIe
Time, NVLink
Muon
1.00
1.00
1.00
1.00
1.00
LionMuon P=2
0.86
0.81
0.75
0.75
0.79
LionMuon P=5
0.86
0.79
0.69
0.70
0.77
SignMuon P=2
0.93
0.88
0.81
0.81
0.85
AdamW
0.93
0.84
0.69
0.70
0.79
Lion
0.95
0.85
0.71
–
–
Appendix
Table 5: What it costs to reach Muon ’s final loss, relative to Muon , at 124M on four GPUs. Steps are the first evaluation at or below that loss, and the other columns multiply them by the per-step costs of Table 1 . Step times are given for the methods of Table 1 . Dion at rank 1/4 and Signum never reach it.
Method
ηM / ηL
Val loss
SignMuon P=2
3×10−3 / 10−4
3.008
LionMuon P=2
10−3 / 10−4
3.020
LionMuon P=5
3×10−3 / 10−4
3.023
Muon
10−3
3.040
Lion
3×10−4
3.050
AdamW
10−3
3.064
Appendix
Table 6: 355M on FineWeb, 8.2 B tokens, every hyperparameter copied from 124M without retuning. Best validation loss, one seed per method.
Model architecture
Number of layers
12
Number of heads
12
Embedding dim
768
Sequence length
512
Vocabulary size
50,304 (GPT-2 BPE)
Architectures
GPT-2 base
Appendix
Table 7: Full experimental configuration.
method
ηM swept
α swept
cells
chosen (ηM,α)
loss (next ηM down / up)
Muon ( P=1 )
0.0003–0.003
–
3+0
0.001
3.528 (3.596 / 3.534)
SignMuon P=2
0.001–0.01
3–300
8+3
(0.003, 30)
3.511 (3.529 / cut)
SignMuon P=5
0.001–0.03
3–1000
13+3
(0.01, 300)
3.508 (3.534 / cut)
SignMuon P=20
0.003–0.1
10–1000
8+4
(0.01, 100)
3.540 (3.581 / 3.551)
LionMuon P=1
0.0003–0.003
–
2+1
0.001
3.511 (3.540 / cut)
LionMuon P=2
0.0003–0.01
1–300
12+2
(0.001, 10)
3.501 (3.538 / 3.507)
Appendix
Table 8: The tuning sweep at 124M on FineWeb, every cell a full 64,000 -step run. ‘‘Cells’’ counts the runs that went to the end plus the ones the pruning rule stopped early. The last column is the loss of the chosen cell, with the losses at the next lower and next higher ηM (same α ) in brackets.
method
ηM swept
α swept
cells
chosen (ηM,α)
loss (next ηM down / up)
Muon ( P=1 )
0.001–0.01
–
2+1
0.003
2.838 (2.884 / cut)
SignMuon P=2
0.0003–0.03
1–3000
12+5
(0.003, 3)
2.835 (2.885 / cut)
SignMuon P=5
0.0001–0.03
1–100
10+5
(0.01, 30)
2.842 (2.890 / cut)
SignMuon P=20
0.001–0.1
3–3000
9+8
(0.03, 100)
2.859 (2.902 / cut)
LionMuon P=1
0.001–0.01
–
2+1
0.003
2.826 (2.861 / cut)
LionMuon P=2
0.0003–0.01
1–3000
8+8
(0.003, 10)
2.827 (2.881 / cut)
Appendix
Table 9: The same sweep on WikiText-103.
β1\β2
0.9
0.95
0.99
0.9
3.511
3.508
3.501
0.95
3.509
3.502
0.99
3.510
Appendix
Table 10: Loss at 124M on FineWeb with P=2 for each momentum pair, after tuning the rate and the ratio for that pair. The diagonal is SignMuon , the rest is LionMuon .
Fixed period
Random period
Lion on 1D parameters
LionMuon P=2
3.498 ( 0.002 )
3.499 ( 0.003 )
3.535 ( 0.001 )
LionMuon P=5
3.498 ( 0.002 )
3.499 ( 0.003 )
3.538 ( 0.002 )
Appendix
Table 11: Loss at 124M on FineWeb, 64,000 steps, mean over three seeds with the spread in brackets. The first column is the tuned setting of Section 5 , the other two are the ablations.
In large-scale optimization, the cheapness and effectiveness of update steps are the most crucial factors for a successful optimizer. Sign-based optimizers like Lion or Signum produce cheap per-step updates, whereas Muon's spectral matrix-sign update gives a much stronger direction at a substantially higher per-step cost. In this work, we propose LionMuon, which retains the effectiveness of Muon steps while considerably cutting the averaged iteration cost, similar to sign-based methods. It alternates between Lion's and Muon's updates on a fixed period P, sharing a single dual-EMA momentum buffer between them. The optimizer state memory therefore matches Lion and is exactly half of AdamW's. A simpler single-EMA variant, SignMuon, by itself already outperforms pure Muon. At P = 2, LionMuon Pareto-dominates Muon, Lion, Signum, and AdamW on every dataset and architecture we tested at 124M model size, reaching lower validation loss at lower compute, and the same advantage persists at 355M and 720M scale. On the theory side, we prove sharp complexity bounds under heavy-tailed noise which are governed by period-averaged smoothness and noise that interpolate between Muon's and Lion's constants. These bounds predict the compute-optimal period and the conditions under which LionMuon outruns Muon and Lion. Code: https://github.com/brain-lab-research/lion-muon
Arman Bolatov, Artem Riabinin, Nikita Kornilov +4
Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) · Basic Research of Artificial Intelligence Laboratory (BRAIn Lab) · Applied Artificial Intelligence Institute +1
Matrix-orthogonalization-based optimizers, exemplified by Muon, have demonstrated strong convergence behavior across a wide range of modern deep learning workloads. The matrix-aware updates offer a compelling alternative to conventional element-wise optimization, particularly as model architectures continue to grow in scale and heterogeneity. Yet contemporary distributed training infrastructure built around the assumption of element-wise optimizers is poorly matched to matrix-level optimizers such as Muon, whose updates couple entire weight matrices and require costly Newton-Schulz iterations. Vanilla Muon implementations incur more than 2x the cost of forward and backward passes. To close this gap, we present DMuon, an open-source distributed Muon implementation that integrates into existing training pipelines as a drop-in module, with no framework-level modifications. Across both embodied foundation model and large language model (LLM) training workloads, DMuon achieves a 1.48x-3.01x speedup in end-to-end step time and a 6.85x-163.00x speedup in optimizer-step time, bringing per-step latency to near-AdamW levels and enabling efficient scaling in our model training.
The Muon optimizer has recently offered a promising alternative to AdamW for large language model training, leveraging matrix orthogonalization to produce geometry-aware updates. However, like all first-order methods, Muon can become trapped in sharp local minima. In this work, we present MONA, an optimizer that bridges Muon's orthogonalization framework with curvature-aware acceleration. MONA adds an acceleration term directly into Muon's gradient processing pipeline. This term is calculated from the exponential moving average of gradient differences. We provide a detailed convergence analysis for MONA, showing that the acceleration term introduces curvature-sensitive corrections while preserving Muon's spectral-norm regularization. Empirically, MONA achieves better convergence and downstream task performance compared to both Muon and AdamW across three scales of Mixture-of-Experts pretraining, spanning from 1B to 68B parameters, with the largest model trained on 1 trillion tokens. Furthermore, we conduct supervised fine-tuning on the MOE-68B-A3B model and evaluate it on general capability, mathematical reasoning, and code generation benchmarks, where MONA achieves SOTA performance.