LionMuon: Alternating Spectral and Sign Descent for Efficient Training
Organizations: Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) · Basic Research of Artificial Intelligence Laboratory (BRAIn Lab) · Applied Artificial Intelligence Institute · Innopolis University
Abstract
Pretraining a language model takes enormous compute, and the right optimizer can save a good part of it. Muon's spectral step gives a stronger direction than a sign step, but it is expensive. Every step runs Newton-Schulz iterations on the full matrix and, in distributed training, an extra all-reduce. Sign steps, as in Lion and Signum, are cheap and stay local to each device. We propose LionMuon, which takes one Muon step every iterations and Lion steps in between, with a single dual-EMA momentum buffer shared by both. Muon's compute and communication are paid once per steps, and the optimizer state is half of AdamW's. A single-EMA variant, SignMuon, already improves on Muon. We prove complexity bounds under heavy-tailed noise in which the period sets an interpolation between Muon's and Lion's smoothness and noise constants, and which say when LionMuon is faster than both. On 124M and 355M models trained on FineWeb, LionMuon with and reaches a lower loss than Muon, AdamW, Lion and Signum at the same number of tokens. Under 4-GPU data-parallel training it reaches Muon's final loss with a third less wall-clock on PCIe, and it beats the communication-efficient Muon variants Dion and MuonBP on loss at no more exposed communication, while keeping the exact gradient. Code: https://github.com/brain-lab-research/lion-muon
Figures & tables
| Exposed | Step time / Muon | ||||
|---|---|---|---|---|---|
| Optimizer | Newton–Schulz | MB/step | State | PCIe | NVLink |
| AdamW | no | 0 | 0.75 | 0.85 | |
| Lion / Signum | no | 0 | – | – | |
| Muon | every step | 170 | 1.00 | 1.00 | |
| LionMuon | every 2nd step | 85 | 0.88 | 0.92 | |
| LionMuon | every 5th step | 34 | 0.81 | 0.89 | |
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| FineWeb | |||||||
|---|---|---|---|---|---|---|---|
| WikiText-103 |
| Optimizer | Momentum | Period |
|---|---|---|
| Signum [ Bernstein et al., 2018 ] | ||
| Lion [ Chen et al., 2023 ] | (dual-EMA) | |
| Muon [ Jordan et al., 2024 ] | ||
| SignMuon (this work) | any | |
| LionMuon (this work) | (dual-EMA) | any |
| Method | / | Loss | MB per step | Steps to Muon ’s loss |
|---|---|---|---|---|
| AdamW | 3.407 | 496 | 137k | |
| Muon | 3.418 | 667 | 147k | |
| Muon | 3.467 | 667 | never | |
| Muon | 3.756 at 96k, stopped | 667 | never | |
| LionMuon | 3.391 | 667 | 133k | |
| SignMuon | / | 3.395 | 581 | 136k |
| Steps | FLOPs | Bytes | Time, PCIe | Time, NVLink | |
|---|---|---|---|---|---|
| Muon | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| LionMuon | 0.86 | 0.81 | 0.75 | 0.75 | 0.79 |
| LionMuon | 0.86 | 0.79 | 0.69 | 0.70 | 0.77 |
| SignMuon | 0.93 | 0.88 | 0.81 | 0.81 | 0.85 |
| AdamW | 0.93 | 0.84 | 0.69 | 0.70 | 0.79 |
| Lion | 0.95 | 0.85 | 0.71 | – | – |
| Method | / | Val loss |
|---|---|---|
| SignMuon | / | |
| LionMuon | / | |
| LionMuon | / | |
| Muon | ||
| Lion | ||
| AdamW |
| Model architecture | |
|---|---|
| Number of layers | 12 |
| Number of heads | 12 |
| Embedding dim | 768 |
| Sequence length | 512 |
| Vocabulary size | 50,304 (GPT-2 BPE) |
| Architectures | GPT-2 base |
| method | swept | swept | cells | chosen | loss (next down / up) |
|---|---|---|---|---|---|
| Muon ( ) | 0.0003–0.003 | – | 3+0 | 0.001 | 3.528 (3.596 / 3.534) |
| SignMuon | 0.001–0.01 | 3–300 | 8+3 | (0.003, 30) | 3.511 (3.529 / cut) |
| SignMuon | 0.001–0.03 | 3–1000 | 13+3 | (0.01, 300) | 3.508 (3.534 / cut) |
| SignMuon | 0.003–0.1 | 10–1000 | 8+4 | (0.01, 100) | 3.540 (3.581 / 3.551) |
| LionMuon | 0.0003–0.003 | – | 2+1 | 0.001 | 3.511 (3.540 / cut) |
| LionMuon | 0.0003–0.01 | 1–300 | 12+2 | (0.001, 10) | 3.501 (3.538 / 3.507) |
| method | swept | swept | cells | chosen | loss (next down / up) |
|---|---|---|---|---|---|
| Muon ( ) | 0.001–0.01 | – | 2+1 | 0.003 | 2.838 (2.884 / cut) |
| SignMuon | 0.0003–0.03 | 1–3000 | 12+5 | (0.003, 3) | 2.835 (2.885 / cut) |
| SignMuon | 0.0001–0.03 | 1–100 | 10+5 | (0.01, 30) | 2.842 (2.890 / cut) |
| SignMuon | 0.001–0.1 | 3–3000 | 9+8 | (0.03, 100) | 2.859 (2.902 / cut) |
| LionMuon | 0.001–0.01 | – | 2+1 | 0.003 | 2.826 (2.861 / cut) |
| LionMuon | 0.0003–0.01 | 1–3000 | 8+8 | (0.003, 10) | 2.827 (2.881 / cut) |
| Fixed period | Random period | Lion on 1D parameters | |
|---|---|---|---|
| LionMuon | ( ) | ( ) | ( ) |
| LionMuon | ( ) | ( ) | ( ) |