Learning-rate (LR) scheduling plays a central role in large language model (LLM) pretraining, yet current practice still relies heavily on hand-crafted heuristics such as Warmup-Cosine-Decay and Warmup-Stable-Decay. Because these schedules are fixed in advance, they cannot adapt to evolving optimization dynamics. Online learned scheduling within the Learning to Optimize (L2O) framework offers a dynamic alternative, but remains brittle at LLM scale due to noisy signals, delayed feedback, and the risk of catastrophic divergence. We propose SOLAR (State-driven Online Learning rAte scheduleR), a stabilized framework for reliable online LR adaptation. SOLAR uses a base schedule as a reference and learns bounded, state-dependent residual corrections for individual parameter groups. Each correction re-anchors to the base at every step, allowing the policy to adapt the LR without relearning the warmup-decay profile. A lightweight state representation and progress-aware reward guide online learning, while a Circuit-Breaker restores training after rare unsafe actions. Across autoregressive language-model pretraining, SOLAR improves final perplexity over tuned static schedules and automatic LR tuners for dense models from 60M to 1B, AdamW and Muon, and two MoE settings up to 3B. Matched 130M controls show that adding base anchoring and action bounds improves a global PPO controller from 27.09 to 23.74 final PPL, while group-wise control reaches 22.87 on the same two seeds. A residual policy trained on a 60M proxy can also be frozen and reused at larger dense scales without target PPO updates, remaining effective across a fourfold base-LR range. These results establish SOLAR as a practical learned LR controller for LLM pretraining.
Figures & tables
Figure 1: Overview of SOLAR. Panel (a) presents the optimization perspective: SOLAR augments a base scheduler with residual learning-rate modulation, applies bounded group-wise LR actions to the LLM pretraining process, and uses an automated Circuit-Breaker to roll back unsafe trajectories when severe loss spikes are detected. Panel (b) illustrates the RL control loop: SOLAR constructs lightweight global and local state features from the current training dynamics, outputs parameter-group-wise residual actions, receives delayed reward feedback after the optimizer update, and improves the scheduler through PPO updates.
Figure 2: Pretraining performance on C4 across model scales. SOLAR consistently achieves the lowest perplexity, outperforming strong baselines in both AdamW and Muon families.
Model Scale (Validation Perplexity ↓ )
Method
60M
130M
350M
1B
Base Optimizer: AdamW
Cosine
30.49
24.52
18.31
16.52
AvgLR Replay
30.68
25.61
21.69
20.06
WSD
29.80
23.96
18.75
16.29
CLR
31.78
28.14
21.43
20.22
Table 1: Seed-52 final-checkpoint validation perplexity on C4. SOLAR-online acquires its policy within the target run; SOLAR-frozen reuses a policy acquired from full-length online runs at 60M, with no target PPO updates. AvgLR Replay replays SOLAR’s step-wise mean LR. † marks official AdamW-based implementations. Multi-seed results appear in Appendix C.2 .
Figure 3: Qwen2-MoE 1B pretraining on The Pile. SOLAR improves validation perplexity, suppresses early gradient-norm spikes, and maintains a higher average LR through group-wise micro-modulation. Panel (d) shows one representative attention projection group; the solid line is a 200-step moving average and the shaded region visualizes the corresponding step-wise variation.
Figure 4: Mechanistic analysis of SOLAR on 130M Llama 2. (a)–(d): average effective LR, single-group LR dynamics, module-wise LR allocation, and per-module LR–GradNorm correlation, respectively. In (b), the solid line is a 200-step moving average and the shaded region visualizes the corresponding step-wise variation.
Appendix figures & tables26 assets
Supplementary material from the paper’s appendix.
Appendix
Quantity
Default
Status
Values evaluated
Short-term progress scale
20
tunable
{10,20,40} and removal
EMA trend scale
2
tunable
{1,2,4} and removal
Severe stability offset
20
tunable
{10,20,40} and removal
Severe stability threshold
3.0
tunable
{2.5,3.0,4.0}
State short-loss EMA decay βs
0.90
tunable
{0.90,0.95}
State-normalization EMA decay
0.99
fixed
fixed across experiments
Appendix
Table 2: Controller-coefficient census on 130M/C4/AdamW.
Variant
Seed 52
Seed 42
Mean PPL
Δ vs. full
CB activations
Full reward
22.79
22.95
22.87
0.00
0/2
− immediate progress
24.91
24.71
24.81
+1.94
0/2
− EMA trend
23.71
23.83
23.77
+0.90
0/2
− stability shaping term
23.24
23.09
23.17
+0.30
1/2
Appendix
Table 3: Reward leave-one-out on 130M/C4/AdamW using seeds 42 and 52.
60M
130M
350M
1B
dmodel
512
768
1024
2048
nlayers
8
12
24
24
nheads
8
12
16
32
dff
1376
2048
2736
5461
Vocab
32000
32000
32000
32000
Appendix
Table 4: Architecture of Llama 2 dense models.
Optimizer
60M
130M
350M
1B
AdamW
3×10−3
1×10−3
1×10−3
5×10−4
Muon
5×10−3
3×10−3
3×10−3
1×10−3
Appendix
Table 5: Selected peak LR for each dense-model scale. All settings use weight decay 0.01 and 10% warmup. SOLAR inherits the corresponding Cosine base and does not run a separate base search.
Method
Selected base LR / WD
Method-specific search → selected
Cosine
10−3/0.01
decay to 0.1× peak
WSD
10−3/0.01
stable fraction {70,80,90}%→80% ; decay starts at 0.9T
CLR
max 10−3 , base 10−4 / 0.01
period {1,2,4} K →2 K; triangular2
Blockwise LR
10−3/0.01
segments {2,3,4}→3 ; decay {0.5,0.3}→0.5
Schedule-Free
10−3/0.01
β∈{0.9,0.95,0.98}→0.9
Prodigy
adaptive ( d0=10−6 ) / 0.01
dcoef∈{0.1,0.5,1.0}→1.0
Appendix
Table 6: Scheduler-specific search at 130M/C4/AdamW. Each method uses the common LR grid unless its update rule determines the LR. No non-scheduler training hyperparameter is retuned.
Parameter
dmodel
nlayers
nheads
nkv_heads
dff
dmoe_ff
dshared_ff
nexperts
Value
768
15
12
12
3072
768
3072
32
Appendix
Table 7: Architecture of the Qwen2-MoE 1B model.
Method
Learned object
Acquisition
Evaluated setting
Daniel et al. (2016)
global step size
repeated episodes
small networks
Xu et al. (2017)
global absolute LR
actor–critic with resets
vision
Xu et al. (2019)
global LR profile
past complete histories
Fashion-MNIST, CIFAR-10
GNS (2022)
global graph policy
separate target episodes
vision, GLUE
Subramanian et al. (2023)
PPO LR schedule
separate controller training
MNIST, CIFAR-100
GANNO (2023)
layer-wise absolute LR
separate environments
vision
Appendix
Table 8: Positioning among representative learned LR controllers. “Same live run” means that policy updates occur inside the final run being improved.
Setting
Method
Seed 42
Seed 52
Seed 62
Mean PPL
60M AdamW
Cosine
30.71
30.49
30.34
30.51
WSD
29.63
29.80
29.91
29.78
SOLAR-online
28.99
28.92
29.14
29.02
130M AdamW
Cosine
24.41
24.52
24.57
24.50
WSD
24.07
23.96
23.90
23.98
SOLAR-online
22.95
22.79
22.84
22.86
Appendix
Table 9: Final-checkpoint PPL by seed. Seed 42 is the selection seed for online and baseline configurations; seeds 52 and 62 use the selected configuration. SOLAR-frozen reuses one source-trained policy across all target seeds.
Full-length 60M source runs K
Frozen final PPL
0
25.04
1
23.26
2
22.89
3
22.58
5
22.41
Appendix
Table 10: Source-side acquisition and 130M/C4/AdamW frozen reuse. Every target evaluation uses seed 52. The K=0 row keeps the policy at its random initialization.
Source → target
Cosine
SOLAR-frozen
SOLAR-online
60M/C4/AdamW → 130M/C4/AdamW
24.52
22.41
22.79
60M/C4/Muon → 130M/C4/Muon
22.55
21.87
21.99
60M/C4/AdamW → 130M/Pile/AdamW
13.96
13.21
12.99
60M/C4/AdamW → 130M/C4/Muon
22.55
22.49
21.99
Appendix
Table 11: Online acquisition and frozen reuse across source–target settings. Rows with C4 targets use seed 52; the Pile-target row reports the mean of seeds 42 and 52.
Method and control structure
Per-seed PPL → summary
Isolated axis
SOLAR-online: anchored group residual, G=111
22.95/22.79→22.87
full method
N1 Global-PPO-Residual: anchored, G=1
23.68/23.80→23.74
group granularity
WSD: tuned static base
24.07/23.96/23.90→23.98±0.09
learned control
N3 Group-MHD: anchored, group-wise, non-RL
24.31/24.21→24.26
learned feedback
Cosine: tuned static base
24.41/24.52/24.57→24.50±0.08
learned control
N4 GANNO-IPPO: layer-wise absolute LR
24.89/25.01→24.95
residual design
Appendix
Table 12: Matched controls that isolate the SOLAR control structure. The matched SOLAR/N1/N2 comparisons use seeds 42 and 52; each arrow gives the mean. Three-run static-baseline summaries include the sample standard deviation.
α
0.8
1.0
1.3
1.6
1.8
Final PPL
23.42
23.05
22.79
22.94
23.61
CB activation
no
no
no
no
yes
Appendix
Table 13: Sensitivity to the residual action bound on 130M/C4/AdamW.
Base peak
130M Cosine
130M frozen
350M Cosine
350M frozen
0.5×
25.05
22.72
18.79
17.61
1.0×
24.52
22.41
18.31
17.37
2.0×
diverged @4.1K
22.68
diverged @6.8K
17.96
Appendix
Table 14: Frozen residual-policy transfer under misspecified target base LRs.
Method
Searched SP base
μ P-transferred base
Cosine
24.50±0.08 ( n=3 )
24.44/24.53→24.49
SOLAR-online
22.86±0.08 ( n=3 )
22.87/22.94→22.91
SOLAR-frozen
22.41 ( n=1 )
22.62/22.74→22.68
Appendix
Table 15: SOLAR on a directly searched standard-parameterization (SP) base and a μ P-transferred base at 130M/C4/AdamW. For the two-run μ P settings, values before the arrow are per-seed final PPL and the arrow gives their mean.
Method
Seed 52
Seed 42
Mean
Gain vs. Cosine
Cosine
13.90
14.02
13.96
–
SOLAR-online
12.94
13.03
12.99
0.97 (6.9%)
SOLAR-frozen (60M/C4 policy)
13.18
13.24
13.21
0.75 (5.4%)
Appendix
Table 16: Cross-corpus evaluation on 130M/The Pile/AdamW.
Method
Group-wise
Time-varying
Live state
Final PPL
Cosine
–
base only
–
24.52
AvgLR Replay
–
✓
–
25.61
Level-corrected scalar replay
–
✓
–
24.63
Static normalized group profile
✓
–
–
24.19
SOLAR-online
✓
✓
✓
22.79
Appendix
Table 17: Replay controls on 130M/C4/AdamW with seed 52. All entries report final PPL.
Setting
Cosine
Static group
Time-varying group
SOLAR-frozen
SOLAR-online
C4, seed 42, 1.0× base
24.41
24.29
23.94
22.96
22.95
C4, seed 52, 0.5× base
25.05
24.83
24.47
22.72
–
C4, seed 52, 2.0× base
div. @4.1K
div. @3.8K
div. @4.4K
22.68
–
Pile, seed 52, 1.0× base
13.90
13.71
13.55
13.18
12.94
Appendix
Table 18: SOLAR-derived replays across seeds, base LRs, and corpora. Both replay profiles come from the seed-52 C4 run at the selected base LR. The Pile row uses seed 52; “–” denotes a setting not evaluated.
State Variant
Final Eval PPL ( ↓ )
Full state (Ours)
22.79
Global-only
23.11
Local-only
23.57
Appendix
Table 19: Ablation on state representation for the 130M model.
Policy Formulation
Final Eval PPL ( ↓ )
Stochastic (Ours)
22.79
Near-deterministic
25.03
Appendix
Table 20: Comparison of stochastic exploration versus near-deterministic scheduling (130M model).
Design
Final Eval PPL ( ↓ )
SOLAR (Residual + Warmup)
22.79
Direct-Abs
24.74
Appendix
Table 21: Ablation on residual LR modulation versus direct absolute LR prediction (130M model).
Figure 5: Sensitivity analysis and mechanism illustration for SOLAR.
Figure 6: Demonstration of the automated Circuit-Breaker mechanism under an aggressive exploration stress test ( α=1.8 ). In both panels, the teal line represents the initial training rollout, which experiences a severe spike and triggers the Circuit-Breaker (marked by the red cross). The orange line represents the automatically resumed trajectory, which restarts from the last safe checkpoint (indicated by the vertical dotted line) and successfully completes the training process.
Figure 7: Additional validation on the DeepSeek-V2 3B MoE model in Megatron. SOLAR is compared against a matched WSD baseline under the same model, data, tokenizer, optimizer, and training setup.
Method
State- conditioned
PPO updates
Stochastic sampling
Action noise
Final Eval PPL ↓
SOLAR
✓
✓
✓
learned
22.79
State-agnostic stochastic scheduler
–
–
✓
selected σ
25.21
Untrained fixed-policy controller
✓
–
✓
fixed init
25.04
Appendix
Table 22: Controls for state-independent stochasticity and untrained policy initialization. All experiments use the 130M Llama 2 model with AdamW on C4.
Method
Step Time (s)
Throughput (tok/s)
Overhead Increase (%)
Baseline LRS
1.976
66329.01
0.00
SOLAR (frozen)
1.992
65799.20
0.81
SOLAR (online)
2.001
65499.35
1.27
Appendix
Table 23: Steady-state step-time overhead on the 1B AdamW setting with 8× RTX 4090. The post-warmup benchmark window excludes evaluation and checkpointing.
Setting
Baseline (h)
SOLAR-online (h)
Online OH
Frozen OH
1B AdamW
56.75
57.45
1.23%
0.76%
1B Muon
65.34
66.04
1.07%
0.69%
Qwen2-MoE 1B
127.38
129.27
1.48%
–
DeepSeek-V2-style MoE 3B
163.38
165.00
0.99%
–
Appendix
Table 24: Full-run wall-clock accounting. These measurements include evaluation, checkpointing, state construction, policy inference, action application, PPO updates for the online mode, and any rollback time. No listed run triggered rollback. Dense and Qwen2-MoE runs use 8× RTX 4090; the 3B MoE runs use 32× A800 80GB.