cs.LGOct 1, 2026

Learning Rate Transfer for Hybrid Transformer-SSM Architectures

Authors: Jimin Seo, Gyubok Lee, Yeonsik Jo, Kiwoong Yoo, Yeongoon Kim, Minhae Oh, Jin Woo Koo, Suhwan Kim, +6 more

Organizations: Seoul National University · SB Intuitions · LG AI Research · Hodoo AI

Abstract

We study learning rate (LR) scaling for hybrid architectures combining Transformer and State-Space Model (SSM) blocks, a class adopted by several recent production language models. In particular, we focus on the gap between the theoretical scaling rules derived for SSMs under zero-order-hold (ZOH) discretization at infinite width with growing state size, and the field-standard practical implementations using simplified-ZOH Mamba at fixed state size. Surprisingly, in this practical regime hybrid architectures achieve a near-zero LR transfer gap across widths 256-2048 and depths 4-32 up to billion-parameter scale using only the original μμP prescription, even though SSM operations fall outside its Tensor Programs representability conditions and every parameterization we test fails the standard coordinate-check diagnostic of μμP correctness. We attribute this to a two-condition decomposition of LR transfer in hybrid architectures: a global update-to-weight invariance, enforced by μμP's initialization and LR scaling; and a local per-component balance, provided by AdamW's per-parameter normalization. Our observations show that the optimal LR is invariant to width up to 8×\times, that this width invariance holds across depth, sequence length, batch size, and Transformer-to-SSM ratio, and that it transfers to Nemotron-H, a production hybrid outside our custom architecture set. We hope these findings fill the gap between theoretical scaling rules and practical hybrid implementations, and stimulate further research toward bridging it.

Figures & tables

Appendix figures & tables47 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate

    May 20, 2026Dayal Singh Kalra, Maissam BarkeshliHyperparameterLarge Language Model Training

  2. Priming: Hybrid State Space Models From Pre-trained Transformers

    May 8, 2026Aditya Chattopadhyay, Elvis Nunez, Prannay Kaul +6Generative Pretrained TransformersPriming

  3. Long-Context Modeling via GSS-Transformer Hybrid Architecture with Learnable Mixing

    Jun 15, 2026Kuzey Torlak, Hüseyin Arda Arslan, Anıl Dervişoğlu +2Efficient Long-Context InferenceHybrid Transformer