Distributed Learning with Selective State Space Models: Architecture-Aware Convergence Analysis
Organizations: Purdue University · University of Colorado Colorado Springs · University of Exeter
Abstract
Modern state space models (SSMs), such as Mamba2, provide a compelling alternative to transformers by combining linear-time sequence modeling with recurrent state-space dynamics. However, the behavior of SSMs in distributed learning settings remains poorly understood. In particular, the existing standard federated learning methods are largely architecture-agnostic, and do not account for the stability, selectivity, and state-space parameterization that characterize modern selective SSMs. To address this, we derive architecture-aware gradient and smoothness bounds for single- and multi-layer selective SSMs, and convergence bounds for FedAvg and FedProx, characterizing how recurrent stability, input-dependent discretization, and state projection norms affect federated optimization. We then numerically validate the single-layer bounds on sequences generated by a teacher SSM, using a learner that follows the analyzed recurrence. We use this analysis to formulate expectations about the effects of local training and client heterogeneity, and examine these expectations by comparing nine federated learning algorithms on Mamba2 language modeling across six text domains. These experiments illustrate how SSM-specific bounds can provide a basis for interpreting the behavior of practical federated learning algorithms.
Figures & tables
| Statement | Inequality checked | Evaluations | Violations | Max. utilization |
| Thm. 4.4 | , per sequence | 1,327,104 | 0 / 0 | / |
| Lemma 4.7 | 41,472 | 0 | ||
| Lemma 4.7 | finite-difference ratio | 41,472 | 0 | |
| Lemma 4.7 | 864 | 0 | ||
| Lemma B.1 | 864 | 0 |
| Expected | PPL (%) | ||||
| Method | Theory-linked mechanism | ||||
| FedProx ( Li et al., 2020 ) | radius, | ||||
| SCAFFOLD ( Karimireddy et al., 2020 ) | corrected disagreement | ||||
| FedDyn ( Acar et al., 2021 ) | heterogeneity; radius (proximal ) | ||||
| FedAlign ( Keçeci et al., 2025 ) | state-coordinate alignment | ||||
| FedSAM ( Qu et al., 2022 ) | expanded evaluation neighborhood | ||||
Appendix figures & tables38 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | Value |
| Input/vocabulary dimension | |
| State dimension | |
| Stability limit | |
| Configurations (datasets) | |
| Sequence length | stored tokens ( burn-in tokens discarded, generated) |
| Training / test sequences | / |
| Criterion | Threshold |
| Global token balance | largest probability in the token marginal |
| Marginal entropy | Shannon entropy of the token marginal (the maximum is attained by the uniform distribution) |
| Per-sequence majority (test split) | for each test sequence, the fraction of its positions occupied by its most frequent token; the th percentile of this fraction over the test sequences |
| Class coverage | each of the tokens occurs at least once in the training split and at least once in the test split |
| Constant sequences | fraction of sequences whose tokens are all identical |
| Sequence diversity | number of distinct sequences divided by the number of sequences |
| Setting | Train cap | Val cap | Train limiter data | Valid limiter data | Server val | |
| Setting 1 | 0.0 | 7,427 | 1,857 | stackexchange | stackexchange | 7,424 |
| Setting 1 | 0.7 | 7,427 | 1,857 | stackexchange | stackexchange | 7,431 |
| Setting 1 | 1.0 | 7,427 | 1,857 | stackexchange | stackexchange | 7,428 |
| Setting 2 | 0.0 | 7,367 | 1,842 | c4 | c4 | 7,360 |
| Setting 2 | 0.7 | 7,367 | 1,842 | c4 | c4 | 7,364 |
| Setting 2 | 1.0 | 7,367 | 1,842 | c4 | c4 | 7,368 |
| Dataset | Task | # Classes | Train cap | Eval cap | max_seq_len |
| ag_news | News topic | 4 | 40,000 | 5,000 | 256 |
| dbpedia_14 | Wikipedia topic | 14 | 40,000 | 5,000 | 256 |
| trec | Question type | 6 | (full) | 5,000 | 256 |
| imdb | Sentiment | 2 | 40,000 | 5,000 | 256 |
| Dataset | Prompt template and labels |
| ag_news | Article: {text}\nCategory: |
| Labels: World , Sports , Business , Tech | |
| dbpedia_14 | <title-aware template>...:Category: |
| Labels: 14-way Wikipedia topic labels | |
| trec | Question: {text}\nType: |
| Labels: Description , Entity , Abbreviation , Person , Location , Number |
| Depth | Total params | SSM (mixer) params | Embedding params | Per-layer mixer |
| 4 | 14,733,664 | 1,860,704 | 12,871,680 | 465,176 |
| 16 | 20,318,848 | 7,442,816 | 12,871,680 | 465,176 |
| Backbone | Layout | SSM fraction | Total params |
| mamba2 | Mamba2 | ||
| hybrid75 | attention at | ||
| hybrid50 | attention at | ||
| hybrid50alt | attention at | ||
| hybrid25 | attention at | ||
| attn0 | attention |
| Checkpoint | Total params |
| state-spaces/mamba2-130m | 130M |
| state-spaces/mamba2-370m | 370M |
| state-spaces/mamba2-780m | 780M |
| Component | Value |
| Loss | mean softmax cross-entropy over the next-token positions |
| Optimizer | SGD, no momentum, weight decay, or gradient clipping |
| Learning rate | |
| Batch size | sequences |
| Centralized updates | (epochs of updates, reshuffled every epoch) |
| Checkpoints and evaluation | every updates; bound and smoothness |
| Component | Value |
| Algorithms | FedAvg; FedProx |
| Base setting | clients, local updates, IID partition |
| Communication rounds | |
| Client-count variation | |
| Local-step variation | |
| Proximal variation (FedProx) |
| Component | Value |
| Number of clients, | |
| Participation rate | ; all clients participate each round |
| Communication rounds, | |
| Local steps per round, | |
| Per-step batch size | |
| Sequence length, |
| Algorithm | Hyperparameters and values |
| FedAvg | none |
| FedProx | |
| SCAFFOLD | ; ; ; |
| FedAdam | ; ; ; |
| FedAlign | align_mode mamba2_ref ; align_mamba2_heads |
| FedDyn |
| Component | Value |
| Optimizer | AdamW |
| Learning rate | |
| Weight decay | 0.0 |
| Gradient clipping | norm, threshold 1.0 |
| Train batch size | 4 for mamba2-130m, mamba2-370m , 8 for mamba2-780m |
| Eval batch size | 4 |
| cells | points | sequences | viol. (conf./tight.) | median util. (conf.) | median util. (tight.) | max util. (tight.) | median | |
| 0.3 ∗ | 1 | 18 | 18,432 | 0 / 0 | 0.0602 | |||
| 0.3 | 8 | 144 | 147,456 | 0 / 0 | 0.0345 | |||
| 0.5 | 9 | 162 | 165,888 | 0 / 0 | 0.146 | |||
| 0.7 | 9 | 162 | 165,888 | 0 / 0 | 0.369 | |||
| 0.9 | 9 | 162 | 165,888 | 0 / 0 | 0.746 | |||
| 0.95 | 9 | 162 | 165,888 | 0 / 0 | 0.868 |
| checkpoint | points | median util. (conf.) | median util. (tight.) | max util. (tight.) | median | median gradient norm |
| frozen_teacher | 216 | 0.177 | ||||
| learner_final | 216 | 0.0612 | ||||
| learner_initialization | 216 | 0.0293 | ||||
| learner_intermediate | 648 | 0.0438 |
| checkpoint | recurrent share of | |||||
| final (step 2000) | 0.3 | 0.148 | 0.428 | 0.431 | 0.000 | 0.994 |
| 0.5 | 0.145 | 0.434 | 0.433 | 0.000 | 0.996 | |
| 0.7 | 0.121 | 0.432 | 0.428 | 0.002 | 0.997 | |
| 0.9 | 0.044 | 0.369 | 0.415 | 0.039 | 0.999 | |
| 0.95 | 0.039 | 0.387 | 0.429 | 0.079 | 0.999 | |
| 0.99 | 0.006 | 0.187 | 0.204 | 0.600 | 1.000 |
| in | in | in | in | in | in | |
| 0.3 | 1.78 [1.71, 2.84] | -0.29 [-1.28, 0.11] | -0.92 [-0.95, 0.41] | 1.96 [1.81, 2.10] | 1.81 [0.82, 2.30] | 0.17 [0.14, 0.22] |
| 0.5 | 1.75 [1.48, 2.80] | 0.38 [-0.79, 0.77] | -0.97 [-1.09, 0.34] | 2.05 [1.76, 2.21] | 1.86 [1.21, 2.44] | 0.19 [0.12, 0.29] |
| 0.7 | 1.39 [1.30, 2.47] | 0.04 [-0.40, 0.89] | -1.23 [-1.49, -0.07] | 1.70 [0.94, 1.80] | 1.41 [1.32, 1.96] | 0.03 [-0.37, 0.27] |
| 0.9 | 1.86 [1.70, 2.19] | 0.48 [0.33, 1.64] | -1.28 [-1.66, -1.09] | 1.32 [0.74, 2.23] | 1.55 [0.30, 3.18] | 0.10 [-0.50, 0.87] |
| 0.95 | 1.74 [1.25, 2.01] | 0.31 [0.10, 1.40] | -1.63 [-1.84, -1.58] | 1.09 [0.87, 1.14] | 1.37 [0.28, 1.55] | 0.05 [-0.07, 0.23] |
| 0.99 | 1.85 [0.88, 2.23] | 1.05 [0.03, 1.96] | -1.38 [-1.58, -1.27] | 1.08 [0.95, 1.11] | 1.38 [0.75, 1.66] | 0.06 [-0.12, 0.12] |
| runs | recovered fraction | median excess NLL | median gap recovered [95% CI] | median final | |
| 0.3 | 27 | 0.93 | 0.0180 | 0.80 [0.52, 0.90] | 0.031 |
| 0.5 | 27 | 0.93 | 0.0212 | 0.79 [0.49, 0.82] | 0.134 |
| 0.7 | 27 | 0.93 | 0.0203 | 0.81 [0.40, 0.87] | 0.357 |
| 0.9 | 27 | 0.93 | 0.0187 | 0.69 [0.64, 0.96] | 0.724 |
| 0.95 | 27 | 1.00 | 0.0179 | 0.74 [0.52, 0.99] | 0.855 |
| 0.99 | 27 | 1.00 | 0.0027 | 0.95 [0.42, 1.03] | 0.966 |
| factor | range (decades) | std (decades) | share of | share of |
| memory | 6.2 | 2.06 | 93% | 102% |
| dimension | 1.2 | 0.35 | 7% | 12% |
| readout | 1.5 | 0.31 | 6% | 6% |
| input | 1.8 | 0.28 | 3% | 3% |
| step | 0.8 | 0.18 | 4% | 6% |
| transition | 2.0 | 0.50 | -13% | -11% |
| Algorithm | Train | Val | Last-10 | Train | Val | Last-10 | Train | Val | Last-10 |
| FedAvg | 75.4 0.5 | 63.7 0.3 | 45.4 0.2 | 99.3 0.2 | 66.7 0.6 | 45.8 0.3 | 143.5 0.7 | 136.5 3.8 | 60.7 0.5 |
| FedProx | 69.0 0.2 | 57.6 0.3 | 36.9 0.4 | 88.0 0.4 | 60.4 0.4 | 36.5 0.2 | 119.3 0.4 | 121.3 2.9 | 40.2 0.3 |
| SCAFFOLD | 78.3 0.5 | 66.3 0.2 | 46.9 0.1 | 102.4 0.4 | 67.3 0.3 | 45.2 0.1 | 149.0 0.5 | 131.6 3.6 | 49.7 0.2 |
| FedAdam | 88.4 0.4 | 164.8 1.6 | 53.0 0.3 | 112.7 0.4 | 160.0 2.1 | 50.8 0.1 | 157.2 0.6 | 355.2 13.8 | 61.7 0.3 |
| FedAlign | 75.3 0.4 | 63.6 0.2 | 45.2 0.0 | 99.2 0.4 | 66.6 0.3 | 45.7 0.2 | 143.5 0.5 | 135.3 5.4 | 60.1 0.4 |
| Algorithm | Train | Val | Last-10 | Train | Val | Last-10 | Train | Val | Last-10 |
| FedAvg | 151.2 0.5 | 108.5 1.7 | 63.9 0.2 | 100.9 0.5 | 85.5 1.1 | 48.7 0.2 | 66.1 0.3 | 72.9 2.0 | 39.2 0.4 |
| FedProx | 137.7 0.3 | 106.6 1.7 | 48.6 0.2 | 85.1 0.4 | 74.9 0.9 | 35.6 0.4 | 53.6 0.3 | 57.8 1.0 | 29.4 0.2 |
| SCAFFOLD | 150.8 0.4 | 105.6 1.7 | 61.0 0.3 | 104.5 0.3 | 85.3 0.7 | 45.4 0.3 | 74.5 0.4 | 74.3 2.0 | 35.4 0.0 |
| FedAdam | 167.3 0.5 | 261.5 5.6 | 69.6 0.5 | 113.0 0.4 | 217.6 6.6 | 53.1 0.1 | 78.1 0.3 | 200.9 5.2 | 42.8 0.1 |
| FedAlign | 151.2 0.5 | 108.3 2.1 | 63.6 0.2 | 100.8 0.4 | 85.1 1.7 | 48.4 0.3 | 66.0 0.3 | 72.0 2.3 | 39.0 0.2 |
| Setting 1 | Setting 2 | Setting 3 | Setting 4 | |||||||||
| Algorithm | Train | Val | Last-10 | Train | Val | Last-10 | Train | Val | Last-10 | Train | Val | Last-10 |
| FedAvg | 45.1 0.2 | 53.2 1.5 | 25.2 0.3 | 107.3 0.0 | 96.5 4.2 | 51.7 0.6 | 140.6 1.0 | 91.4 3.3 | 56.9 0.1 | 131.3 0.5 | 114.8 1.1 | 68.7 0.8 |
| FedProx | 42.3 0.2 | 46.4 0.2 | 19.0 0.2 | 93.8 0.3 | 86.4 0.3 | 38.9 0.3 | 117.7 0.6 | 87.3 3.1 | 43.0 0.3 | 114.6 0.5 | 98.9 1.7 | 50.5 0.9 |
| SCAFFOLD | 45.9 0.2 | 51.6 1.5 | 22.6 0.4 | 112.0 0.0 | 95.6 3.4 | 47.5 0.2 | 146.5 0.7 | 93.1 3.2 | 55.6 0.1 | 135.2 0.5 | 113.3 1.3 | 63.3 0.7 |
| FedAdam | 51.4 0.1 | 169.0 4.3 | 27.2 0.2 | 118.2 0.4 | 260.3 6.6 | 54.2 0.2 | 158.0 1.1 | 237.0 6.4 | 62.2 0.2 | 150.2 0.2 | 240.3 7.6 | 77.1 1.3 |
| FedAlign | 45.0 0.0 | 50.8 1.6 | 24.8 0.2 | 107.2 0.3 | 96.7 3.4 | 51.4 0.3 | 140.6 0.9 | 91.8 2.7 | 56.7 0.1 | 131.3 0.6 | 114.6 1.2 | 68.5 0.9 |
| Depth ( M) | Depth ( M) | |||||
| Algorithm | Train | Val | Last-10 | Train | Val | Last-10 |
| FedAvg | 109.8 0.5 | 88.8 0.5 | 54.8 0.3 | 102.3 0.4 | 89.1 2.6 | 46.4 0.2 |
| FedProx | 93.2 0.4 | 79.0 1.3 | 39.5 0.4 | 91.0 0.3 | 80.5 1.2 | 36.2 0.1 |
| SCAFFOLD | 112.8 0.4 | 88.0 1.4 | 50.1 0.0 | 107.0 0.4 | 88.8 2.0 | 44.5 0.2 |
| FedAdam | 122.6 0.3 | 235.1 4.1 | 59.8 0.3 | 116.2 0.5 | 218.2 7.4 | 50.5 0.3 |
| FedAlign | 109.7 0.5 | 89.1 1.9 | 54.4 0.2 | 102.4 0.4 | 87.9 2.2 | 46.3 0.3 |
| Algorithm | Train | Val | Last-10 | Train | Val | Last-10 | Train | Val | Last-10 | Train | Val | Last-10 |
| FedAvg | 144.4 0.3 | 89.7 0.9 | 64.2 0.4 | 141.4 0.4 | 87.3 0.5 | 61.6 0.2 | 136.9 0.6 | 86.3 0.7 | 60.9 0.3 | 129.0 0.3 | 86.2 0.6 | 61.1 0.2 |
| FedProx | 131.6 0.2 | 82.4 0.8 | 49.0 0.2 | 129.1 0.5 | 80.6 0.3 | 47.3 0.3 | 125.2 0.4 | 79.9 0.5 | 46.8 0.2 | 118.5 0.3 | 80.5 0.4 | 47.6 0.1 |
| SCAFFOLD | 143.7 0.2 | 87.6 0.9 | 62.4 0.4 | 139.5 0.5 | 83.8 0.5 | 58.7 0.2 | 134.2 0.7 | 81.2 0.6 | 55.7 0.4 | 126.8 0.4 | 79.7 0.5 | 53.0 0.2 |
| FedAdam | 159.9 0.1 | 187.1 2.0 | 69.5 0.6 | 156.2 0.6 | 183.0 1.2 | 67.0 0.4 | 151.3 0.5 | 181.7 2.0 | 66.1 0.4 | 143.7 0.2 | 182.8 1.7 | 66.5 0.5 |
| FedAlign | 144.4 0.2 | 89.7 0.9 | 64.2 0.4 | 141.4 0.4 | 87.3 0.5 | 61.6 0.2 | 136.9 0.6 | 86.3 0.7 | 60.9 0.3 | 129.0 0.3 | 86.2 0.6 | 61.1 0.2 |
| mamba2 | hybrid75 | hybrid50 | hybrid50alt | hybrid25 | attn0 | |||||||
| SSM 1.00 | SSM 0.75 | SSM 0.50 | SSM 0.50 | SSM 0.25 | SSM 0.00 | |||||||
| Algorithm | Val | Last-10 | Val | Last-10 | Val | Last-10 | Val | Last-10 | Val | Last-10 | Val | Last-10 |
| FedAvg | 43.4 0.4 | 30.7 0.3 | 41.5 0.8 | 26.5 0.5 | 43.3 0.6 | 29.3 0.4 | 52.6 1.2 | 31.3 0.2 | 54.4 1.2 | 32.7 0.7 | 69.6 1.1 | 42.3 1.1 |
| FedProx | 41.9 0.4 | 25.4 0.2 | 39.8 0.4 | 21.0 0.2 | 41.4 0.3 | 23.6 0.2 | 50.4 0.7 | 24.7 0.4 | 52.2 1.3 | 26.8 0.9 | 67.7 1.0 | 34.8 0.9 |
| SCAFFOLD | 42.2 0.4 | 29.4 0.3 | 39.2 1.0 | 22.5 0.5 | 41.3 0.7 | 25.4 0.8 | 50.3 1.2 | 27.1 0.3 | 52.5 1.2 | 29.6 1.3 | 68.3 1.0 | 41.8 0.9 |
| FedAdam | 116.0 1.1 | 33.2 0.2 | 116.5 3.4 | 30.7 0.5 | 120.4 2.8 | 32.5 0.5 | 141.7 3.9 | 38.8 0.2 | 144.4 2.3 | 40.3 0.6 | 180.8 4.6 | 59.4 1.1 |
| Checkpoint | Model | epoch | epochs | epochs | Selected / |
| Mamba2-130M | Zero-shot (pretrained) | 25.7 0.0 | – | ||
| Specialists (own task) | 81.8 1.8 | 88.4 0.5 | 89.2 0.2 | – | |
| Specialists (other tasks) | 23.0 0.2 | 23.6 0.1 | 23.4 0.3 | – | |
| Weight averaging | 34.4 2.1 | 36.3 1.7 | 37.4 0.6 | – | |
| Task Arithmetic | 53.5 4.6 | 61.9 5.1 | 61.3 3.8 | / / | |
| TIES-Merging | 47.0 2.0 | 55.0 3.9 | 56.6 3.6 | / / | |
| Mamba2-130M | Mamba2-370M | Mamba2-780M | ||||||||||
| Model | AG News | DBpedia | TREC | IMDB | AG News | DBpedia | TREC | IMDB | AG News | DBpedia | TREC | IMDB |
| Zero-shot (pretrained) | 25.0 0.0 | 9.9 0.0 | 16.8 0.0 | 51.1 0.2 | 24.3 0.0 | 8.1 0.0 | 13.4 0.2 | 51.1 0.1 | 25.3 0.0 | 8.7 0.0 | 15.1 0.1 | 49.9 0.1 |
| Specialists (own task) | 89.5 0.3 | 96.4 0.2 | 86.7 0.6 | 83.9 0.4 | 90.8 0.2 | 97.6 0.2 | 88.7 0.5 | 86.0 0.1 | 90.8 0.1 | 97.9 0.2 | 88.4 1.0 | 86.1 0.5 |
| Weight averaging | 37.9 2.6 | 21.6 3.9 | 34.9 5.3 | 55.3 0.7 | 42.4 2.3 | 36.8 0.9 | 63.5 4.4 | 59.1 5.4 | 44.4 0.6 | 41.1 3.3 | 54.5 1.6 | 66.8 3.2 |
| Task Arithmetic | 64.7 8.5 | 53.3 13.8 | 61.7 7.3 | 65.5 2.6 | 75.3 7.0 | 74.4 2.8 | 65.1 2.8 | 72.5 2.2 | 80.6 2.3 | 65.9 7.9 | 57.9 7.7 | 81.1 1.7 |
| TIES-Merging | 72.2 8.4 | 51.7 17.9 | 46.2 13.1 | 56.2 0.1 | 77.3 6.8 | 73.4 3.6 | 62.9 3.0 | 69.9 0.8 | 83.4 1.4 | 69.5 8.2 | 52.4 11.6 | 83.1 1.1 |
| Checkpoint | Epochs | ||||||
| Mamba2-130M | 31.2 1.2 | 38.1 3.2 | 48.3 3.3 | 52.4 4.2 | 53.5 4.6 | 51.7 3.8 | |
| 32.6 1.1 | 41.3 1.9 | 55.5 2.0 | 61.9 4.1 | 61.9 5.1 | 55.8 5.4 | ||
| 32.8 0.8 | 42.7 0.8 | 55.8 2.1 | 61.3 3.8 | 57.2 6.2 | 49.3 7.0 | ||
| Mamba2-370M | 40.2 1.9 | 49.5 0.6 | 59.4 2.9 | 61.1 2.5 | 58.6 3.4 | 54.1 5.6 | |
| 43.5 1.5 | 54.7 0.4 | 67.6 1.0 | 71.0 1.7 | 68.2 2.2 | 59.6 2.7 | ||
| 44.5 1.5 | 55.4 0.5 | 68.6 0.5 | 71.8 1.2 | 66.3 2.5 | 54.0 3.3 |
| Checkpoint | |||||
| Mamba2-130M | 34.7 1.5 | 33.7 0.9 | 32.6 0.8 | 32.1 0.8 | |
| 46.1 1.3 | 44.0 1.1 | 42.2 1.1 | 41.2 1.4 | ||
| 51.7 1.6 | 49.6 1.7 | 47.9 1.9 | 46.9 1.3 | ||
| 56.6 3.6 | 54.7 2.6 | 53.7 2.3 | 52.9 1.8 | ||
| Mamba2-370M | 45.4 0.6 | 46.4 0.3 | 45.0 1.5 | 42.0 1.3 | |
| 63.0 1.7 | 62.6 0.8 | 60.9 1.1 | 57.5 1.3 |
| Specialists (own task) | Weight averaging | |||||
| epoch | epochs | epochs | epoch | epochs | epochs | |
| (reference) | 81.8 1.8 | 88.4 0.5 | 89.2 0.2 | 34.4 2.1 | 36.3 1.7 | 37.4 0.6 |
| 82.6 3.0 | 88.1 0.9 | 88.9 0.6 | 33.3 0.9 | 35.7 0.6 | 37.3 0.6 | |
| 82.5 3.5 | 87.6 0.3 | 88.1 0.3 | 32.5 2.0 | 33.1 1.1 | 34.0 0.8 | |
| 80.6 2.6 | 85.3 0.6 | 86.1 0.1 | 30.4 1.1 | 30.3 0.5 | 30.6 0.5 | |
| 75.3 3.1 | 79.7 0.9 | 81.9 0.4 | 27.1 0.7 | 27.6 0.1 | 28.0 0.1 | |