Aligning Inductive Bias for Data-Efficient Generalization in State Space Models
Organizations: School of Computer Science, National Key Laboratory for Multimedia Information Processing, Peking University, China · School of Physics, Peking University, China
Abstract
The remarkable success of modern AI has been closely tied to scaling laws, yet the finite supply of high-quality data makes data efficiency--learning more from less--an increasingly important frontier. A model's inductive bias is a critical lever for data efficiency, but foundational sequence models such as State Space Models (SSMs) often rely on fixed, task-agnostic biases. When this fixed prior is misaligned with the underlying structure of a task, the model may require additional samples to overcome its own bias before learning the relevant signal. In this work, we introduce a principled framework for understanding and aligning the inductive bias of linear time-invariant SSMs. We first formalize this bias through an SSM-induced kernel and show theoretically and empirically that its spectrum is governed by the model's frequency response. This characterization motivates Task-Dependent Initialization (TDI), a fast power-spectrum matching method that aligns the initial SSM bias with the task's spectral characteristics before downstream training. Across controlled synthetic experiments, trainable one-layer SSMs, and deep SSMs on diverse real-world benchmarks, TDI can improve data-efficient generalization primarily when task-relevant spectral structure is present and the default SSM bias is spectrally mismatched. Our results provide both a theoretical lens and a practical tool for task-adaptive inductive bias, suggesting a path toward more data-efficient sequence modeling.
Figures & tables
| Low-data regime | High-data regime | |||||||
|---|---|---|---|---|---|---|---|---|
| Dataset | Base Acc. | TDI Acc. | Base Loss | TDI Loss | Base Acc. | TDI Acc. | Base Loss | TDI Loss |
| Binary Freq. | % | % | % | % | ||||
| ECG5000 | % | % | % | % | ||||
| FordA | % | % | % | % | ||||
| FordB | % | % | % | % | ||||
| sMNIST | % | % | % | % | ||||
| Low-data regime | High-data regime | |||||||
|---|---|---|---|---|---|---|---|---|
| Dataset | Base Acc. | TDI Acc. | Base Loss | TDI Loss | Base Acc. | TDI Acc. | Base Loss | TDI Loss |
| CIFAR-10 | 33.9 15.7 % | 35.4 14.5 % | 2.026 0.190 | 2.012 0.184 | 73.7 7.2 % | 69.4 8.5 % | 0.948 0.344 | 1.094 0.406 |
| Freq. Classification | 88.1 18.9 % | 91.1 11.5 % | 0.385 0.590 | 0.331 0.439 | 99.2 0.4 % | 98.9 0.5 % | 0.024 0.014 | 0.026 0.012 |
| ListOps | 16.6 7.1 % | 16.2 6.6 % | 5.419 0.754 | 5.597 0.941 | 44.7 6.9 % | 43.2 6.2 % | 1.928 0.786 | 1.816 0.584 |
| sMNIST | 56.3 42.2 % | 67.4 38.1 % | 1.145 1.084 | 0.897 0.991 | 99.1 0.2 % | 99.0 0.2 % | 0.034 0.011 | 0.038 0.011 |
| Pathfinder | 51.4 2.4 % | 51.7 2.6 % | 3.482 0.853 | 3.343 0.864 | 66.6 10.4 % | 68.4 8.1 % | 0.935 0.567 | 0.891 0.407 |
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | First stable ratio | Observation |
|---|---|---|
| Freq. Cls. | Identifiable transition | |
| sMNIST | Identifiable transition | |
| Speech Commands | Identifiable transition | |
| SC10 | Identifiable transition | |
| CIFAR-10 | Negative TDI gains on both sides of the transition | |
| PathFinder | Not reached | No transition identified on the evaluated grid |
| Experiment | SSM parameterization | TDI variant |
|---|---|---|
| Kernel regression | Complex diagonal SSM | Peak initialization + gradient refinement (Algorithm 3 ) |
| One-layer RNN | Real-valued matrix SSM | Peak initialization + gradient refinement (Algorithm 4 ) |
| Deep SSM | Complex diagonal SSM | Direct construction (Algorithm 3 ) |
| Dataset | Train | Test | Classes | Seq. len. | Input dim. |
|---|---|---|---|---|---|
| Binary Freq. | 40000 | 10000 | 2 | 256 | 1 |
| ECG5000 | 500 | 4500 | 5 | 140 | 1 |
| FordA | 3601 | 1320 | 2 | 500 | 1 |
| FordB | 3636 | 810 | 2 | 500 | 1 |
| MNIST | 60000 | 10000 | 10 | 784 | 1 |
| TwoPatterns | 1000 | 4000 | 4 | 128 | 1 |
| Dataset | Epochs |
|---|---|
| Binary Freq. | 15 |
| ECG5000 | 15 |
| FordA | 10 |
| FordB | 15 |
| MNIST | 15 |
| TwoPatterns | 15 |
| Dataset | Train | Val | Test | Classes | Seq. len. | Input dim. |
|---|---|---|---|---|---|---|
| Freq. Classification | 10000 | 1000 | 1000 | 10 | 1024 | 1 |
| sMNIST | 54000 | 6000 | 10000 | 10 | 784 | 1 |
| pMNIST | 54000 | 6000 | 10000 | 10 | 784 | 1 |
| PathFinder | 480001 | 59999 | 59999 | 2 | 1024 | 1 |
| CIFAR-10 | 45000 | 5000 | 10000 | 10 | 1024 | 1 |
| ListOps | 96000 | 2000 | 2000 | 10 | 2048 | – |
| Dataset | Depth | Features | State Size | Norm | Pre-norm | Dropout | LR | Batch Size | Epochs | WD | |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Freq. Classification | 2 | 64 | 64 | BN | True | 0 | 0.005 | 64 | 10 | 0.01 | |
| sMNIST | 4 | 256 | 64 | LN | False | 0.1 | 0.01 | 128 | 30 | 0.05 | |
| pMNIST | 4 | 256 | 64 | LN | False | 0.1 | 0.01 | 128 | 30 | 0.05 | |
| CIFAR-10 | 4 | 256 | 64 | LN | False | 0.1 | 0.01 | 128 | 30 | 0.05 | |
| Speech Commands | 4 | 256 | 64 | BN | False | 0 | 0.01 | 16 | 30 | 0.05 | |
| SC10 | 4 | 256 | 64 | BN | False | 0 | 0.01 | 16 | 30 | 0.05 |
| Final test accuracy | First-5 test accuracy | ||||||
|---|---|---|---|---|---|---|---|
| Dataset | Ratio | Std. RNN | Baseline | TDI | Std. RNN | Baseline | TDI |
| Binary Freq. | 0.01 | % | % | % | % | % | % |
| Binary Freq. | 0.02 | % | % | % | % | % | % |
| Binary Freq. | 0.04 | % | % | % | % | % | % |
| Binary Freq. | 0.08 | % | % | % | % | % | % |
| Binary Freq. | 0.16 | % | % | % | % | % | % |
| Final test accuracy | First-5 test accuracy | ||||||
|---|---|---|---|---|---|---|---|
| Dataset | Ratio | Baseline | TDI | TDI-Frozen | Baseline | TDI | TDI-Frozen |
| CIFAR-10 | 0.005 | % | % | % | % | % | % |
| CIFAR-10 | 0.010 | % | % | % | % | % | % |
| CIFAR-10 | 0.022 | % | % | % | % | % | % |
| CIFAR-10 | 0.046 | % | % | % | % | % | % |
| CIFAR-10 | 0.100 | % | % | % | % | % | % |
| Time-to-80% (epochs) | Early-10 (%) | Final Acc. | |||||
| Task | Train ratio | Epochs | Baseline | TDI | Baseline | TDI | (pp) |
| sMNIST | 2.15% | 100 | 34.2 ± 4.5 | 20.0 ± 4.6 | 10.19 | 14.15 | +0.12 ± 0.47 |
| sMNIST | 21.5% | 100 | 4.6 ± 0.9 | 3.0 ± 0.0 | 69.71 | 83.69 | −0.08 ± 0.07 |
| Freq. Cls. | 2.15% | 50 | 11.6 ± 1.7 | 9.6 ± 4.0 | 32.72 | 63.25 | −0.69 ± 0.73 |
| Freq. Cls. | 10.0% | 50 | 2.0 ± 0.0 | 1.2 ± 0.4 | 92.04 | 96.03 | +0.29 ± 0.29 |
| Freq. Cls. | 21.5% | 50 | 1.0 ± 0.0 | 1.0 ± 0.0 | 98.04 | 97.93 | +0.13 ± 0.15 |
| Dataset | Train ratio | First-layer TDI | All-layer TDI | (pp) |
|---|---|---|---|---|
| sMNIST | 0.464% | 10.47 ± 0.85 | 30.39 ± 14.76 | +19.92 ± 14.81 |
| sMNIST | 1.000% | 34.97 ± 23.93 | 80.41 ± 5.61 | +45.45 ± 23.50 |
| sMNIST | 2.154% | 91.64 ± 3.74 | 95.22 ± 1.28 | +3.58 ± 4.15 |
| CIFAR-10 | 0.464% | 15.85 ± 1.36 | 17.79 ± 1.65 | +1.94 ± 2.35 |
| CIFAR-10 | 1.000% | 22.63 ± 2.37 | 27.58 ± 3.32 | +4.95 ± 3.43 |
| CIFAR-10 | 2.154% | 37.03 ± 1.92 | 33.93 ± 0.97 | −3.09 ± 2.20 |
| Dataset | Train ratio | TDI-True | TDI-Permuted | Baseline |
|---|---|---|---|---|
| sMNIST | 2.15% | 20.0 ± 4.6 | 28.6 ± 8.7 | 34.2 ± 4.5 |
| sMNIST | 4.64% | 11.4 ± 2.6 | 13.4 ± 2.9 | 15.4 ± 6.3 |
| sMNIST | 10% | 5.2 ± 1.3 | 7.0 ± 2.1 | 10.6 ± 1.8 |
| Freq. Cls. | 2.15% | 9.6 ± 4.0 | 21.6 ± 13.8 | 11.6 ± 1.7 |
| Freq. Cls. | 4.64% | 2.8 ± 1.6 | 7.0 ± 3.7 | 4.2 ± 0.8 |
| Freq. Cls. | 10% | 1.2 ± 0.4 | 3.8 ± 1.9 | 2.0 ± 0.0 |
| Asset | Use in this paper | License / Terms |
|---|---|---|
| MNIST [ 17 ] | One-layer SSMs and Deep SSMs | MIT License |
| CIFAR-10 [ 29 ] | Deep SSMs | Official source |
| UCR Time Series Archive [ 16 ] | One-layer SSMs | Official source |
| PathFinder [ 43 ] | Deep SSMs | Apache-2.0 License |
| ListOps [ 32 , 43 ] | Deep SSMs | Apache-2.0 License |
| Speech Commands / SC10 [ 44 ] | Deep SSMs | CC-BY 4.0 License |