Multivariate time series forecasting (MTSF) is critical across many real-world domains. Existing deep learning approaches fall into two paradigms with distinct limitations: channel-independent (CI) methods unconditionally ignore cross-variable dependencies and model only temporal dynamics, while channel-dependent (CD) methods consider both but typically rely on architectural compromises to mitigate overfitting and computational overhead. We therefore propose Chameleon, a specialized CD state space model (SSM) that enables data-dependent, fine-grained interactions across variables while scaling linearly with their number. By connecting selective SSMs with the Kalman filter, we leverage the missing measurement update in the former for cross-variable modeling while preserving the SSM backbone for robust temporal modeling. We further identify favorable inductive biases of GatedDeltaNet for time series, adapt it as our backbone, and improve generalization through additional techniques, including a previously unexplored stochastic perturbation of reversible instance normalization. On strongly dependent ODE and PEMS datasets, Chameleon achieves the best MSE and MAE across all settings, while its CI ablation and prior CD methods incur 61-178% higher MSE on average. Across 28 standard benchmark settings, Chameleon also achieves better MSE and MAE than each baseline in at least 27 and 22 cases, respectively. Training-time and peak-memory analyses on Traffic and ETT further demonstrate competitive efficiency and favorable memory scalability across different variable counts.
Figures & tables
Figure 1: Overview of Chameleon, which preserves the recurrent GatedDeltaNet backbone for temporal modeling and introduces a Kalman-filter-inspired measurement update for cross-variable modeling through data-dependent observation signals and covariances (cov.).
Model
Chameleon
Chameleon (CI)
S-Mamba+ACN
TimeFilter
DUET
Crossformer
Dataset
(F)
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
Double Pendulum
96
0.0678
0.1176
0.1258
0.1880
0.0927
0.1642
0.0870
0.1616
0.2510
0.3271
0.1208
0.1975
192
0.3053
0.3216
0.3845
0.3836
0.3325
0.3701
0.3522
0.3787
0.5044
0.5052
0.4774
0.4717
Lorenz Coupled
48
0.0179
0.0389
0.0493
0.0761
0.0708
0.1247
0.0424
0.0934
0.1570
0.2137
0.0581
0.1152
96
0.1569
0.1842
0.3368
0.3213
0.3666
0.3837
0.2935
0.3230
0.5026
0.4758
0.3019
0.3299
PEMS03
48
0.0905
0.1920
0.1045
0.1991
0.0927
0.1968
0.1048
0.2099
0.1028
0.2050
0.0968
0.1989
Table 1: Results on strongly dependent datasets. Chameleon (CI) removes the proposed cross-variable measurement update. Best and second-best results are highlighted. Underlining denotes statistical significance over all other baselines ( p<0.05 ). The Avg./Max. increase rows report each baseline’s average/maximum relative error increase over Chameleon across all settings, with larger values indicating worse performance.
Model
Chameleon
S-Mamba+ACN
TimeFilter
DUET
SFNN
FITS
Dataset
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
Traffic
0.3636
0.2409
0.4012
0.2744
0.3738
0.2580
0.3868
0.2548
0.3836
0.2656
0.4074
0.2783
Solar
0.1783
0.2204
0.1936
0.2531
0.2020
0.2438
0.1884
0.2150
0.1831
0.2434
0.2086
0.2495
Electricity
0.1522
0.2418
0.1590
0.2576
0.1573
0.2541
0.1575
0.2469
0.1554
0.2452
0.1609
0.2552
ETTh1
0.3870
0.4168
0.4484
0.4538
0.4477
0.4431
0.4078
0.4214
0.3964
0.4208
0.4087
0.4207
ETTh2
0.3310
0.3773
0.3850
0.4074
0.3928
0.4155
0.3440
0.3864
0.3481
0.3832
0.3331
0.3820
Table 2: Results on long-horizon benchmarks, averaged across four forecasting horizons. Highlighting and Avg./Max. increase conventions follow Table 1 .
Dataset
Chameleon
GatedDeltaNet
Chameleon (CI)
w/o RevIN Pert
w/o Decomp
w/ Mamba2
PEMS08
0.1141
+26.8%
+28.5%
+2.16%
+1.09%
+0.38%
Traffic
0.3636
+2.47%
+2.31%
+2.67%
+2.53%
+1.82%
Solar
0.1783
+7.20%
+3.89%
+1.07%
+3.05%
+3.76%
ETTh2
0.3310
+0.77%
+0.55%
+0.77%
+0.80%
+0.85%
ETTm2
0.2356
+1.00%
+0.04%
+0.77%
+0.27%
+0.83%
Table 3: Ablation studies across five datasets. Chameleon reports the average MSE across horizons, while the other columns report the relative MSE increase over Chameleon. Chameleon (CI) deactivates the cross-variable measurement update, w/o RevIN Pert disables the stochastic RevIN perturbation, and w/o Decomp removes decomposition. GatedDeltaNet denotes the vanilla backbone without these three designs. w/ Mamba2 replaces the GatedDeltaNet backbone with Mamba2.
Figure 2: Learned effective Kalman gains ωKl on ETTh2, Traffic, and Double Pendulum. Larger off-diagonal magnitudes indicate stronger interactions across variables, while ω∈(0,1) is a learned gate controlling the reliance on the raw Kalman gain Kl . The same Double Pendulum patch yields different gains when serving as the first and last patch in two temporally offset input windows.
Figure 3: Efficiency and scalability analyses on Traffic ( F=96 ) with a batch size of one.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Description
Double Pendulum
Simulated double pendulum dynamics with angles and angular velocities.
( Gilpin, 2021 )
Lorenz Coupled
Two coupled simulated Lorenz systems.
( Gilpin, 2021 )
PEMS
Traffic flow collected from hundreds of sensors in the Caltrans PeMS.
( Song et al., 2020 )
Traffic
Road occupancy rates from 862 sensors on Bay Area freeways.
( Lai et al., 2018 )
Solar
Solar power production from 137 photovoltaic plants in Alabama.
( Lai et al., 2018 )
Electricity
Electricity consumption of 321 clients recorded in kWh from 2012 to 2014.
( Lai et al., 2018 )
Appendix
Table 4: Dataset descriptions. The last four are standard long-horizon benchmarks.
Dataset
Total Length
Split
Freq.
Var.
Look-back Window Size ( T )
Horizon ( F )
Double Pendulum
60,000
7:1:2
N/A
4
{96, 192, 336}
{96, 192}
Lorenz Coupled
60,000
7:1:2
N/A
6
{96, 192, 336}
{48, 96}
PEMS03
26,208
6:2:2
5 min
358
{96, 288}
{48, 96}
PEMS08
17,856
6:2:2
5 min
170
{96, 288}
{48, 96}
Traffic
17,544
7:1:2
1 hour
862
{168, 336, 672, 1344}
{96, 192, 336, 720}
Solar
52,560
7:1:2
10 min
137
{144, 288, 576, 1008}
{96, 192, 336, 720}
Appendix
Table 5: Experimental setup. Split: train/validation/test ratio by time; Var.: number of variables; Freq.: sampling frequency.
Dataset
Trend
Mean Shift
Std. Shift
KS
Double Pendulum
0.058
0.093
1.105
0.074
Lorenz Coupled
0.089
0.076
1.010
0.043
PEMS03
0.471
0.213
1.173
0.174
PEMS08
0.815
0.467
1.283
0.270
Traffic
0.782
0.507
1.604
0.303
Solar
0.222
0.056
1.058
0.072
Appendix
Table 6: Trend and distribution-shift diagnostics computed from the training and validation splits. All values report the 95th percentile across variables. Values exceeding the selection thresholds (Trend/Mean Shift/KS >0.2 ; Std. Shift >1.2 ) are highlighted.
Dataset
ETTh2
ETTm2
Solar
Double Pendulum
RevIN
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
Yes
0.2628
0.3271
0.1528
0.2409
0.1534
0.1981
0.1158
0.1753
No
0.2968
0.3623
0.1695
0.2671
0.1725
0.2110
0.0678
0.1176
Appendix
Table 7: RevIN ablation results at F=96 . Better results in each column are highlighted.
Figure 4: Traffic periodicity analysis, identifying 168 time steps (one week) as the strongest period.
Figure 5: PEMS08 periodicity analysis, identifying 288 time steps (one day) as the strongest period.
Figure 6: Double Pendulum periodicity analysis. No clear dominant period is identified.
Figure 7: Lorenz Coupled periodicity analysis. No clear dominant period is identified, with even weaker periodicity than Double Pendulum.
Model
Chameleon
S-Mamba+ACN
TimeFilter
DUET
SFNN
FITS
Dataset
(F)
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
Traffic
96
± 0.3309
± 0.2247
± 0.3653
± 0.2563
± 0.3428
± 0.2412
± 0.3493
± 0.2367
± 0.3484
± 0.2478
± 0.3820
± 0.2674
± 0.0018
± 0.0021
± 0.0009
± 0.0007
± 0.0003
± 0.0003
± 0.0008
± 0.0003
± 0.0002
± 0.0003
± 0.0008
± 0.0011
192
± 0.3485
± 0.2328
± 0.3755
± 0.2664
± 0.3609
± 0.2505
± 0.3691
± 0.2456
± 0.3697
± 0.2582
± 0.3944
± 0.2713
± 0.0026
± 0.0014
± 0.0029
± 0.0019
± 0.0003
± 0.0002
± 0.0010
± 0.0002
± 0.0031
± 0.0013
± 0.0003
± 0.0003
336
± 0.3651
± 0.2418
± 0.4055
± 0.2746
± 0.3775
± 0.2600
± 0.3922
± 0.2592
± 0.3853
± 0.2661
± 0.4077
± 0.2777
Appendix
Table 8: Detailed experimental results (mean ± std) on long-horizon benchmarks. Highlighting and underlining follow Table 1 .
Model
Chameleon
Chameleon (CI)
S-Mamba+ACN
TimeFilter
DUET
Crossformer
Dataset
(F)
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
Double Pendulum
96
± 0.0678
± 0.1176
± 0.1258
± 0.1880
± 0.0927
± 0.1642
± 0.0870
± 0.1616
± 0.2510
± 0.3271
± 0.1208
± 0.1975
± 0.0029
± 0.0026
± 0.0043
± 0.0040
± 0.0019
± 0.0013
± 0.0097
± 0.0101
± 0.0053
± 0.0046
± 0.0074
± 0.0039
192
± 0.3053
± 0.3216
± 0.3845
± 0.3836
± 0.3325
± 0.3701
± 0.3522
± 0.3787
± 0.5044
± 0.5052
± 0.4774
± 0.4717
± 0.0051
± 0.0033
± 0.0035
± 0.0027
± 0.0032
± 0.0024
± 0.0144
± 0.0123
± 0.0050
± 0.0036
± 0.0209
± 0.0217
Lorenz Coupled
48
± 0.0179
± 0.0389
± 0.0493
± 0.0761
± 0.0708
± 0.1247
± 0.0424
± 0.0934
± 0.1570
± 0.2137
± 0.0581
± 0.1152
Appendix
Table 9: Detailed experimental results (mean ± std) on strongly dependent datasets. Highlighting and underlining follow Table 1 .
Figure 8: Efficiency and scalability analyses on ETT at F=96 with a batch size of one.
Figure 9: Additional parameter-count analyses at F=96 with a batch size of one.
Multivariate time series forecasting is fundamental to numerous domains such as energy, finance, and environmental monitoring, where complex temporal dependencies and cross-variable interactions pose enduring challenges. Existing Transformer-based methods capture temporal correlations through attention mechanisms but suffer from quadratic computational cost, while state-space models like Mamba achieve efficient long-context modeling yet lack explicit temporal pattern recognition. Therefore we introduce UniMamba, a unified spatial-temporal forecasting framework that integrates efficient state-space dynamics with attention-based dependency learning. UniMamba employs a Mamba Variate-Channel Encoding Layer enhanced with FFT-Laplace Transform and TCN to capture global temporal dependencies, and a Spatial Temporal Attention Layer to jointly model inter-variate correlations and temporal evolution. A Feedforward Temporal Dynamics Layer further fuses continuous and discrete contexts for accurate forecasting. Comprehensive experiments on eight public benchmark datasets demonstrate that UniMamba consistently outperforms state-of-the-art forecasting models in both forecasting accuracy and computational efficiency, establishing a scalable and robust solution for long-sequence multivariate time-series prediction.
Xingsheng Chen, Xianpei Mu, Deyu Yi +6
School of Computing and Data Science, The University of Hong Kong, Hong Kong, China · School of Information Engineering, Beijing Institute of Graphic Communication, Beijing, China · Innovation Engineering College, Macau University of Science and Technology, Macau, China +1
Multivariate time series forecasting is critical in many real-world systems, and thus modeling cross-channel dependencies is essential. Although existing methods improve overall accuracy by enhancing representations and cross-channel interactions, it remains challenging to reliably capture inter-variable dependencies under specific conditions. We observe that dependencies in real data are often state-dependent and noisy; in such cases, dense interactions can amplify spurious correlations and lead to representation over-smoothing, which may yield unreliable predictions in certain scenarios. Motivated by this, we propose MS-FLOW, a sparse-bottleneck framework that explicitly models inter-variable interaction as capacity-limited information flow. Specifically, MS-FLOW replaces fully connected communication with selective sparse routing, retaining only a few critical dependency paths and injecting cross-variable signals under a strict communication budget, thereby suppressing redundant connections and spurious-correlation propagation. Extensive experiments demonstrate that MS-FLOW learns more reliable multivariate correlations, achieving state-of-the-art forecasting accuracy on 12 real-world benchmarks while producing fewer yet more reliable dependencies, shifting multivariate forecasting from "more interaction" to "more effective interaction".
Fan Zhang, Shiming Fan, Hua Wang
Shandong Technology and Business University, Yantai, Shandong, China · Ludong University, Yantai, Shandong, China
Multivariate time-series analysis involves extracting informative representations from sequences of multiple interdependent variables, supporting tasks such as forecasting, imputation, and anomaly detection. In real-world scenarios, these variables are typically collected from a shared context or underlying phenomenon, suggesting the presence of latent dependencies across time and channels that can be leveraged to improve performance. However, recent findings show that channel-independent (CI) models, which assume no inter-variable dependencies, often outperform channel-dependent (CD) models that explicitly model such relationships. This surprising result indicates that current CD models may not fully exploit their potential due to limitations in how dependencies are captured. Recent studies have revisited channel dependence modeling with various approaches; however, these methods often employ indirect modeling strategies, which can lead to meaningful dependencies being overlooked. To address this issue, we introduce XCTFormer, a transformer-based channel-dependent (CD) model that explicitly captures cross-temporal and cross-channel dependencies via an enhanced attention mechanism. The model operates in a token-to-token fashion, modeling pairwise dependencies between every pair of tokens across time and channels. The architecture comprises (i) a data processing module, (ii) a novel Cross-Relational Attention Block (CRAB) that increases capacity and expressiveness, and (iii) an optional Dependency Compression Plugin (DeCoP) that improves scalability. Through extensive experiments on three time-series benchmarks, we show that XCTFormer achieves strong results compared to widely recognized baselines; in particular, it attains state-of-the-art performance on the imputation task, outperforming the second-best method by an average of 20.8% in MSE and 15.3% in MAE.
Israel Zexer, Omri Azencot
The Stein Faculty of Computer and Information Science Ben-Gurion University of the Negev