Short-horizon realized volatility forecasting requires the integration of market information that evolves at incompatible temporal resolutions, from second-level order book dynamics to weekly regime drift. Our conference work introduced HAN-T, a hierarchical architecture in which scale-specific Transformer encoders process short, mid, and long-horizon streams and a learned attention fuser weighs their contributions. This article replaces the quadratic attention encoders with selective state space (Mamba) encoders while retaining attention only in the fuser, where the input is a three-token set rather than a long sequence. The resulting hybrid, HAN-Mamba, summarizes each stream through a recurrent state whose input-dependent gating matches two structural properties of volatility: persistent but decaying memory and abrupt regime shifts. On the Optiver Realized Volatility Prediction benchmark under time-aware five-fold cross-validation, HAN-Mamba improves mean RMSPE over HAN-T (0.1942 vs. 0.1965) with 33% fewer parameters. Its linear-time encoders further allow the high-frequency context to be extended from 60 to 240 buckets, reducing error to 0.1927 where the attention variant saturates, and support constant-time streaming updates at inference. Ablations attribute the gains to the encoder swap, confirm that the hierarchical prior transfers across sequence-model families, and show that the permutation-invariant attention fuser remains the correct mechanism for cross-scale integration.
Figures & tables
Figure 1: The HAN-Mamba architecture. Adapted from the conference version of this work [ 14 ] : the scale-specific Transformer encoders are replaced by asymmetric selective state space encoders read out through their final recurrent state h(L) , while the three-token attention fuser, mean pooling, and regression head are retained.
Model
Fold 1
Fold 2
Fold 3
Fold 4
Fold 5
Mean
Std
GARCH(1,1) [ 1 a]
0.2753
0.2911
0.2805
0.3051
0.2745
0.2853
0.0118
LightGBM [ 1 a]
0.2150
0.2112
0.2046
0.2144
0.2206
0.2132
0.0053
LightGBM + Optuna [ 1 a]
0.2083
0.2049
0.2020
0.2097
0.2130
0.2076
0.0044
Flat Transformer [ 1 a]
0.2039
0.1972
0.1925
0.2027
0.1981
0.1989
0.0041
HAN-T [ 1 a]
0.1971
0.1939
0.1948
0.2009
0.1958
0.1965
0.0025
Flat Mamba
0.2012
0.1955
0.1938
0.2001
0.1949
0.1971
0.0033
Table 1: Forecasting performance, per-fold RMSPE under time-aware five-fold cross-validation, lower is better. Rows marked [ 1 a] are reproduced from the conference version of this work [ 14 ] . Best per column in bold
Model
Ls=60
Ls=120
Ls=240
HAN-T
0.1965 ± 0.0025
0.1959 ± 0.0027
0.1963 ± 0.0031
HAN-Mamba
0.1942 ± 0.0022
0.1931 ± 0.0021
0.1927 ± 0.0022
Table 2: Effect of the high-frequency context length Ls , mean ± std RMSPE over five folds. The Ls=60 entry for HAN-T is from the conference version [ 14 ]
Figure 2: Context-length scaling. The attention encoders saturate and then regress as Ls grows, while the selective state space encoders improve monotonically with diminishing returns
Group
Variant
Mean RMSPE
Fusion design
Attention fuser (HAN-Mamba default)
0.1942
Mamba fuser, order long → mid → short
0.1958
Mamba fuser, order short → mid → long
0.1963
Non-learnable average fuser
0.1966
Encoder placement
Mamba on all three streams (HAN-Mamba default)
0.1942
Mamba on short stream only
0.1949
Table 3: Ablations, mean RMSPE over five folds at Ls=60 . The HAN-T reference and its average-fuser variant are from the conference version [ 14 ]
GFLOPs
Epoch time (s)
Peak mem. (GB)
Model
Params
Ls=60
Ls=240
Ls=60
Ls=240
Ls=60
Ls=240
HAN-T
1.43M
0.16
0.67
41
152
5.8
18.9
HAN-Mamba
0.96M
0.06
0.23
33
78
4.1
8.7
Table 4: Efficiency at Ls=60 and Ls=240 . FLOPs are per forward sample. Peak memory is per GPU during training. Streaming latency is the time to update one prediction when a new 10-minute bucket arrives at inference
Multivariate time series forecasting is fundamental to numerous domains such as energy, finance, and environmental monitoring, where complex temporal dependencies and cross-variable interactions pose enduring challenges. Existing Transformer-based methods capture temporal correlations through attention mechanisms but suffer from quadratic computational cost, while state-space models like Mamba achieve efficient long-context modeling yet lack explicit temporal pattern recognition. Therefore we introduce UniMamba, a unified spatial-temporal forecasting framework that integrates efficient state-space dynamics with attention-based dependency learning. UniMamba employs a Mamba Variate-Channel Encoding Layer enhanced with FFT-Laplace Transform and TCN to capture global temporal dependencies, and a Spatial Temporal Attention Layer to jointly model inter-variate correlations and temporal evolution. A Feedforward Temporal Dynamics Layer further fuses continuous and discrete contexts for accurate forecasting. Comprehensive experiments on eight public benchmark datasets demonstrate that UniMamba consistently outperforms state-of-the-art forecasting models in both forecasting accuracy and computational efficiency, establishing a scalable and robust solution for long-sequence multivariate time-series prediction.
Xingsheng Chen, Xianpei Mu, Deyu Yi +6
School of Computing and Data Science, The University of Hong Kong, Hong Kong, China · School of Information Engineering, Beijing Institute of Graphic Communication, Beijing, China · Innovation Engineering College, Macau University of Science and Technology, Macau, China +1
Financial volatility is regime dependent, yet incorporating regime information into neural networks can also destabilize training. This paper asks where such information should enter a neural cross-sectional volatility forecasting model. We study five-day realized-volatility forecasts for 1,027 U.S. equities using a rolling walk-forward evaluation framework in which information, model capacity, hyperparameter tuning, and random seeds are matched across architectures. We propose RG-ResMoE, a regime-gated residual mixture-of-experts architecture in which regime information is used only for expert routing rather than for direct forecasting. The base predictor models volatility from stock features, while a gating network uses regime state variables to route residual corrections. RG-ResMoE consistently outperforms a capacity-matched MLP in both forecasting accuracy and training stability in the main U.S. study. Similar gains are observed on an independent Japanese panel. The integration pathway is decisive: appending the same regime variables directly to the forecasting input degrades both predictive performance and training stability, whereas restricting them to the routing gate improves accuracy and Value-at-Risk calibration. Hard routing consistently underperforms soft routing. The results suggest that, in compact neural volatility forecasting models, the primary value of mixture-of-experts models lies less in increasing model capacity than in controlling how nonstationary regime information influences prediction.
Junyi Ye, Gargi Vijay Borde
School of Computing, Montclair State University · Montclair, New Jersey, USA
In financial forecasting, predictive performance depends not only on which model is trained, but also on how the trained model is deployed. We study this issue in multi-horizon volatility forecasting. Our starting point is that a trained multi-output (MIMO) forecaster does not define a single deployable predictor: by changing the inference-time rollout rule, the same trained model induces a family of forecasts with different accuracy and cost profiles. Across 20 stock-volatility series, three forecast horizons, and architectures ranging from linear models to PatchTST, we find that non-default rollout rules often improve over standard MIMO deployment. However, the best fixed rule varies substantially across architectures and horizons, making any single static replacement unreliable. We therefore evaluate validation-based deployment policies over the induced rule family. Under the primary MSE objective, validation-selected singletons provide a low-cost improvement over default MIMO, while small rule subsets recover much of the benefit of larger ensembles at substantially lower inference cost. We also find that policy rankings are metric-sensitive: MSE-selected policies do not transfer uniformly to QLIKE, a finance-standard volatility loss. These results show that inference-time deployment is a meaningful source of adaptiveness in financial forecasting, and that trained volatility forecasters should be evaluated not only by their architecture, but also by their deployment policy.
Riku Green, Zahraa S. Abdallah, Telmo M Silva Filho