Organizations: School of Computer Science and Technology, Beijing Institute of Technology, No.5 Zhongguancun South Street, Beijing, 100081, China · School of Computing and Information Systems, Singapore Management University, 80 Stamford Road, 178902, Singapore
Large Language Models (LLMs) have shown strong potential in multivariate time series forecasting and anomaly detection. Existing studies predominantly inject temporal information into LLMs via direct numerical tokenization or heuristic textual descriptions. However, LLMs still face difficulty in perceiving the underlying structural patterns of numerical time series, particularly the seasonal and trend components obscured by discrete numerical tokens. To bridge this gap, we propose FreSia, a frequency-aware framework that establishes an effective alignment between the semantic space of LLMs and the frequency space of time series. Specifically, FGPrompt, a Frequency-Guided Prompt mechanism within FreSia, distills the frequency-domain structures of time series and projects them into prompts tailored to the semantic space of LLMs. Furthermore, we introduce a Global-driven Context Learning (GCL) component, which uses a global CLS-driven probe to generate global context to bridge the time-frequency domain gap and fuse the multi-modal information. Experiments on eight forecasting benchmarks show that FreSia achieves average improvements of 13.48% and 8.06% in MSE and MAE, respectively.
Figures & tables
Figure 1 : (a) Representation gap between raw numerical prompts and frequency-guided semantic prompts. (b) Top-K frequency components explain a large portion of spectral energy across eight benchmark datasets. This motivates the use of dominant frequency components as compact structural primitives for frequency-guided semantic instantiation.
Figure 2 : Framework of FreSia . ❶ Frequency-Guided Prompt (FGPrompt) injects frequency-domain information into prompts and the frequency pool. ❷ TSEncoder captures temporal dependencies and generates a CLS representation. ❸ Global-driven Context Learning (GCL) bridges the time-frequency domain gap and fuses multi-modal information.
Property
Numerical Range
Linguistic Descriptor
Trend ( k^ )
k^>0.5
strong upward trend
0.1<k^≤0.5
slight upward trend
−0.1≤k^≤0.1
stable trend
−0.5≤k^<−0.1
slight downward trend
k^<−0.5
strong downward trend
Volatility ( v )
v>1.0
high volatility
Table 1: Mapping rules for linguistic descriptors in Frequency-Guided Prompts.
Models
FreSia
LangTime
TimeCMA
Time-LLM
UniTime
Metrics
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
ETTm1
96
0.310
0.348
0.319
0.348
0.312
0.351
0.359
0.381
0.322
0.363
192
0.352
0.374
0.368
0.375
0.361
0.378
0.383
0.393
0.366
0.387
336
0.378
0.396
0.413
0.402
0.392
0.401
0.416
0.414
0.398
0.407
720
0.430
0.427
0.487
0.439
0.453
0.438
0.483
0.449
0.454
0.440
Avg
0.367
0.386
0.397
0.391
0.380
0.392
0.410
0.409
0.385
0.399
Table 2 : Results of multivariate time series forecasting. The input sequence length is 36 for the ILI dataset and 96 for others. The best results are in red and the second best are blue.
Models
FreSia
OFA
iTransformer
TimesNet
RLinear
Metrics
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
ETTm1
96
0.310
0.348
0.335
0.369
0.334
0.368
0.338
0.375
0.355
0.376
192
0.352
0.374
0.374
0.385
0.377
0.391
0.374
0.387
0.387
0.392
336
0.378
0.396
0.407
0.406
0.426
0.420
0.410
0.411
0.424
0.415
720
0.430
0.427
0.469
0.442
0.491
0.459
0.478
0.450
0.487
0.450
Avg
0.367
0.386
0.396
0.401
0.407
0.410
0.400
0.406
0.413
0.408
Table 3 : Results of multivariate time series forecasting. The input sequence length is 36 for the ILI dataset and 96 for others. The best results are in red and the second best are blue.
Dataset
Metric
FreSia
TSINR
FPT
TimesNet
Anomaly.
LightTS
DLinear
SMD
P
88.56
83.09
87.27
87.98
78.72
87.04
87.27
R
84.07
80.46
81.08
81.54
65.43
78.39
80.99
F1
86.25
81.76
84.06
84.64
71.46
82.49
84.01
PSM
P
97.61
99.21
98.55
98.51
98.76
98.29
98.66
R
96.22
89.37
95.79
96.27
83.25
93.60
94.70
F1
96.91
94.04
97.15
97.38
90.35
95.89
96.64
Table 4 : Results of anomaly detection. The input sequence length is 100. The best results are in red and the second best are blue.
Figure 3 : Sensitivity analysis of frequency pool size.
Datasets
ETTh2
ETTm2
ECL
Metrics
MSE
MAE
MSE
MAE
MSE
MAE
w/o LLM
0.357
0.388
0.288
0.331
0.186
0.271
w/o TSE
0.393
0.417
0.321
0.358
0.382
0.424
w/o PE
0.353
0.387
0.279
0.325
0.177
0.274
w/o CA
0.356
0.390
0.291
0.333
0.179
0.268
w/o CLS
0.356
0.388
0.282
0.327
0.179
0.276
Table 5 : Ablation studies of FreSia
Dataset
FreSia
TimeCMA-style
w/o Freq. Path.
w/o LLM
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
Weather
0.236
0.263
0.237
0.264
0.238
0.264
0.252
0.273
Exchange
0.339
0.394
0.383
0.417
0.392
0.420
0.404
0.425
ILI
1.563
0.779
1.577
0.775
1.601
0.782
1.696
0.787
Avg.
0.713
0.479
0.732
0.485
0.743
0.489
0.784
0.495
Table 6: Disentangling the effects of frequency-domain information and LLM semantic encoding.
Figure 4 : Effect of mixed prompting strategy on Exchange.
Figure 5 : Attention difference analysis of CLS-driven calibration.
Figure 6 : ETTh1 frequency pool analysis case1.
Figure 7 : ETTh1 frequency pool analysis case2.
Figure 8 : Frequency pool Gram matrix.
Figure 9 : Comparison with different losses.
Model
Test Time (s)
GPU Memory (MiB)
Params (M)
TimeCMA
16.78
2690
29.53
OFA
1697.66
916
3.92
iTransformer
129.98
2248
4.96
TimeKAN
44.32
3556
0.044
FreSia
11.29
2936
28.29
Table 7: Efficiency comparison. We report test time, peak GPU memory usage, and model parameters under the same evaluation setting.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Time Residual
Time-Freq. Residual
Freq. Residual
SMD
0.863
0.392
0.162
PSM
0.969
0.875
0.781
SWaT
0.943
0.913
0.846
MSL
0.856
0.171
0.080
SMAP
0.787
0.641
0.637
Avg.
0.884
0.598
0.501
Appendix
Table 1: F1 comparison of anomaly scoring rules. Time Residual denotes the standard reconstruction-error score used by FreSia. Time-Freq. Residual combines time-domain and frequency-domain residuals with equal weight. Freq. Residual uses only the frequency residual.
Models
FreSia
TimeKAN
PaiFilter
TexFilter
FredFormer
Metrics
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
ETTm1
96
0.310
0.348
0.322
0.361
0.318
0.358
0.321
0.361
0.328
0.363
192
0.352
0.374
0.357
0.383
0.364
0.383
0.367
0.387
0.367
0.382
336
0.378
0.396
0.383
0.402
0.396
0.406
0.401
0.409
0.395
0.403
720
0.429
0.427
0.446
0.436
0.456
0.444
0.477
0.448
0.454
0.440
Avg
0.367
0.386
0.377
0.395
0.384
0.398
0.392
0.401
0.386
0.397
Appendix
Table C.1 : Comparison with frequency-domain baselines. The input sequence length is 36 for the ILI dataset and 96 for others. The best results are in red and the second best are blue.
Variant
ETTm2
ECL
Avg.
MSE
MAE
MSE
MAE
MSE
MAE
Adaptive-Quantile
0.276
0.321
0.177
0.274
0.227
0.298
Continuous
0.274
0.320
0.178
0.274
0.226
0.297
Freq-only
0.276
0.322
0.179
0.276
0.228
0.299
FreSia
0.274
0.320
0.175
0.271
0.224
0.296
Appendix
Table D.1: Robustness to auxiliary trend/volatility mappings. The results are averaged over all horizons.
Recent advances in Large Language Models (LLMs) have spurred cross-modal solutions for time-series forecasting. However, existing methods rely heavily on textual prompts for modality alignment-introducing nontrivial computational overhead and failing to leverage the rich spectral dynamics inherent in time-series data. To enable prompt-free, frequency-aware adaptation of frozen LLMs, we propose FM-LLM (Frequency-Enhanced Mixture-of-Experts for adapting LLMs to Time Series Forecasting), an autoregressive framework grounded in constrained asymmetric coupling. A Fourier Analysis Network (FAN)-based spectral token aligner injects structured harmonic representations directly into the frozen LLM with numerical compatibility. An asymmetric Mixture-of-Experts (MoE) decoder enforces role separation: shared experts with lightweight FAN layers reconstruct the global periodic backbone, while routed experts-restricted to standard FFNs-specialize in modeling non-periodic residual dynamics. A time-frequency hybrid loss function jointly optimizes temporal accuracy and spectral consistency, mitigating error accumulation during long-horizon autoregressive rollouts. Evaluated across eleven public benchmarks, FM-LLM achieves state-of-the-art performance on 59 out of 78 evaluation metrics. Compared to the strongest autoregressive LLM-based baseline, it delivers average improvements of 5.3% in MSE and 5.6% in MAE, with maximum gains reaching 8.0% for MSE and 8.4% for MAE. FM-LLM also demonstrates robust transferability, maintaining superior performance in 10% few-shot and zero-shot forecasting scenarios.
Rentao Gu, Yihang Ding, Junjie Li +5
Beijing University of Posts and Telecommunications, Beijing, 100876, China · China Telecom Research Institute, Beijing, 100032, China
Recent advances in Large Language Models (LLMs) have opened new possibilities for time series forecasting by enabling alignment between temporal patterns and pretrained word embeddings. However, most LLM-based methods overlook the heterogeneous nature of time series, where dynamic fluctuations and invariant semantics are entangled. This entanglement introduces spurious correlations during the alignment, as dynamic components act as confounders by simultaneously influencing invariant components and the resulting aligned embeddings. To address this issue, a variable-level alignment framework CVAformer is proposed. CVAformer explicitly disentangles each variable into invariant and dynamic components just before alignment, and applies causal intervention to mitigate the confounding effect of the dynamics. To better support variable-level alignment, CVAformer replaces the standard causal attention in LLMs with a non-causal attention mechanism that captures interactions among variables at each time step. Extensive experiments across long-term, short-term, few-shot, and zero-shot forecasting settings indicate that CVAformer matches or exceeds state-of-the-art performance on most datasets, and in some cases achieves notably better accuracy. Experimental results validate the effectiveness of variable-level alignment and dynamic disentanglement in CVAformer, offering a new perspective for LLM-based time series tasks.
Cross-domain multimodal time series forecasting is a challenging task, requiring models to integrate precise numerical comprehension, cross-domain semantic understanding, and effective multimodal fusion. Existing approaches either build Time Series Foundation Models (TSFMs) from scratch or leverage pretrained Large Language Models (LLMs). However, TSFMs often overlook semantic understanding and lack the ability to perform future-oriented semantic reasoning, and LLMs struggle with numerical comprehension and accurate quantitative forecasting. To overcome these limitations, we propose KairosAgent, a novel agentic framework for multimodal time series forecasting, including an LLM-based reasoner and a TSFM-based forecaster. KairosAgent unifies textual reasoning and numerical forecasting by dynamically invoking analytical tools to enhance the numerical understanding and semantic reasoning capabilities of LLMs. The reasoning results are subsequently fused into the TSFM pipeline, enabling more accurate and reliable future predictions. To further improve the reasoning, we curate a large-scale corpus of high-quality trajectories, alongside a reinforcement learning from forecasting paradigm with multi-turn refinement and turn-level credit assignment. Experiments demonstrate that KairosAgent achieves superior zero-shot forecasting performance while maximizing the utility of pretrained LLMs and TSFMs, presenting a promising direction for efficient and interpretable time series agents. The project page is at https://foundation-model-research.github.io/KairosAgent .