Organizations: School of Computer Science and Technology, Beijing Institute of Technology, No.5 Zhongguancun South Street, Beijing, 100081, China · School of Computing and Information Systems, Singapore Management University, 80 Stamford Road, 178902, Singapore
Large Language Models (LLMs) have shown strong potential in multivariate time series forecasting and anomaly detection. Existing studies predominantly inject temporal information into LLMs via direct numerical tokenization or heuristic textual descriptions. However, LLMs still face difficulty in perceiving the underlying structural patterns of numerical time series, particularly the seasonal and trend components obscured by discrete numerical tokens. To bridge this gap, we propose FreSia, a frequency-aware framework that establishes an effective alignment between the semantic space of LLMs and the frequency space of time series. Specifically, FGPrompt, a Frequency-Guided Prompt mechanism within FreSia, distills the frequency-domain structures of time series and projects them into prompts tailored to the semantic space of LLMs. Furthermore, we introduce a Global-driven Context Learning (GCL) component, which uses a global CLS-driven probe to generate global context to bridge the time-frequency domain gap and fuse the multi-modal information. Experiments on eight forecasting benchmarks show that FreSia achieves average improvements of 13.48% and 8.06% in MSE and MAE, respectively.
Figures & tables
Figure 1 : (a) Representation gap between raw numerical prompts and frequency-guided semantic prompts. (b) Top-K frequency components explain a large portion of spectral energy across eight benchmark datasets. This motivates the use of dominant frequency components as compact structural primitives for frequency-guided semantic instantiation.
Figure 2 : Framework of FreSia . ❶ Frequency-Guided Prompt (FGPrompt) injects frequency-domain information into prompts and the frequency pool. ❷ TSEncoder captures temporal dependencies and generates a CLS representation. ❸ Global-driven Context Learning (GCL) bridges the time-frequency domain gap and fuses multi-modal information.
Property
Numerical Range
Linguistic Descriptor
Trend ( k^ )
k^>0.5
strong upward trend
0.1<k^≤0.5
slight upward trend
−0.1≤k^≤0.1
stable trend
−0.5≤k^<−0.1
slight downward trend
k^<−0.5
strong downward trend
Volatility ( v )
v>1.0
high volatility
Table 1: Mapping rules for linguistic descriptors in Frequency-Guided Prompts.
Models
FreSia
LangTime
TimeCMA
Time-LLM
UniTime
Metrics
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
ETTm1
96
0.310
0.348
0.319
0.348
0.312
0.351
0.359
0.381
0.322
0.363
192
0.352
0.374
0.368
0.375
0.361
0.378
0.383
0.393
0.366
0.387
336
0.378
0.396
0.413
0.402
0.392
0.401
0.416
0.414
0.398
0.407
720
0.430
0.427
0.487
0.439
0.453
0.438
0.483
0.449
0.454
0.440
Avg
0.367
0.386
0.397
0.391
0.380
0.392
0.410
0.409
0.385
0.399
Table 2 : Results of multivariate time series forecasting. The input sequence length is 36 for the ILI dataset and 96 for others. The best results are in red and the second best are blue.
Models
FreSia
OFA
iTransformer
TimesNet
RLinear
Metrics
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
ETTm1
96
0.310
0.348
0.335
0.369
0.334
0.368
0.338
0.375
0.355
0.376
192
0.352
0.374
0.374
0.385
0.377
0.391
0.374
0.387
0.387
0.392
336
0.378
0.396
0.407
0.406
0.426
0.420
0.410
0.411
0.424
0.415
720
0.430
0.427
0.469
0.442
0.491
0.459
0.478
0.450
0.487
0.450
Avg
0.367
0.386
0.396
0.401
0.407
0.410
0.400
0.406
0.413
0.408
Table 3 : Results of multivariate time series forecasting. The input sequence length is 36 for the ILI dataset and 96 for others. The best results are in red and the second best are blue.
Dataset
Metric
FreSia
TSINR
FPT
TimesNet
Anomaly.
LightTS
DLinear
SMD
P
88.56
83.09
87.27
87.98
78.72
87.04
87.27
R
84.07
80.46
81.08
81.54
65.43
78.39
80.99
F1
86.25
81.76
84.06
84.64
71.46
82.49
84.01
PSM
P
97.61
99.21
98.55
98.51
98.76
98.29
98.66
R
96.22
89.37
95.79
96.27
83.25
93.60
94.70
F1
96.91
94.04
97.15
97.38
90.35
95.89
96.64
Table 4 : Results of anomaly detection. The input sequence length is 100. The best results are in red and the second best are blue.
Figure 3 : Sensitivity analysis of frequency pool size.
Datasets
ETTh2
ETTm2
ECL
Metrics
MSE
MAE
MSE
MAE
MSE
MAE
w/o LLM
0.357
0.388
0.288
0.331
0.186
0.271
w/o TSE
0.393
0.417
0.321
0.358
0.382
0.424
w/o PE
0.353
0.387
0.279
0.325
0.177
0.274
w/o CA
0.356
0.390
0.291
0.333
0.179
0.268
w/o CLS
0.356
0.388
0.282
0.327
0.179
0.276
Table 5 : Ablation studies of FreSia
Dataset
FreSia
TimeCMA-style
w/o Freq. Path.
w/o LLM
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
Weather
0.236
0.263
0.237
0.264
0.238
0.264
0.252
0.273
Exchange
0.339
0.394
0.383
0.417
0.392
0.420
0.404
0.425
ILI
1.563
0.779
1.577
0.775
1.601
0.782
1.696
0.787
Avg.
0.713
0.479
0.732
0.485
0.743
0.489
0.784
0.495
Table 6: Disentangling the effects of frequency-domain information and LLM semantic encoding.
Figure 4 : Effect of mixed prompting strategy on Exchange.
Figure 5 : Attention difference analysis of CLS-driven calibration.
Figure 6 : ETTh1 frequency pool analysis case1.
Figure 7 : ETTh1 frequency pool analysis case2.
Figure 8 : Frequency pool Gram matrix.
Figure 9 : Comparison with different losses.
Model
Test Time (s)
GPU Memory (MiB)
Params (M)
TimeCMA
16.78
2690
29.53
OFA
1697.66
916
3.92
iTransformer
129.98
2248
4.96
TimeKAN
44.32
3556
0.044
FreSia
11.29
2936
28.29
Table 7: Efficiency comparison. We report test time, peak GPU memory usage, and model parameters under the same evaluation setting.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Time Residual
Time-Freq. Residual
Freq. Residual
SMD
0.863
0.392
0.162
PSM
0.969
0.875
0.781
SWaT
0.943
0.913
0.846
MSL
0.856
0.171
0.080
SMAP
0.787
0.641
0.637
Avg.
0.884
0.598
0.501
Appendix
Table 1: F1 comparison of anomaly scoring rules. Time Residual denotes the standard reconstruction-error score used by FreSia. Time-Freq. Residual combines time-domain and frequency-domain residuals with equal weight. Freq. Residual uses only the frequency residual.
Models
FreSia
TimeKAN
PaiFilter
TexFilter
FredFormer
Metrics
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
ETTm1
96
0.310
0.348
0.322
0.361
0.318
0.358
0.321
0.361
0.328
0.363
192
0.352
0.374
0.357
0.383
0.364
0.383
0.367
0.387
0.367
0.382
336
0.378
0.396
0.383
0.402
0.396
0.406
0.401
0.409
0.395
0.403
720
0.429
0.427
0.446
0.436
0.456
0.444
0.477
0.448
0.454
0.440
Avg
0.367
0.386
0.377
0.395
0.384
0.398
0.392
0.401
0.386
0.397
Appendix
Table C.1 : Comparison with frequency-domain baselines. The input sequence length is 36 for the ILI dataset and 96 for others. The best results are in red and the second best are blue.
Variant
ETTm2
ECL
Avg.
MSE
MAE
MSE
MAE
MSE
MAE
Adaptive-Quantile
0.276
0.321
0.177
0.274
0.227
0.298
Continuous
0.274
0.320
0.178
0.274
0.226
0.297
Freq-only
0.276
0.322
0.179
0.276
0.228
0.299
FreSia
0.274
0.320
0.175
0.271
0.224
0.296
Appendix
Table D.1: Robustness to auxiliary trend/volatility mappings. The results are averaged over all horizons.