Organizations: School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences, Beijing 100049, China · State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, Beijing 100190, China · School of Intelligent Systems Engineering, Sun Yat-sen University, Shenzhen, Guangdong 518107, China · School of Computer Science and Technology, Huazhong University of Science and Technology, Wuhan, Hubei 430074, China · School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing 100049, China · College of Computer Science, Chongqing University, Chongqing 401331, China
Media-bridged time series forecasting is expanding to encompass traditional "multivariate" and emerging "multimodal" (e.g., through textual assistance). Existing Time Series Forecasting (TSF) models still rely on paradigm-specific relation, fusion, and temporal modules, hindering a common forecasting backbone across numerical and pre-aligned narrative-flow settings. To explore this, we propose the Multimedia Identity-Aware Prism Network (MIDAPN), a unified spatiotemporal forecasting backbone based on media-general graph adaptation and automatic temporal learning: (1) Following media pre-alignment, our Multimedia Identity-Aware Graph (MIDAG) revisits identity through static essence, dynamic behavior, and latent commonality, inducing affinities that extend variable-specific dependencies across media. Contextual Identity Modulation (CIM) further refines discriminative aggregation. (2) We develop Spectral Prism Convolution (SPConv) to automatically perform hierarchical temporal analysis, balancing coarse trends and fine-grained details. Meanwhile, its Adaptive Search Guidance configures a scale-efficient architecture for temporal-dimension reconstruction. These decoupled yet synergistic components jointly address media identity disentanglement and temporal-scale mismatch. Comprehensive evaluations involving 16 SOTA TSF models across 13 "multivariate" and 12 "multimodal" datasets, alongside targeted long-context comparisons against 14 time series foundation models and fused pretrained language models, demonstrate MIDAPN's consistent superiority and broad shared backbone compatibility. The code is available at https://github.com/MIDAPN.
Figures & tables
Fig. 1: Motivations from identity awareness and temporal capturing (spatiotemporal learning). MIDAPN achieves cross-media compatibility across multivariate and multimodal forecasting, while controlled replacements of MIDAG or SPConv with generic counterparts consistently increase errors
Method
Year
External Paired Modality
Pretrained PLM/TSFM
Explicit Cross- Variate Relation
Intermediate Cross- Modal Interaction
Identity/Group-Aware Relation Modeling
Frequency-Domain Modeling
Explicit Multi-Resolution Temporal Modeling
Search-Guided Hierarchical Reconstruction
Native multivariate models and their TaTS-based extensions
MSGNet [ 17 ]
2024
✗
✗
✓
✗
✗
✓
✓
✗
+ TaTS [ 15 ]
2026
✓
✓
✓
✓
✗
✓
✓
✗
TimeFilter [ 29 ]
2025
✗
✗
✓
✗
✗
✗
✗
✗
+ TaTS [ 15 ]
2026
✓
✓
✓
✓
✗
✗
✗
✗
DUET [ 16 ]
2025
✗
✗
✓
✗
✓
✓
✗
✗
TABLE I: Comparison with representative multivariate and multimodal models.“+TaTS” denotes its TaTS-based multimodal extension.
Fig. 2: MIDAPN architecture. (a) Workflow for "multivariate" scenarios. (b) Workflow for TaTS-style "multimodal" scenarios. In both settings, TSEM-his first extracts spatiotemporal representations through Decoupled Frequency Fusion, Multimedia Identity-Aware Graph, and Spectral Prism Convolution from historical windows, then a channel-independent MLP maps the history to the future, followed by TSEM-pred for forecast refinement (details in Sec. IV-E )
Fig. 3: Core TSEM modules: (a) Decoupled Frequency Fusion (DFF); (b) Multimedia Identity-Aware Graph (MIDAG); and (c) Spectral Prism Convolution (SPConv), whose architecture is configured by Adaptive Search Guidance (ASGM). Here, C and T denote the numbers of aligned variates and temporal steps, respectively. MIDAPN learns multivariate dependencies and multimodal heterogeneity within a unified backbone for Spatiotemporal decoupling
TABLE III: Statistics of Datasets and the Actual Use of RevIN/PRReg for MIDAPN. All time series are standardized
Models
MIDAPN
TimeFilter
VPNet
CPiRi
SEMixer
DUET
TimeKAN
FilterTS
TimePro
P-sLSTM
ModernTCN
TimeMixer
iTransformer
Time-LLM
MSGNet
PatchTST
TimesNet
(Years)
(Ours)
(ICML’25)
(ICLR’26)
(ICLR’26)
(WWW’26)
(KDD’25)
(ICLR’25)
(AAAI’25)
(ICML’25)
(AAAI’25)
(ICLR’24)
(ICLR’24)
(ICLR’24)
(ICLR’24)
(AAAI’24)
(ICLR’23)
(ICLR’23)
Metric
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
ETT (Avg)
0.318
0.353
0.318
0.359
0.334
0.365
0.332
0.367
0.318
0.363
0.326
0.361
0.330
0.364
0.324
0.359
0.329
0.365
0.336
0.371
0.329
0.361
0.329
0.365
0.336
0.369
0.332
0.366
0.342
0.373
0.334
0.366
0.352
0.380
Flight
0.161
0.262
0.161
0.268
0.161
0.269
0.160
0.268
0.160
0.267
0.163
0.271
0.179
0.289
0.165
0.269
0.164
0.272
0.171
0.281
0.172
0.282
0.172
0.274
0.180
0.292
0.179
0.291
0.208
0.321
0.163
0.275
0.191
0.304
Weather
0.235
0.259
0.239
0.269
0.246
0.275
0.253
0.276
0.256
0.281
0.251
0.273
0.243
0.272
0.244
0.274
0.251
0.276
0.261
0.283
0.245
0.273
0.245
0.275
0.258
0.278
0.263
0.284
0.249
0.278
0.259
0.281
0.259
0.287
Traffic
0.461
0.276
0.409
0.269
0.448
0.291
0.450
0.306
0.505
0.324
0.451
0.269
0.601
0.363
0.471
0.315
0.440
0.286
0.453
0.286
0.490
0.326
0.484
0.297
0.428
0.282
0.541
0.358
0.660
0.381
0.481
0.304
0.620
0.336
TABLE IV: Averaged Comparison across 25 “multivariate” and “multimodal” datasets in multiple future horizons. Red indicates best, while blue indicates second. Prefixes: [MV] Multivariate, [MM] Multimodal. For fairness: (1) In multivariate, 96 predict {12, 24, 48, 96} length for PEMS, 36 predict {24, 36, 48, 60} length for ILI and NASDAQ, other datasets use 96 predict {96, 192, 336, 720} in long-terms. (2) In multimodal, we use 8 predict {8, 10, 12} length for unified-granularity evaluation within TaTS [ 15 ]
Fig. 6: TSFMs, fused-PLMs and multimodal frameworks comparison (sorted by averaged MAE from low to high). MIDAPN consistently achieves the lowest MAE across the selected benchmarks, highlighting its accuracy advantage over pretrained forecasting approaches under extended look-back settings
Fig. 7: Forecasting across media-bridged domains with periodic, volatile, and trend-dominated patterns. MIDAPN generally tracks major trends and recurring oscillations, with localized deviations around abrupt changes
Task Types
Multivariate
Multimodal
Multivariate + Text
Datasets
Electricity
Weather
Traffic
Energy
LEU
Strategies
Avg
Drop
Avg
Drop
Avg
Drop
Avg
Drop
Avg
Drop
MIDAPN
0.2135
—
0.2468
—
0.1963
—
0.2478
—
0.5665
—
w/o DFF
0.2389
11.89%
0.2541
2.99%
0.2122
8.06%
0.2560
3.30%
0.6233
10.03%
w/o MIDAG
0.2384
11.65%
0.2526
2.38%
0.2113
7.64%
0.2555
3.09%
0.7015
23.83%
w/o SPConv
0.2428
13.70%
0.2546
3.19%
0.2117
7.81%
0.2508
1.21%
0.5792
2.24%
TABLE V: Ablation studies (For averaged horizons, multivariate: {96, 192, 336, 720}, multimodal: {8, 10, 12}). The "Drop" is the error percentage increase in Average of MAE and MSE
Fig. 11: Text encoder evaluation. (a): Easily predictable domains; (b): Challenging domains. (Embedding dimension: [BERT:768, BERT-XL:1024, GPT2:768, GPT2-XL:1600, LLaMA3.2-1B:2048, LLaMA2-7B/LLaMA3-8B:4096, Qwen3-0.6B:1024, Qwen2/Deepseek-distil-1.5B:1536, Qwen3-8B:4096], Input-8-predict-12). The contribution of PLMs remains inconclusive
Fig. 12: Performance, efficiency and memory footprint comparison. (a): Input-96-predict-720 for Electricity; (b): Input-8-predict-12 for Health. By selectively ablating TSEMhis and TSEMpred , MIDAPN can achieve the most expeditious training speed compared to other SOTA TSF models
Datasets
Variates
RevIN
PRReg
Timestamps
Frequency
Properties
Date
ETTm2
7
×
✓
69680
15mins
Volatility & Periodicity
2016.7.1-2018.6.26
ETTh2
7
×
✓
17420
1h
Volatility & Periodicity
2016.7.1-2018.6.26
Flight
7
✓
×
26304
1h
Volatility
2019.1.1-2021.12.31
Weather
21
✓
×
52696
10mins
Fluctuation
2020.1.1-2021.1.1
Traffic
862
✓
×
17544
1h
Stable Periodicity
2016.7.1-2018.7.2
Electricity
321
✓
×
26304
1h
Stable Periodicity
2016.7.1-2019.7.2
TABLE VII: Statistics of Datasets and the Actual Use of RevIN/PRReg for MIDAPN
MIDAPN -nonloss
MIDAPN -loss-0.001
MIDAPN -loss-0.01
Dataset
MSE
MAE
MSE
MAE
MSE
MAE
Agriculture
0.1314
0.2317
0.1315
0.2320
0.1323
0.2334
Climate
0.9061
0.7618
0.9059
0.7617
0.9116
0.7638
Energy
0.1834
0.3359
0.1968
0.3436
0.1877
0.3407
Environment
0.2691
0.3677
0.2694
0.3719
0.2691
0.3682
Health
0.8322
0.6260
0.8365
0.6287
0.8716
0.6423
TABLE VIII: Performance comparison with the auxiliary assignment. Red / Blue indicate 1st / 2nd
Dataset
Model
MSE
MAE
Multivariate datasets
ETTh2
MIDAPN
0.368±0.003
0.392±0.002
ETTm2
MIDAPN
0.271±0.001
0.316±0.001
Flight
MIDAPN
0.161±0.000
0.262±0.000
Weather
MIDAPN
0.236±0.001
0.260±0.001
PEMS03
MIDAPN
0.096±0.001
0.201±0.001
TABLE IX: Forecasting results over five random seeds
Dataset
Metric
Text-Embedding Dimension
4
8
12
24
48
96
Climate
MSE
0.910
0.910
0.890
0.905
0.906
0.907
MAE
0.766
0.765
0.757
0.763
0.763
0.764
Energy
MSE
0.182
0.176
0.172
0.178
0.178
0.177
MAE
0.330
0.327
0.324
0.327
0.328
0.327
Environment
MSE
0.269
0.268
0.267
0.268
0.270
0.270
TABLE X: Performance under different textual dimensions
Models
MIDAPN
TimeFilter
VPNet
CPiRi
SEMixer
DUET
TimeKAN
FilterTS
TimePro
P_sLSTM
ModernTCN
TimeMixer
iTransformer
TimeLLM
MGSNET
PatchTST
TimesNet
(Years)
(Ours)
(2025)
(2026)
(2026)
(2026)
(2025)
(2025)
(2025)
(2025)
(2025)
(2024)
(2024)
(2024)
(2024)
(2024)
(2023)
(2023)
Metric
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
ETTm2
96
0.168
0.250
0.169
0.255
0.173
0.254
0.178
0.259
0.176
0.262
0.174
0.255
0.174
0.255
0.172
0.255
0.178
0.260
0.181
0.267
0.175
0.256
0.174
0.258
0.180
0.264
0.178
0.262
0.177
0.262
0.175
0.259
0.187
0.267
192
0.230
0.294
0.235
0.299
0.247
0.305
0.245
0.305
0.243
0.310
0.243
0.302
0.239
0.299
0.237
0.299
0.242
0.303
0.250
0.312
0.239
0.298
0.239
0.302
0.250
0.309
0.243
0.304
0.247
0.307
0.241
0.302
0.249
0.309
336
0.295
0.330
0.293
0.336
0.322
0.349
0.306
0.344
0.306
0.351
0.304
0.341
0.301
0.340
0.299
0.398
0.303
0.342
0.315
0.353
0.298
0.336
0.296
0.340
0.311
0.348
0.308
0.345
0.312
0.346
0.305
0.343
0.321
0.351
720
0.391
0.389
0.390
0.393
0.430
0.412
0.403
0.401
0.407
0.416
0.399
0.397
0.395
0.396
0.397
0.394
0.400
0.399
0.418
0.413
0.401
0.395
0.393
0.397
0.412
0.407
0.408
0.407
0.414
0.403
0.402
0.400
0.408
0.403
TABLE XI: Comprehensive performance comparison in long-term forecasting. Results are highlighted strictly following the Dense Ranking principle. Red / Blue indicate 1st / 2nd
Models
MIDAPN
TimeFilter
VPNet
CPiRi
SEMixer
DUET
TimeKAN
FilterTS
TimePro
P_sLSTM
ModernTCN
TimeMixer
iTransformer
TimeLLM
MGSNET
PatchTST
TimesNet
(Years)
(Ours)
(2025)
(2026)
(2026)
(2026)
(2025)
(2025)
(2025)
(2025)
(2025)
(2024)
(2024)
(2024)
(2024)
(2024)
(2023)
(2023)
Metric
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
Pems03
12
0.060
0.161
0.067
0.169
0.066
0.170
0.077
0.190
0.081
0.191
0.064
0.166
0.093
0.206
0.094
0.202
0.076
0.180
0.077
0.183
0.069
0.174
0.076
0.188
0.071
0.174
0.095
0.208
0.078
0.187
0.099
0.216
0.085
0.192
24
0.076
0.181
0.085
0.192
0.090
0.201
0.097
0.213
0.122
0.236
0.081
0.186
0.155
0.269
0.112
0.222
0.101
0.210
0.109
0.219
0.093
0.203
0.113
0.226
0.093
0.201
0.135
0.248
0.108
0.218
0.142
0.259
0.118
0.223
48
0.103
0.211
0.126
0.236
0.138
0.252
0.140
0.253
0.195
0.302
0.114
0.222
0.236
0.337
0.160
0.265
0.137
0.248
0.163
0.270
0.141
0.250
0.191
0.292
0.125
0.236
0.198
0.302
0.178
0.272
0.211
0.319
0.155
0.260
96
0.143
0.248
0.173
0.283
0.201
0.311
0.174
0.289
0.255
0.354
0.175
0.283
0.306
0.399
0.208
0.307
0.184
0.292
0.209
0.309
0.208
0.306
0.288
0.363
0.164
0.275
0.243
0.342
0.238
0.328
0.269
0.370
0.228
0.317
TABLE XII: Comprehensive performance comparison in short-term forecasting. Results are highlighted strictly following the Dense Ranking principle. Red / Blue indicate 1st / 2nd
Models
MIDAPN
TimeFilter
VPNet
CPiRi
SEMixer
DUET
TimeKAN
FilterTS
TimePro
P_sLSTM
ModernTCN
TimeMixer
iTransformer
TimeLLM
MGSNET
PatchTST
TimesNet
(Years)
(Ours)
(2025)
(2026)
(2026)
(2026)
(2025)
(2025)
(2025)
(2025)
(2025)
(2024)
(2024)
(2024)
(2024)
(2024)
(2023)
(2023)
Metric
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
Agriculture
8
0.080
0.182
0.083
0.191
0.082
0.188
0.079
0.182
0.083
0.189
0.099
0.220
0.107
0.220
0.082
0.185
0.082
0.195
0.084
0.187
0.080
0.180
0.107
0.221
0.078
0.187
0.081
0.187
0.084
0.194
0.080
0.188
0.086
0.202
10
0.102
0.207
0.105
0.213
0.103
0.211
0.102
0.210
0.103
0.211
0.105
0.221
0.124
0.238
0.105
0.210
0.105
0.216
0.109
0.214
0.103
0.208
0.121
0.234
0.101
0.210
0.106
0.214
0.112
0.229
0.104
0.214
0.104
0.216
12
0.131
0.232
0.134
0.237
0.131
0.236
0.135
0.244
0.134
0.236
0.134
0.239
0.154
0.261
0.133
0.238
0.133
0.241
0.137
0.238
0.136
0.234
0.149
0.262
0.132
0.235
0.137
0.240
0.141
0.248
0.133
0.240
0.136
0.248
Average
0.104
0.207
0.107
0.214
0.105
0.212
0.105
0.212
0.107
0.212
0.113
0.227
0.128
0.240
0.107
0.211
0.107
0.217
0.110
0.213
0.106
0.207
0.126
0.239
0.104
0.211
0.108
0.214
0.112
0.224
0.106
0.214
0.109
0.222
TABLE XIII: Comprehensive performance comparison across “multimodal scenarios for different prediction horizons (8, 10, 12). Results are highlighted strictly following the Dense Ranking principle. Red / Blue indicate 1st / 2nd
Models
MIDAPN
GTM
SE-LLM
CORA
Aurora
DualSG
Time-VLM
SEMPO
Sundial
CALF
MOIRAI
(Years)
(Ours)
(2026)
(2026)
(2026)
(2026)
(2025)
(2025)
(2025)
(2025)
(2025)
(2024)
Metric
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
Weather
96
0.140
0.179
0.147
0.197
0.150
0.200
0.149
0.194
0.160
0.207
0.156
0.201
0.148
0.200
0.154
0.205
0.157
0.205
0.160
0.208
0.199
0.211
192
0.190
0.228
0.192
0.241
0.194
0.243
0.193
0.237
0.202
0.247
0.196
0.239
0.193
0.240
0.198
0.246
0.205
0.251
0.205
0.252
0.246
0.251
336
0.245
0.275
0.250
0.291
0.247
0.285
0.240
0.274
0.252
0.288
0.248
0.280
0.243
0.281
0.249
0.285
0.253
0.289
0.253
0.288
0.274
0.291
720
0.316
0.328
0.310
0.334
0.323
0.339
0.313
0.322
0.307
0.327
0.323
0.334
0.312
0.332
0.321
0.337
0.320
0.336
0.329
0.341
0.337
0.340
TABLE XIV: Comprehensive performance comparison on Weather and Electricity datasets for different prediction horizons (96, 192, 336, 720). Results are highlighted strictly following the Dense Ranking principle. Red / Blue indicate 1st / 2nd
Models
MIDAPN
Aurora
SE-LLM
VoT
Time-VLM
CALF
ChatTime
GPT4MTS
MM-TSFLib
(Years)
(Ours)
(2026)
(2026)
(2026)
(2025)
(2025)
(2025)
(2024)
(2024)
Metric
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
Traffic
8
0.157
0.224
0.158
0.286
0.172
0.229
0.167
0.231
0.212
0.313
0.176
0.232
0.257
0.347
0.195
0.256
0.195
0.250
12
0.176
0.236
0.168
0.294
0.191
0.232
0.181
0.239
0.222
0.322
0.193
0.243
0.261
0.349
0.218
0.268
0.261
0.272
Average
0.167
0.230
0.163
0.290
0.182
0.231
0.174
0.235
0.217
0.318
0.185
0.238
0.259
0.348
0.207
0.262
0.228
0.261
Health
24
0.958
0.648
1.332
0.796
1.232
0.774
1.210
0.707
1.491
0.839
1.451
0.749
1.758
0.888
1.513
0.802
1.562
0.768
TABLE XV: Performance comparison on Traffic, Health, and Environment datasets. Results are highlighted strictly following the Dense Ranking principle. Red / Blue indicate 1st / 2nd
Most existing time series forecasting methods rely solely on numerical observations, overlooking rich contextual information from auxiliary texts. Recent multimodal approaches attempt to incorporate textual signals, but they often treat text as static features or use large language models as forecasting backbones, limiting their ability to capture temporal dynamics and increasing computational cost. To address these challenges, we propose TAC-Time, a unified framework that transforms textual information into additional temporal channels. By modeling text features jointly with numerical sequences in a shared temporal backbone, TAC-Time preserves temporal continuity and periodic structures while remaining efficient and scalable. This formulation also enables systematic interpretability analyses. We show strong cross-modal dependencies through attention and frequency-domain analyses, and identify predictive textual signals whose correlation-aware alignment yields partial forecasting improvements. Extensive experiments on real-world multimodal benchmarks demonstrate that TAC-Time outperforms prior methods.
Real-world time series come with text: metadata, descriptions, news, reports. Yet time series foundation models process numerical sequences in isolation, and the multimodal text-and-time-series models that attempt to bridge the two all adapt a pretrained language model post hoc, inheriting representations shaped without ever seeing temporal data. These models are also evaluated almost exclusively against other multimodal baselines, not against the strongest unimodal foundation models in either domain, leaving open whether joint training is needed at all. We present Chronicle, a compact 324M-parameter decoder-only transformer trained from scratch on natural language and time series within a single unified architecture. Both modalities share the same transformer blocks, attention mechanism, and residual stream; the bulk of pretraining uses unimodal batches so cross-modal capability emerges purely from shared parameters, with a short alignment stage that interleaves the two. To our knowledge, Chronicle is the first model jointly pretrained on text and time series from scratch, and the first multimodal model evaluated against dedicated foundation models in both domains. It matches Gemma-3-270M-PT on 19 NLU tasks, sets a new bar for frozen-embedding time series classification on 24 UCR/UEA datasets, and produces multimodal forecasts on Time-MMD that beat every supervised fusion baseline, all from a single backbone.
Paul Quinlan, Jeremy Levasseur, Qingguo Li +1
1InertialAI · Department of Electrical and Computer Engineering, Queen’s University · Department of Mechanical and Materials Engineering, Queen’s University
Multivariate time-series forecasting is essential to many real-world applications. Recent large vision models (LVMs) offer a promising paradigm by transferring cross-domain visual priors to time-series forecasting. However, existing LVM-based methods face two key challenges: balancing independent visual representation spaces with cross-variable dependency modeling, and adapting vision backbones pretrained on natural images to the distinct temporal semantics of time-series images. To address these challenges, we propose MUSE, a dependency-aware adaptation framework built on a fully frozen pretrained MAE. First, the Variable Context Refinement Module (VCR) aggregates shared temporal information within each variable and models cross-variable contextual dependencies while preserving independent visual spaces. Second, the Temporal-Periodic Refinement Module (TPR) performs lightweight refinement at different encoder depths and explicitly models across-period temporal dependencies and within-period periodic dependencies. The two modules independently produce forecasts, which are fused through a learnable prediction-level gate. Experiments on 10 real-world datasets demonstrate that MUSE achieves state-of-the-art performance.