cs.CVOct 8, 2026

Revisiting Identity and Spectra Dispersion in Media-Bridged Time Series Forecasting: Linking Multivariate Signals and Narrative Flows

Authors: Jierui Lei, Wenjian Zhang, Qingyi Yang, Yuyang Hong, Fangzheng Chen, Zhengbo Zhang, Haina Tang, Shiming Xiang

Organizations: School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences, Beijing 100049, China · State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, Beijing 100190, China · School of Intelligent Systems Engineering, Sun Yat-sen University, Shenzhen, Guangdong 518107, China · School of Computer Science and Technology, Huazhong University of Science and Technology, Wuhan, Hubei 430074, China · School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing 100049, China · College of Computer Science, Chongqing University, Chongqing 401331, China

Abstract

Media-bridged time series forecasting is expanding to encompass traditional "multivariate" and emerging "multimodal" (e.g., through textual assistance). Existing Time Series Forecasting (TSF) models still rely on paradigm-specific relation, fusion, and temporal modules, hindering a common forecasting backbone across numerical and pre-aligned narrative-flow settings. To explore this, we propose the Multimedia Identity-Aware Prism Network (MIDAPN), a unified spatiotemporal forecasting backbone based on media-general graph adaptation and automatic temporal learning: (1) Following media pre-alignment, our Multimedia Identity-Aware Graph (MIDAG) revisits identity through static essence, dynamic behavior, and latent commonality, inducing affinities that extend variable-specific dependencies across media. Contextual Identity Modulation (CIM) further refines discriminative aggregation. (2) We develop Spectral Prism Convolution (SPConv) to automatically perform hierarchical temporal analysis, balancing coarse trends and fine-grained details. Meanwhile, its Adaptive Search Guidance configures a scale-efficient architecture for temporal-dimension reconstruction. These decoupled yet synergistic components jointly address media identity disentanglement and temporal-scale mismatch. Comprehensive evaluations involving 16 SOTA TSF models across 13 "multivariate" and 12 "multimodal" datasets, alongside targeted long-context comparisons against 14 time series foundation models and fused pretrained language models, demonstrate MIDAPN's consistent superiority and broad shared backbone compatibility. The code is available at https://github.com/MIDAPN.

Figures & tables

Explore similar work

CardsList
  1. TAC-Time: Texts as Channels For Multimodal Time Series Forecasting

    Sep 21, 2026Jiayi Liang, Xiaotian Gu, Xinyu Xie +2Multivariate Time Series ForecastingExplainable Time Series Forecasting

  2. Chronicle: A Multimodal Foundation Model for Joint Language and Time Series Understanding

    May 18, 2026Paul Quinlan, Jeremy Levasseur, Qingguo Li +1Cross-Modal LearningMultivariate Time Series Forecasting