Just for FUNS: LLM-Guided Spatio-Temporal Graph Node Generation for Forecasting Unobserved Node States
Authors: Shuhao Li, Weidong Yang, Changan Liu, Wei Zhuo, Yingbo Zhou, Fan Zhang, Siqiang Luo
Organizations: Fudan University Shanghai, China · Fudan University Shanghai, China Zhuhai Fudan Innovation Research Institute Zhuhai, China · College of Computing and Data Science Nanyang Technological University Singapore · Shanghai Key Laboratory of Data Science Fudan University Shanghai, China · GZHU-SCHB Intelligent Transportation Joint Lab Guangzhou University Guangzhou, China
Spatio-temporal forecasting is a cornerstone of logistics, urban planning, and intelligent transportation systems. However, constrained by deployment costs and maintenance resources, sensor networks often lack comprehensive spatial coverage, rendering Forecast Unobserved Node States (FUNS) a critical yet formidable challenge. Conventional models rely on historical observations and typically falter when encountering nodes without prior records. To address this, we redefine the problem as a conditional generation task on spatio-temporal graphs and propose GenST, a framework that introduces Large Language Models (LLMs) as a semantic bridge, leveraging a pre-trained LLM fine-tuned to extract rich semantic features from node descriptions, such as functional zones and road network structures, to compensate for missing spatio-temporal signals. Specifically, we design a two-stage generative architecture: a Spatio-Temporal VAE first compresses spatio-temporal dynamics into a latent space, followed by a Generative Transformer (GenT) that reconstructs the future states of unobserved nodes from noise, guided by multi-modal conditions including semantics, geographic coordinates, and neighborhood contexts. Experiments on six traffic and two non-traffic datasets show GenST significantly outperforms existing baselines in zero-shot prediction tasks, demonstrating the practical potential of semantic-guided generation for mitigating spatio-temporal data sparsity.
Figures & tables
Figure 1 . Illustrative scenarios and spatial heterogeneity of unobserved nodes for FUNS.
Figure 2 . Reformulating FUNS as unobserved node generation based on observed graphs.
Figure 3 . The GenST architecture. The architecture features a unified generative pipeline centered on Stage 2, where latent foundation is unified with multi-modal context alignment to drive the primary engine for direct Stage 3 zero-shot inference.
Models
Datasets
METR-LA
PeMS-Bay
PeMS03
PeMS04
PeMS07
PeMS08
Tasks
MAE
RMSE
MAPE
MAE
RMSE
MAPE
MAE
RMSE
MAPE
MAE
RMSE
MAPE
MAE
RMSE
MAPE
MAE
RMSE
MAPE
K-NN
Global
16.84
26.76
30.38%
13.79
19.05
27.50%
81.28
124.99
91.34%
107.81
141.79
62.36%
107.62
176.30
79.96%
101.17
135.04
148.28%
Unobserved
20.90
33.01
35.72%
14.83
20.62
34.78%
147.77
230.74
275.44%
180.46
247.05
136.68%
243.90
323.19
268.02%
194.22
259.33
316.14%
Kriging
Global
15.50
21.09
28.21%
13.21
22.43
28.42%
82.89
124.68
92.26%
99.81
128.27
82.80%
132.57
159.28
141.37%
107.12
141.78
170.59%
Unobserved
18.82
31.07
32.41%
14.77
24.03
29.48%
116.74
141.40
210.65%
122.15
157.50
144.07%
165.09
206.04
170.74%
145.50
181.34
297.98%
HA
Global
20.97
24.87
32.56%
12.44
22.43
27.26%
92.78
116.64
214.90%
111.10
138.07
165.64%
133.26
162.86
150.08%
98.42
122.49
123.92%
Table 1 . Main performance comparison. Best Global/Unobserved results are highlighted in bold with dark blue and dark orange colors, respectively; second-best results are indicated with light blue and light orange backgrounds.
The efficient operation of modern cellular networks hinges on the accurate analysis of spatio-temporal traffic data. Mastering these patterns is essential for core network functions, chiefly forecasting future load to pre-empt congestion and imputing missing values caused by sensor failures or transmission errors to ensure data continuity. While deeply connected, forecasting and imputation have historically evolved as separate sub-fields. The dominant paradigm, Spatio-Temporal Graph Neural Networks (STGNNs), while effective, are often specialized, computationally intensive, and exhibit limited generalization. Concurrently, adapting large pre-trained language models (LLMs) offers a powerful alternative for sequence modeling, yet existing approaches provide weak structural guidance, leading to unstable convergence and a narrow focus on forecasting. To bridge these gaps, we propose U-STS-LLM, a unified framework built on a spatio-temporally steered LLM. Our core innovation is a Dynamic Spatio-Temporal Attention Bias Generator that synthesizes a persistent functional graph with transient nodal states to explicitly steer the LLM's attention. Coupled with a partially frozen backbone tuned via Low-Rank Adaptation (LoRA) and a Gated Adaptive Fusion mechanism, the model achieves stable, parameter-efficient adaptation. Trained under a unified multi-task objective, U-STS-LLM learns a holistic data representation. Extensive experiments on real-world cellular datasets demonstrate that U-STS-LLM establishes new state-of-the-art performance in both long-horizon forecasting and high-missing-rate imputation, while maintaining remarkable training efficiency and stability, offering a novel blueprint for harnessing foundation models in structured, non-linguistic domains.
Yichen Zhang, Jun Li
School of Information Science and Engineering, Southeast University, Nanjing, 210096, CHINA.
Large language models (LLMs) have shown promise for spatio-temporal forecasting, but existing approaches often rely on regularly sampled token sequences and struggle with irregular observations because of temporal asynchrony, representation-space misalignment, and limited context windows. We propose LLMODE, a token-efficient framework for irregular spatio-temporal forecasting with a frozen LLM backbone. LLMODE first uses a graph-aware ODE encoder to reconstruct irregular graph observations as a continuous-time latent trajectory. A Fixed-Budget Perceiver Resampler then compresses this variable-length trajectory into a fixed number of dynamic memory tokens. In parallel, compact statistical descriptors are encoded and resampled into context memory tokens. A dual-source gated cross-attention module injects both memories into the frozen LLM, enabling controlled utilization of external spatio-temporal evidence. Experiments on three real-world urban datasets and two physical-dynamics benchmarks show competitive overall performance, with clearer advantages under sparse or dynamically complex irregular sampling. Additional evaluations on unseen urban regions further demonstrate strong zero-shot generalization without adaptation.
Spatiotemporal prediction aims to learn discriminative representations from correlated temporal signals over spatial structures for accurate future inference. A central challenge is \emph{spatial indistinguishability}: different nodes may share similar historical patterns yet evolve toward divergent futures, severely degrading forecasting performance in real-world sensor networks. Existing embedding-based and graph neural network (GNN)-based approaches can partially detect such ambiguous nodes but rely on historical similarity, struggling to capture \emph{future behavioral divergence}. We propose \textbf{STOT} (\textbf{S}patio\textbf{T}emporal \textbf{O}ptimal \textbf{T}ransport), a self-supervised framework that resolves spatiotemporal ambiguity via structured masking guided by optimal transport. Our key idea treats indistinguishability as a \emph{disambiguation} problem: future states are inferred by exploiting concurrent spatial correlations and their time-varying similarity. We design a similarity-aware metric for dynamic inter-node relationships and an optimal transport-based masking strategy to emphasize ambiguous positions during pre-training. A batch consistency constraint preserves semantic coherence, while a random-walk masking mechanism promotes structured context exploration. Experiments on six real-world datasets show that STOT performs competitively with state-of-the-art baselines on the evaluated benchmarks and improved interpretability through transport-plan visualizations.
Guangyu Wang, Jiawei Tong
Data Science and Artificial Intelligence, Dongbei University of Finance and Economics, Dalian, Liaoning, China · Institute of Systems, Molecular & Integrative Biology, University of Liverpool, Liverpool, United Kingdom