cs.CVOct 8, 2026

TAM: Task-Aware Memory Distillation for Efficient Spatiotemporal Prediction

Authors: Yuqi Li, Xiaoqin Feng, Fan Xu, Weilun Feng, Chuanguang Yang, Yingli Tian, Hao Wu

Organizations: CUNY City College of NY · Wyze Labs, Inc. · University of Science and Technology of China · Institute of Computing Technology, Chinese Academy of Sciences · Tsinghua University

Abstract

Knowledge distillation enables efficient spatiotemporal prediction by transferring knowledge from an accurate teacher to a compact student. However, matching outputs or features independently for each sample leaves cross-sample predictive structure underused. Exploiting this structure requires representations and historical references that reflect the dynamics of each task. We propose TAM, a Task-Aware Memory Distillation framework that organizes a frozen teacher's knowledge into a bounded, retrievable history. Memory entries encode latent features, forecast changes, or flow residuals, while task-specific selection rules identify relevant historical references. The student either matches the teacher's similarity distribution over shared references or regresses observation-conditioned residual prototypes. These objectives complement supervised prediction and conventional distillation. The teacher, memory, and auxiliary adapters are used only during training, leaving student inference unchanged. We evaluate TAM on video prediction, weather forecasting, and traffic flow prediction across multiple teacher-student configurations. Averaged over four paired runs, adding TAM improves SSIM on all six video datasets and reduces MSE on five relative to the corresponding KD baselines. Mean paired MSE reductions reach 1.86% on KittiCaltech, 1.93% on WeatherBench with a gSTA teacher, and 1.01% on TaxiBJ. These results demonstrate the utility of historical teacher supervision across distinct forecasting tasks without additional student inference cost.

Figures & tables

Explore similar work

CardsList
  1. MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction

    Apr 11, 2026Wenchang Duan, Zhenguo Gao, Jinguo Xian +1Model CompressionMulti-Agent Trajectory Prediction

  2. Robust Trajectory Distillation: Hybrid Reweighting Meets Teacher-Inspired Targets

    Jun 29, 2026Kaifeng Chen, Lechao Cheng, Jiyang Li +6Dataset DistillationLearning with Noisy Labels