cs.CVSep 27, 2026

Unifying Video Tasks via Spatiotemporal Analogy

Authors: Chia-Hsiang Kao, Belinda Zeng, Bharath Hariharan, Menglin Jia

Organizations: Cornell University · Meta

Abstract

Adapting video models to new tasks typically requires dedicated data curation and fine-tuning. While visual analogy provides a training-free alternative by specifying tasks in-context, it remains restricted to the image domain. To explore whether analogy-based methods can unify diverse video tasks and generalize to out-of-distribution scenarios, we introduce ViGeo, a framework that extends visual in-context learning to the video domain via spatiotemporal canvas completion. Evaluated on a diverse task taxonomy with a strict train-test split, ViGeo generalizes to unseen video manipulations and zero-shot modalities (e.g., event cameras). Finally, we identify task internalization, where a query format associated with a pretrained task overrides the demonstration, and show that this shortcut can be removed with a small amount of task-unrelated data, highlighting the need to decorrelate prompt format from task identity.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. How Far Are Video Models from True Multimodal Reasoning?

    Apr 21, 2026Xiaotian Zhang, Jianhui Wei, Yuan Wang +9Multimodal ReasoningVideo Understanding

  2. Find, Fix, Reason: Context Repair for Video Reasoning

    Apr 17, 2026Haojian Huang, Chuanyu Qin, Yinchuan Li +1Reasoning SkillsLarge Multimodal Models