Video Representation Learning

Latest papers 87

All topics
CardsList
  1. Self-Supervised Learning of Structured Dynamics from Videos

    Jul 23, 2026Lukas Knobel, Andrew Zisserman, Yuki M. AsanoVideo Representation LearningSelf-Supervised Visual Representation Learning

  2. Sidewalk Moments: Are Richer Representations Always More Human-Aligned? Evidence from City-Walk Videos

    Jul 23, 2026Liu Liu, Freya Huying Tan, Fábio DuarteEgocentric Video UnderstandingVideo Representation Learning

  3. Latent-Action-Guided Video-Language Feature Learning for Surgical Instrument-Tissue Interaction Recognition

    Jul 22, 2026Jiajun Cheng, Sainan Liu, Subarna Tripathi +2Latent Action LearningSurgical Video Understanding

  4. VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders

    Jul 15, 2026Zhihao Xie, Junfeng Wu, Xinting Hu +2Video Diffusion ModelsVideo Representation Learning

  5. HyperGS: Fast and Generalizable Gaussian Video Representation

    Jul 13, 2026Fatimah Zohra, Chen Zhao, Shuming Liu +23D Gaussian SplattingVideo Representation Learning

  6. `Attention-Guided Cross-Temporal Clustering for Self-Supervised Video Object Segmentation

    Jul 8, 2026Waqas Arshid, Mohammad Awrangjeb, Alan Wee-Chung Liew +1Video Object SegmentationSelf-Supervised Learning

  7. Gen4U: Unifying Video Generation and Understanding via Diffusion

    Jul 7, 2026Michael King, Aravindh Mahendran, Matthew Koichi Grimes +5Video Diffusion ModelsVideo Representation Learning

  8. From Pixels to Temporal Correlations: Learning Informative Representations for Reinforcement Learning Pre-training

    Jul 1, 2026Jinwen Wang, Youfang Lin, Xiaobo Hu +4Contrastive LearningReinforcement Learning

  9. Local Motion Matters: A Deconstruct-Recompose Paradigm for Reinforcement Learning Pre-training from Videos

    Jul 1, 2026Jinwen Wang, Youfang Lin, Xiaobo Hu +2Cross-Embodiment Robot LearningRobotic RL

  10. P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture

    Jun 22, 2026Felix Tristram, Stefano Gasperini, Benjamin Killeen +4Temporal Action SegmentationVideo Understanding

  11. LEViL: Label-Efficient Video Learning via Zero-Shot Distillation over VLM-Generated Pseudo-Label Spaces

    Jun 19, 2026Aslı ÇelikVLM DistillationVideo Action Recognition

  12. UNIEGO: Proxies as Mediators for Unified Egocentric Video Representation Learning

    Jun 18, 2026Wenhao Chi, Arkaprava Sinha, Dominick Reilly +2Cross-Modal Knowledge DistillationMulti-View Learning

  13. DREAM: Extending Vision-Language Models with Dual-Objective Encoding for Cross-Modal Retrieval

    Jun 17, 2026Kaleem Ullah, Altaf Hussain, Muhammad Munsif +1Cross-Modal RetrievalVideo Representation Learning

  14. Selective Synergistic Learning for Video Object-Centric Learning

    Jun 14, 2026WonJun Moon, Jae-Pil HeoObject-Centric Representation LearningVideo Representation Learning

  15. MVEB: Massive Video Embedding Benchmark

    Jun 12, 2026Adnan El Assadi, Roman Solomatin, Isaac Chung +13Multimodal EmbeddingVideo Representation Learning

  16. MLT-Dedup: Efficient Large-Scale Online Video Deduplication via Multi-Level Representations and Spatial-Temporal Matching

    Jun 10, 2026David Yuchen Wang, Haoying Li, Hailun Xu +6Video Representation Learning

  17. Momentum-Guided Semantic Forecasting (MoFore) for Self-Supervised Video Representation Learning

    Jun 8, 2026Qinwu XuVideo UnderstandingVideo Representation Learning

  18. Do Video Foundation Models Understand Intuitive Physics? A Layerwise Probing Analysis

    Jun 8, 2026Samuele Punzo, Niccolò Caselli, Ippokratis Pantelidis +3Vision Foundation ModelsVideo Representation Learning

  19. Self-supervised Learning Matters: A Simple Ensemble Solution for Micro-Gesture Recognition

    Jun 8, 2026Tingyi Liu, Kun Li, Fei Wang +5Video Representation LearningGesture Recognition

  20. What Makes Video World Model Latents Action-Relevant: Prediction over Reconstruction

    Jun 5, 2026Jewon Yeom, Hanseul Kim, Jeongjae Park +3Video Representation LearningPredictive Representation Learning

  21. The Invisible Hand of Physics: When Video Diffusion Models Know More Than They Show

    Jun 3, 2026Parsa Esmati, Somjit Nath, Katja Hofmann +3Video Diffusion ModelsVideo Representation Learning

  22. TrAction: Action Recognition with Sparse Trajectories

    Jun 2, 2026Jan F. Meier, Felix B. Mueller, Alexander Ecker +1Video Action RecognitionVideo Representation Learning

  23. Paving the Way for Point Cloud Video Representation Learning Using A PDE Model

    Jun 1, 2026Zhuoxu Huang, Zhenkun Fan, Jungong Han +1Point Cloud LearningVideo Representation Learning

  24. CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection

    May 31, 2026Xin Dong, Wenjia Geng, Wenfeng Deng +1Temporal Video GroundingVideo Representation Learning

  25. VidPrism: Heterogeneous Mixture of Experts for Image-to-Video Transfer

    May 27, 2026Rui Lin, Chuanming Wang, Huadong MaVLM AdaptationVideo Representation Learning

  26. Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Models

    May 27, 2026Haozhan Shen, Tiancheng Zhao, Kangjia Zhao +1Vision-Language ModelsVisual Spatial Reasoning