Video Representation Learning

Latest papers 87

All topics
CardsList
  1. HuC-VideoMAE: Human-Centric Video Masked Autoencoding from synthetic data

    Oct 6, 2026Ricardo Pizarro, Roberto Valle, José M. Buenaposada +2Synthetic Data PretrainingMasked Autoencoders

  2. Video Encoders Built on Image Representations

    Oct 5, 2026Jusheng Zhang, Wenhao Wang, Longqi Cai +3Video-Language ModelsVideo Representation Learning

  3. Affine-Aligned Atlas for Canonical Gaussian Construction in Video Representation

    Oct 1, 2026Masaya Takabe, Hiroshi Watanabe, Sujun Hong +13D Gaussian SplattingDynamic Scene Reconstruction

  4. Image Classifiers are Efficient Self-Supervised Video Representation Learners

    Sep 30, 2026Owais Iqbal, Sudipta Sarkar, Shyam Marjit +3Self-Supervised Pre-TrainingVideo Representation Learning

  5. I Have a Stream: Making Self-Supervised Learning Work on Continuous Video

    Sep 30, 2026Ivan Martinović, Lukas Knobel, Yuki M. AsanoMasked AutoencodersSelf-Supervised Learning

  6. CoVisco: Codec-Native Vision Encoder with Native Token Compression for Unified Image-Video Understanding

    Sep 30, 2026Yulong Liu, Xiaotian Han, Junyuan Shang +6Video UnderstandingEfficient VLM Inference

  7. RESUME: Recurrent State Updates from Motion and Residual Signals for Efficient Video Language Modeling

    Sep 30, 2026Can Zhang, Xiaotian Han, Junyuan Shang +5Video-Language ModelsVideo Representation Learning

  8. CurvSpec: Adaptive Multi-Curvature Learning for Partial Relevant Video Retrieval

    Sep 29, 2026Zhen Liu, Letian Li, Jinpeng Wang +4Cross-Modal RetrievalVideo Representation Learning

  9. W2Rep: Learning Visual Representations by Watching the World Change

    Sep 28, 2026Wen Huang, Hang Guo, Jiarui Yang +3Video Representation LearningSelf-Supervised Visual Representation Learning

  10. Generative Uncertainty as a Self-supervised Signal for Semantic Similarity Learning

    Sep 28, 2026Enrico Pallotta, Sina Raoufi, Lars Doorenbos +2Video Representation LearningFeature Selection

  11. λλ-JEPA Spectral Anti-Collapse Regularization for Self-Supervised Learning

    Sep 28, 2026Berker Demirel, Clémentine Dominé, Valentino Maiorca +3Spectral RegularizationJoint-Embedding Predictive Architecture

  12. ReDrive: Shaping Representations with World Modeling for End-to-End Driving

    Sep 27, 2026Yueting Zhu, Shaoyu Chen, Yuehao Song +4World Model LearningEnd-to-End Autonomous Driving

  13. TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining

    Sep 27, 2026Shih-Ying Yeh, Daniel Z. Kaplan, Xuehai Wang +3Efficient ViTsVideo Representation Learning

  14. Video-STLayout Pre-training

    Sep 21, 2026Akash Abdu Jyothi, Greg MoriVideo Action RecognitionVideo Representation Learning

  15. Resolution-Flexible Decoding for Hybrid Neural Video Representations

    Sep 20, 2026Taiga Hayami, Masaya Takabe, Hiroshi WatanabeVideo Representation Learning

  16. GeoLAM: Learning Geometry-Grounded Latent Actions from Unlabeled Human Videos

    Sep 15, 2026Yifan Xie, Hekun Tian, Jinkun Liu +3Latent Action LearningGeometric Representation Learning

  17. Arti-JEPA: Adapting Video World Model to Real-Time MRI of the Vocal Tract for Speech-Production Analysis

    Sep 9, 2026Hong Nguyen, Sean Foley, Christina Hagedorn +4Unsupervised Domain AdaptationJoint-Embedding Predictive Architecture

  18. VidaForge: Open Research Infrastructure for Video Pretraining Data Recipes

    Sep 6, 2026Yan Ma, Jiadi Su, Zhulin Hu +4Vision Foundation ModelsTraining Data Curation

  19. What, Where, and How: Probing Spatiotemporal Representations in Video Foundation Models

    Sep 1, 2026Sharon S. Musa, Fereshteh Forghani, Harrish Thasarathan +3Vision Foundation ModelsVideo Representation Learning

  20. MR-JEPA: A General Purpose Video Foundation Model for Cardiac MRI

    Aug 31, 2026Athira J. Jacob, Puneet Sharma, Dorin Comaniciu +1Medical Imaging Foundation ModelsSelf-Supervised Pre-Training

  21. V-RAE: Rethinking Video Latent Spaces for Generation

    Aug 13, 2026Minghui Guo, Shengqiong Wu, Hao FeiAutoencodersVideo Representation Learning

  22. AVA-Encoder: Towards Agent-Native Video Representation Learning

    Aug 12, 2026Chuyue Li, Jinpeng Yu, Haozhe Wang +7Video Representation Learning

  23. Explicit Layer Modeling for Video Object Insertion and Video Layer Decomposition

    Jul 28, 2026Kyujin Han, Seungjoo Shin, Sunghyun ChoVideo Diffusion ModelsVideo Editing