Video Retrieval

Momentum

3 papers in the last four weeks, level with the four weeks before. 0.0% of all new papers.

Jul 13Week of Sep 28

Latest papers 41

All topics
CardsList
  1. VEDJE: Video-Efficient Discriminative Joint Encoder for Scalable Video-Text Retrieval

    Oct 8, 2026Shahaf Wagner, Gabriele Serussi, Dan Ben Ami +2Video RetrievalVision-Language Retrieval

  2. VideoEvolve: Co-Evolving Memory and Retrieval for Long Video Understanding

    Oct 7, 2026Yongchao Xu, Bowen Ye, Jiefeng Gan +5Long-Video UnderstandingMemory-Augmented Video Understanding

  3. SoccerNet-FoulRet: Retrieving Semantically Similar Soccer Foul Videos

    Oct 7, 2026Jacobus Arthur, Ahmad Sait, Batool Hani +6Video RetrievalSports Video Analysis

  4. Concept Driven Domain Adaptation: Finding an Abstract Needle in a Haystack

    Oct 1, 2026Haiming Zhao, Tai Wang, Kun Zhang +2VLM AdaptationVision-Language Alignment

  5. MultiVENT-Raw: A Benchmark for Retrieval and Reasoning over Raw Videos

    Sep 23, 2026Reno Kriz, David Etter, Alexander Martin +11Video UnderstandingVideo Retrieval

  6. Concentrate After Imagination: Text-Conditioned Evidence Grounding for Partially Relevant Video Retrieval

    Sep 8, 2026Shuaiqi Cheng, Siyu You, Yanbi Wu +3Temporal Video GroundingInformation Retrieval

  7. Allocate Before You Embed: Adaptive Visual Input Allocation for Video Embeddings

    Sep 1, 2026Song Jin, Zhongtao Jiang, Chenglei Shen +5Multimodal EmbeddingInformation Retrieval

  8. From Intent to Evidence: Policy-Steered Multi-Strategy Retrieval for Long-Video Agents

    Aug 31, 2026Can Zhang, Baofeng Zhang, Xiaotian Han +5Long-Video UnderstandingLong-Video QA

  9. TraVEL: Trajectory-Guided Video Embedding Learning for Driving-Video Retrieval

    Aug 13, 2026Yi-Chung Chen, Philip Jacobson, Tom Lampo +6VLMs for Autonomous DrivingMultimodal Embedding

  10. Keep It Simple: Multi-Key Episodic Memory Retrieval for Ultra-Long Video Understanding

    Aug 7, 2026Yeeun Choi, Youngbeom Yoo, Joon-Young Lee +2Memory-Augmented VLMsEpisodic Memory

  11. Predict, Then Retrieve: Cross-Instance Future-State Retrieval from Video Prefixes

    Aug 5, 2026Quynh Vo, Thong Nguyen, Vinh-Hien Do +2Cross-Modal RetrievalVideo Prediction

  12. Prompting-MammAlps: Fine-Grained Text-to-Video Retrieval for Camera-Trap Data

    Jul 10, 2026Valentin Gabeff, Baptiste Maquignaz, Jennifer Shan +5Vision-Language RetrievalWildlife Monitoring

  13. QSVideo: Query-Conditioned Semantic Temporal Retrieval for Video Understanding

    Jul 6, 2026Wei Ao, Lan Wang, Vishnu Naresh BoddetiTemporal Video GroundingStreaming Video Understanding

  14. VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement

    Jul 1, 2026Seohyun Lee, Seoung Choi, Dohwan Ko +2Temporal Video GroundingTemporal IR

  15. Vortex: Multi-Modal Fusion System for Intelligent Video Retrieval

    Jun 18, 2026Duc-Tho Nguyen, Hieu-Hoc Tran-Minh, Khanh-Hoa Lam +4Multimodal IRInformation Retrieval

  16. Rethinking RAG in Long Videos: What to Retrieve and How to Use It?

    Jun 11, 2026Yuho Lee, Jisu Shin, Nicole Hee-Yeon Kim +5Multimodal RAGEgocentric Video Understanding

  17. Findings of the MAGMaR 2026 Shared Task

    Jun 10, 2026Alexander Martin, Dengjia Zhang, Joel Brogan +7Multimodal RAGMultimodal Grounding

  18. MAVIS: Multi-Agent Video Retrieval via Structured Video Understanding

    Jun 8, 2026Jie Zhang, Qilang Ye, Hao Zhou +2Video UnderstandingMulti-Agent Collaboration

  19. Driving Video Retrieval for Complex Queries with Structured Grounding

    Jun 8, 2026Manyi Yao, Sparsh Garg, Christian Shelton +2Autonomous DrivingVision-Language Retrieval

  20. To Be Multimodal or Not to Be: Query-Adaptive Audio-Visual Person Retrieval via Active Modality Detection

    Jun 4, 2026Erfan Loweimi, Mengjie Qian, Kate Knill +7Missing-Modality LearningAudio-Visual Understanding

  21. VidMsg: A Benchmark for Implicit Message Inference in Short Videos

    Jun 2, 2026Issar Tzachor, Michael Green, Rami Ben-AriVLM EvaluationVideo QA

  22. Reason-Then-Retrieve for CoVR-R with Structured Edit Prompts and Dense-Sparse Fusion

    Jun 1, 2026DongQing Liu, MengShi Qi, HongWei JiCross-Modal RetrievalVideo Reasoning

  23. Training-Free Composed Video Retrieval via Visual Representation-Guided Video-LLM Reasoning

    Jun 1, 2026Yang Liu, Qianqian Xu, Peisong Wen +2Cross-Modal RetrievalComposed Video Retrieval

  24. R^3: Composed Video Retrieval via Reasoning-Guided Recalling and Re-ranking

    May 31, 2026Zixu Li, Yupeng Hu, Zhiheng Fu +3Cross-Modal RetrievalComposed Video Retrieval

  25. Dual-Route Top-K Retrieval with 1v1 VLM Reranking for the CoVR-R

    May 31, 2026Yuyang Sun, Yongliang Wu, Xingyu Zhu +8Multimodal RerankingCross-Modal Retrieval

  26. Not All Inputs Are Valid: Towards Open-Set Video Moment Retrieval Using Language

    May 28, 2026Xiang Fang, Wanlong Fang, Daizong Liu +8Temporal Video GroundingOpen-Set Recognition