Vision-Language Retrieval

Latest papers 62

All topics
CardsList
  1. VCG: A Multimodal Retrieval Framework for E-Commerce Video Feeds under Extreme Cold-Start Conditions

    Jun 17, 2026Katya Mirylenka, Egor Malykh, Mahdyar Ravanbakhsh +10Multimodal IRCold-Start Recommendation

  2. ReSiReg: Towards Spatially Consistent Semantics in Language-Conditioned Robotic Tasks

    Jun 17, 2026Simon Schwaiger, David Seyser, Alessandro Scherl +2Vision-Language ModelsVision-Language Grounding

  3. DREAM: Extending Vision-Language Models with Dual-Objective Encoding for Cross-Modal Retrieval

    Jun 17, 2026Kaleem Ullah, Altaf Hussain, Muhammad Munsif +1Cross-Modal RetrievalVideo Representation Learning

  4. LARE: Low-Attention Region Encoding for Text-Image Retrieval

    Jun 17, 2026Abdulmalik Alquwayfili, Faisal Almeshal, Jumanah Almajnouni +8Visual Representation LearningFine-Grained Image Retrieval

  5. FARM: Find Anything using Relational Spatial Memory

    Jun 13, 2026Siming He, Leo Huang, Adam Lilja +6Visual Spatial ReasoningRelational Reasoning

  6. Beyond Scalar Distances: Semantic Attribute Gradients from Frozen MLLMs for Visual Embeddings

    Jun 13, 2026Shubhang Bhatnagar, Dheeraj Baiju, Narendra AhujaVisual Representation LearningFine-Grained Image Retrieval

  7. MAVIS: Multi-Agent Video Retrieval via Structured Video Understanding

    Jun 8, 2026Jie Zhang, Qilang Ye, Hao Zhou +2Video UnderstandingMulti-Agent Collaboration

  8. Driving Video Retrieval for Complex Queries with Structured Grounding

    Jun 8, 2026Manyi Yao, Sparsh Garg, Christian Shelton +2Autonomous DrivingVision-Language Retrieval

  9. Decoupling Semantics and Logic: A Training-Free Coarse-to-Fine Pipeline for Video Retrieval-Augmented Generation

    Jun 6, 2026Jiaxin Dai, Zehang Wei, Jiamin Yan +1Multimodal RAGAgentic RAG

  10. TEVI: Text-Conditioned Editing of Visual Representations via Sparse Autoencoders for Improved Vision-Language Alignment

    Jun 5, 2026Sweta Mahajan, Sukrut Rao, Jiahao Xie +2Vision-Language RetrievalVision-Language Alignment

  11. VidMsg: A Benchmark for Implicit Message Inference in Short Videos

    Jun 2, 2026Issar Tzachor, Michael Green, Rami Ben-AriVLM EvaluationVideo QA

  12. Reason-Then-Retrieve for CoVR-R with Structured Edit Prompts and Dense-Sparse Fusion

    Jun 1, 2026DongQing Liu, MengShi Qi, HongWei JiCross-Modal RetrievalVideo Reasoning

  13. Dual-Route Top-K Retrieval with 1v1 VLM Reranking for the CoVR-R

    May 31, 2026Yuyang Sun, Yongliang Wu, Xingyu Zhu +8Multimodal RerankingCross-Modal Retrieval

  14. Can Retrieval Heads See Images? Multimodal Retrieval Heads in Long-Context Vision-Language Models

    May 26, 2026Aaron Branson Cigres Li, Zhaowei Wang, Yu Zhao +9Vision-Language ModelsVLM Interpretability

  15. Supervised Classification Heads as Semantic Prototypes: Unlocking Vision-Language Alignment via Weight Recycling

    May 21, 2026David Méndez, Roberto Confalonieri, Natalia Díaz RodríguezCross-Modal LearningVLM Adaptation

  16. USV: Towards Understanding the User-generated Short-form Videos

    May 20, 2026Haoyue Cheng, Su Xu, Liwei Jin +3Vision-Language Retrieval

  17. What Matters for Grocery Product Retrieval with Open Source Vision Language Models

    May 18, 2026Emmanuel G. Maminta, Rowel O. AtienzaVLM EvaluationFine-Grained Image Retrieval

  18. Memory-Augmented Query Intent Understanding for Efficient Chat-based Image Retrieval

    May 17, 2026Xianke Chen, Daizong Liu, Yushuo Lou +5Vision-Language RetrievalImage Retrieval

  19. Neuroscience-inspired Staged Representation Learning with Disentangled Coarse- and Fine-Grained Semantics for EEG Visual Decoding

    May 16, 2026Xiang Gao, Hui Tian, Yanming Zhu +2Disentangled Representation LearningCross-Modal Representation Learning

  20. GeoFlowVLM: Geometry-Aware Joint Uncertainty for Frozen Vision-Language Embedding

    May 13, 2026Mayank Nautiyal, Li Ju, Andreas Hellander +2Cross-Modal LearningFlow Matching

  21. From Clouds to Hallucinations: Atmospheric Retrieval Hijacking in Remote Sensing Vision-Language RAG

    May 8, 2026Jiaju Han, Chao Li, Chengyin Hu +8Multimodal RAGRetrieval-Augmented Generation

  22. Zero-Shot Satellite Image Retrieval through Joint Embeddings: Application to Crisis Response

    May 6, 2026James Walsh, William Fawcett, Grace Colverd +1Geospatial Foundation ModelsDisaster Response

  23. Open-SAT: LLM-Guided Query Embedding Refinement for Open-Vocabulary Object Retrieval in Satellite Imagery

    May 6, 2026Md Adnan Arefeen, Biplob Debnath, Ravi K. Rajendran +2Remote Sensing Image UnderstandingImage-Text Retrieval

  24. Interactive Multi-Turn Retrieval for Health Videos

    May 2, 2026Chengzheng Wu, Ke Qiu, Baoming Zhang +3HealthcareCross-Modal Retrieval

  25. UniCVR: From Alignment to Reranking for Unified Zero-Shot Composed Visual Retrieval

    Apr 22, 2026Haokun Wen, Xuemeng Song, Haoyu Zhang +3Cross-Modal LearningLearning to Rank

  26. KIRA: Knowledge-Intensive Image Retrieval and Reasoning Architecture for Specialized Visual Domains

    Apr 18, 2026Parthaw Goswami, Jaynto Goswami DeepMultimodal RAGVisual Reasoning

  27. LaVPR: Benchmarking Language and Vision for Place Recognition

    Feb 3, 2026Ofer Idan, Dan Badur, Yosi Keller +1Multimodal RobustnessVisual Place Recognition

  28. R3G: A Reasoning-Retrieval-Reranking Framework for Vision-Centric Answer Generation

    Jan 25, 2026Zhuohong Chen, Zhengxian Wu, Zirui Liao +6Multimodal RAGVisual Question Answering