Multimodal Grounding

Momentum

18 papers in the last four weeks, up 200% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 120

All topics
CardsList
  1. SYNCR: A Cross-Video Reasoning Benchmark with Synthetic Grounding

    May 8, 2026Sara Ghazanfari, Siddharth Garg, Prashanth Krishnamurthy +1Multimodal GroundingVideo Reasoning

  2. GeoDial: A Multimodal Conversational Tutoring Dataset for Geometry Problem-Solving with Visual Tutor Turns

    May 8, 2026Sankalan Pal Chowdhury, Junling Wang, Donya Rooein +2Multimodal GroundingIntelligent Tutoring Systems

  3. Uncovering Entity Identity Confusion in Multimodal Knowledge Editing

    May 7, 2026Shu Wu, Xiaotian Ye, Xinyu Mou +3Multimodal GroundingMultimodal Large Language Models

  4. The Cost of Context: Mitigating Textual Bias in Multimodal Retrieval-Augmented Generation

    May 7, 2026Hoin Jung, Xiaoqian WangMultimodal RAGMultimodal Grounding

  5. Mitigating Multimodal LLMs Hallucinations via Relevance Propagation at Inference Time

    May 3, 2026Itai Allouche, Joseph KeshetMultimodal GroundingInference-Time Optimization

  6. Affordance Agent Harness: Verification-Gated Skill Orchestration

    May 1, 2026Haojian Huang, Jiahao Shi, Yinchuan Li +1Multimodal GroundingLLM Agent Orchestration

  7. From Mirage to Grounding: Towards Reliable Multimodal Circuit-to-Verilog Code Generation

    Apr 30, 2026Guang Yang, Xing Hu, Xiang Chen +1Multimodal GroundingCode Generation

  8. DualFact+: A Multimodal Fact Verification Framework for Procedural Video Understanding

    Apr 28, 2026Cennet Oguz, Yasser Hamidullah, Josef van Genabith +1Multimodal GroundingVideo Understanding

  9. MEG-RAG: Quantifying Multi-modal Evidence Grounding for Evidence Selection in RAG

    Apr 27, 2026Xihang Wang, Zihan Wang, Chengkai Huang +2Multimodal RAGMultimodal IR

  10. Multimodal QUD: Inquisitive Questions from Scientific Figures

    Apr 26, 2026Yating Wu, William Rudman, Venkata S Govindarajan +2Multimodal GroundingScientific QA

  11. One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition

    Apr 25, 2026Balaji Darur, Amanmeet Garg, Makarand TapaswiMultimodal GroundingVideo Understanding

  12. CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding

    Apr 24, 2026Lihao Zheng, Zhenwei Shao, Yu Zhou +5Multimodal GroundingMultimodal Large Language Models

  13. Using Machine Mental Imagery for Representing Common Ground in Situated Dialogue

    Apr 22, 2026Biswesh Mohapatra, Giovanni Duca, Laurent Romary +1Multimodal GroundingConversational Agents

  14. SAKE: Self-aware Knowledge Exploitation-Exploration for Grounded Multimodal Named Entity Recognition

    Apr 22, 2026Jielong Tang, Xujie Yuan, Jiayang Liu +6Multimodal GroundingNamed Entity Recognition

  15. A-MAR: Agent-based Multimodal Art Retrieval for Fine-Grained Artwork Understanding

    Apr 21, 2026Shuai Wang, Hongyi Zhu, Jia-Hong Huang +6Multimodal RAGMultimodal Grounding

  16. ZSG-IAD: A Multimodal Framework for Zero-Shot Grounded Industrial Anomaly Detection

    Apr 20, 2026Qiuhui Chen, Jiaxiang Song, Shuai Tan +1Multimodal GroundingInterpretable Anomaly Detection

  17. E2E-GMNER: End-to-End Generative Grounded Multimodal Named Entity Recognition

    Apr 19, 2026Meng Zhang, Jinzhong Ning, Xiaolong Wu +2Multimodal GroundingNamed Entity Recognition

  18. BOOKAGENT: Orchestrating Safety-Aware Visual Narratives via Multi-Agent Cognitive Calibration

    Apr 17, 2026Bo Gao, Chang Liu, Yuyang Miao +2Multimodal GroundingMulti-Agent Collaboration

  19. GIST: Multimodal Knowledge Extraction and Spatial Grounding via Intelligent Semantic Topology

    Apr 16, 2026Shivendra Agrawal, Bradley HayesMultimodal GroundingEmbodied Navigation

  20. Rethinking Patient Education as Multi-turn Multi-modal Interaction

    Apr 16, 2026Zonghai Yao, Zhipeng Tang, Chengtao Lin +5Multimodal GroundingHealthcare

  21. SemConFlow: Semantic Grounding of Holistic Co-Speech Gesture Generation with Contrastive Flow-Matching

    Mar 27, 2026Lanmiao Liu, Esam Ghaleb, Aslı Özyürek +1Multimodal GroundingFlow Matching

  22. OmniACBench: A Benchmark for Evaluating Context-Grounded Acoustic Control in Omni-Modal Models

    Mar 25, 2026Seunghee Kim, Bumkyu Park, Kyudan Jung +5Multimodal GroundingUnified Multimodal Models

  23. SONIC-O1: A Real-World Benchmark for Evaluating Multimodal Large Language Models on Audio-Video Understanding

    Jan 29, 2026Ahmed Y. Radwan, Christos Emmanouilidis, Hina Tabassum +2Audio-Language Model EvaluationMultimodal Grounding

  24. GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding

    Nov 2, 2025Shijie Zhou, Viet Dac Lai, Hao Tan +4Multimodal GroundingGUI Grounding

  25. Evidence-Grounded Trustworthy Multimodal Reasoning and Evaluation Benchmark in Complex Urban Scenes

    Jun 11, 2025Zhaoyang Wei, Bowen Jiang, Xumeng Han +6Multimodal GroundingMultimodal Robustness

  26. V-Retrver: Evidence-Driven Agentic Reasoning for Universal Multimodal Retrieval

    Date pendingDongyang Chen, Chaoyang Wang, Dezhao Su +6Multimodal IRMultimodal Grounding