Multimodal Large Language Models

Also known as MLLM

Latest papers 676

All topics
CardsList
  1. DICA: Dual-Indicator Guided Contrastive Alignment in Multimodal Large Language Models

    Jul 27, 2026Hao Yang, Jin Wang, Xuejie ZhangLLM AlignmentMultimodal Large Language Models

  2. OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models

    Jul 25, 2026Jinsen Su, Yongdong Luo, Yuexiao Ma +3Efficient Multimodal InferenceAudio Token Compression

  3. Filling Before Advancing: Capability-Gap-Driven Post-Training for Scenario-Specialized Remote Sensing MLLMs

    Jul 24, 2026Yuheng Zong, Minghua Wang, Xin Zhao +3Remote Sensing Image UnderstandingFine-Tuning

  4. Zero-Shot Mission-Level Evaluation for Aerial MLLM Agents

    Jul 24, 2026Suman Navaratnarajah, Taehyoung Kim, Jona Ruthardt +5Zero-Shot LearningMultimodal Large Language Models

  5. Toward Generalizable Cognitive Impairment Detection with Speech-Based Multimodal Large Language Models

    Jul 23, 2026Yingchao Huang, Xin Wang, Yuhan Su +1Multimodal Large Language ModelsSpeech-Based Dementia Detection

  6. Unlearning Under Imbalance: Benchmarking Fairness in Multimodal LLM Unlearning

    Jul 23, 2026Lorenzo Orsingher, Thomas De Min, Massimiliano Mancini +2VLM UnlearningAlgorithmic Fairness

  7. Out of Sight, Still in Mind: Token Compression for Omni-LLMs

    Jul 23, 2026Suho Yoo, Youngjoon Jang, Hyebin Cho +1Efficient Multimodal InferenceMultimodal Large Language Models

  8. Geo3R: Mitigating Spatial Reasoning Hallucination in Multimodal Large Language Models

    Jul 23, 2026Mingyu Wang, Weilin Jin, Wenbo Li +33D Spatial ReasoningHallucination in Language Models

  9. EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent Specialization

    Jul 23, 2026Lihuang Fang, Yuchen Zou, kebing Jin +1RL for Language ModelsAgentic RL

  10. Enhancing Explainable Cardiac Diagnosis with Guide-Grounded Multimodal LLMs

    Jul 23, 2026Hai-Nam Duy Vuong, Duy-Anh Bui, Trong-Nghia Nguyen +5Multimodal GroundingMedical Report Generation

  11. Look Less, Think Faster: Joint Token-Compute Adaptation for Multimodal LLMs

    Jul 22, 2026Pengcheng Wang, Zhiquan Wang, Jayoung Lee +5Efficient Multimodal InferenceEfficient VLM Inference

  12. Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?

    Jul 22, 2026Qiwei Ma, Chunping Qiu, Xinjun Cheng +5Remote Sensing Image UnderstandingRemote Sensing VQA

  13. OmniReasoner: Thinking with Long Audio-Video via Native Tool Use

    Jul 21, 2026Yu Chen, Caorui Li, Ziyu Xiong +8Temporal Video GroundingAudio-Visual Understanding

  14. MeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party Meetings

    Jul 21, 2026Ziyi Wang, Yuhang Wu, Dongxu Piao +3Theory of MindMultimodal Large Language Models

  15. Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges

    Jul 21, 2026Tuo Liang, Zhe Hu, Disheng Liu +2Multimodal Large Language ModelsComputational Creativity

  16. In-Context Learning for Wound Classification with Small Multimodal Language Models

    Jul 21, 2026George Martvel, Oskar Gustafsson, John Pavia +1Multimodal ICLMultimodal Large Language Models

  17. Querying Multimodal Scientific Papers with AI: Practices and Preferences Across Blind, Low-Vision, and Sighted Scientists

    Jul 20, 2026Arnavi Chheda-Kothary, Lucy Lu Wang, Joseph Chee Chang +1Multimodal Large Language ModelsScientific QA

  18. PGN: Design and Implementation of a Vision-Language Navigation System Based on Pangu Multimodal Foundation Model

    Jul 20, 2026Li Xian, Mingxi Li, Yizheng Wang +3VLM AdaptationMultimodal Large Language Models

  19. Pailitao-MMSearch: Building Native E-Commerce Multimodal Search Foundation

    Jul 20, 2026Xiaohan Ye, Xu Chen, Zihan Gong +15Multimodal IRMultimodal Pretraining

  20. TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs

    Jul 19, 2026Yuhan Zhu, Changlian Ma, Xiangyu Zeng +12Temporal Video GroundingReinforcement Learning

  21. SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing

    Jul 19, 2026Kaiwen Jing, Ruixu Jia, Bingyao Li +3Referring Video Object SegmentationMultimodal Large Language Models

  22. Multi-Dimensional Quality Assessment for AI-Generated Human-Centric Videos: Dataset and Model

    Jul 18, 2026Sijing Wu, Yunhao Li, Huiyu Duan +4Mixture of ExpertsMultimodal Large Language Models

  23. Can Multimodal Large Language Models Understand OCT?

    Jul 18, 2026Baochen Fu, Wenzhi Deng, Baihao Jin +5Medical ImagingMedical Image Analysis

  24. Knowing the Self, Understanding the World: A Dual-Cognition Benchmark for UAV Spatio-temporal Reasoning with MLLMs

    Jul 17, 2026Like Liu, Zhengzheng Xu, Haitao He +3Multimodal GroundingAerial Robotics

  25. An Exam for Active Observers

    Jul 17, 2026Jiarui Zhang, Muzi Tao, Shangshang Wang +3VLM EvaluationMultimodal Large Language Models