cs.CVOct 15, 2025

MMLongCite: A Benchmark for Evaluating Faithfulness of Long-Context Vision-Language Models

Authors: Keyan Zhou, Zecheng Tang, Lingfeng Ming, Qiguang Chen, Wangjie You, Guanghao Zhou, Dan Qiao, Zheming Yang, +4 more

Organizations: Soochow University · ByteDance · Harbin Institute of Technology · Central South University

Abstract

The rapid advancement of long-context vision language models (LCVLMs) has led to a significant expansion of their context windows. However, an extended context window does not guarantee the effective utilization of the context, posing a critical challenge for real-world applications. Current evaluations of such long-context faithfulness in multimodal settings remain limited to short contexts. To bridge this gap, we introduce MMLongCite, the first benchmark evaluating the faithfulness of LCVLMs via multimodal citation generation. MMLongCite features 2,280 examples across 8 tasks and diverse modalities (image, video, interleaved), with context lengths scaled from 16K to 128K tokens. To test spatial localization capabilities of LCVLMs, we also introduce MMLongCite-HR, evaluating fine-grained visual grounding amidst dense pixel spaces. Through extensive benchmarking of cutting-edge LCVLMs, we provide a systematic analysis of current multimodal citation capabilities. Our results reveal a significant discrepancy between answer correctness and citation faithfulness. We also conduct attention pattern investigations and in-depth error analyses to reveal the underlying phenomena of failures in LCVLMs. MMLongCite establishes a rigorous foundation for diagnosing and advancing the faithfulness of LCVLMs. We hope our findings provide meaningful insights to drive further improvements in the long-context capabilities of LCVLMs.

Figures & tables

Appendix figures & tables29 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. MMLongEmbed: Benchmarking Multimodal Embedding Models in Long-Context Scenarios

    Jun 5, 2026Haitian Wang, Ruoxi Sun, Quantong Qiu +5Multimodal EmbeddingsQuestion-Answering Benchmarks

  2. MemLens: Benchmarking Multimodal Long-Term Memory in Large Vision-Language Models

    May 14, 2026Xiyu Ren, Zhaowei Wang, Yiming Du +11Multimodal MemoryRecent Vision-Language Models

  3. Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Context

    May 13, 2026Zhaowei Wang, Lishu Luo, Haodong Duan +9Efficient Long-Context InferenceRecent Vision-Language Models