MMLongCite: A Benchmark for Evaluating Faithfulness of Long-Context Vision-Language Models
Organizations: Soochow University · ByteDance · Harbin Institute of Technology · Central South University
Abstract
The rapid advancement of long-context vision language models (LCVLMs) has led to a significant expansion of their context windows. However, an extended context window does not guarantee the effective utilization of the context, posing a critical challenge for real-world applications. Current evaluations of such long-context faithfulness in multimodal settings remain limited to short contexts. To bridge this gap, we introduce MMLongCite, the first benchmark evaluating the faithfulness of LCVLMs via multimodal citation generation. MMLongCite features 2,280 examples across 8 tasks and diverse modalities (image, video, interleaved), with context lengths scaled from 16K to 128K tokens. To test spatial localization capabilities of LCVLMs, we also introduce MMLongCite-HR, evaluating fine-grained visual grounding amidst dense pixel spaces. Through extensive benchmarking of cutting-edge LCVLMs, we provide a systematic analysis of current multimodal citation capabilities. Our results reveal a significant discrepancy between answer correctness and citation faithfulness. We also conduct attention pattern investigations and in-depth error analyses to reveal the underlying phenomena of failures in LCVLMs. MMLongCite establishes a rigorous foundation for diagnosing and advancing the faithfulness of LCVLMs. We hope our findings provide meaningful insights to drive further improvements in the long-context capabilities of LCVLMs.
Figures & tables
| Tasks | Source | Context | Length Distribution | Total | |||
| 0 16k | 16 32k | 32 64k | 64 128k | ||||
| Single-Source Visual QA | |||||||
| LongDocURL | ( Deng et al., 2025 ) | Image | 80 | 80 | 80 | 80 | 320 |
| MMLongBench-Doc | ( Ma et al., 2024 ) | Image | 80 | 80 | 80 | 80 | 320 |
| Multi-Source Visual QA | |||||||
| SlideVQA | ( Tanaka et al., 2023 ) | Image | 70 | 70 | 70 | 70 | 280 |
| Models | Single-Source Visual QA | Multi-Source Visual QA | Visual Grounding | Video Understanding | ||||||||||||
| Proprietary Models | ||||||||||||||||
| Gemini-3-Pro(think) | 84.49 | 83.58 | 83.36 | 80.86 | 83.75 | 84.97 | 84.02 | 85.85 | 80.48 | 72.54 | 75.07 | 87.28 | 77.97 | 69.39 | 71.75 | 85.31 |
| Gemini-3-Flash(think) | 79.61 | 83.96 | 81.06 | 82.73 | 86.22 | 89.13 | 87.22 | 85.40 | 78.54 | 86.29 | 81.32 | 84.52 | 72.40 | 74.24 | 71.74 | 84.69 |
| Seed-2.0-Pro | 76.26 | 68.38 | 70.57 | 73.67 | 81.70 | 74.34 | 76.56 | 78.71 | - | - | - | - | 73.14 | 46.34 | 54.63 | 82.87 |
| Seed-2.0-Lite | 53.41 | 51.50 | 51.22 | 71.30 | 65.78 | 68.41 | 66.08 | 84.71 | - | - | - | - | 50.74 | 40.33 | 42.72 | 79.72 |
| Models | LongDocURL | 2WikiMultiHopQA | Visual Haystack | LongVideoBench | ||||||||||||
| Easy | Hard | Easy | Hard | Easy | Hard | Easy | Hard | Easy | Hard | Easy | Hard | Easy | Hard | Easy | Hard | |
| Gemini-3-Pro | 87.07 | 85.78 | 76.68 | 83.12 | 78.08 | 76.75 | 67.56 | 80.62 | ||||||||
| 84.68 | 63.78 | 85.31 | 70.31 | 78.17 | 76.77 | 83.75 | 85.47 | 77.11 | 74.81 | 74.12 | 62.75 | 65.02 | 67.51 | 79.06 | 75.62 | |
| Qwen2.5-VL-7B-Ins | 25.69 | 49.06 | 21.18 | 41.25 | 28.08 | 52.62 | 36.92 | 65.31 | ||||||||
| 15.19 | 12.81 | 50.00 | 45.62 | 11.25 | 13.22 | 38.44 | 38.91 | 21.38 | 26.15 | 54.00 | 53.12 | 30.12 | 26.34 | 61.25 | 61.25 | |
| Model | Version | Ctx. Size | #Param | Architecture | Open-source | Think |
| Gemini-3-Pro [ 18 ] | gemini-3-pro-preview | 1M | ✗ | |||
| Gemini-3-Flash [ 19 ] | gemini-3-flash-preview | 1M | ✗ | |||
| Seed-2.0-Pro [ 6 ] | doubao-seed-2-0-pro-260215 | 256K | ✗ | |||
| Seed-2.0-Lite [ 6 ] | doubao-seed-2-0-lite-260215 | 256K | ✗ | |||
| Seed-2.0-Mini [ 6 ] | doubao-seed-2-0-mini-260215 | 256K | ✗ | |||
| Seed-1.6 [ 5 ] | doubao-seed-1-6-251015 | 256K | ✗ |
Appendix figures & tables29 assets
Supplementary material from the paper’s appendix.
Appendix
| Models | LongDocURL | MMLongBench-Doc | SlideVQA | 2WikiMultiHopQA | ||||||||||||
| Proprietary Models | ||||||||||||||||
| Gemini-3-Pro | 88.07 | 87.35 | 87.07 | 85.78 | 80.91 | 79.82 | 79.65 | 75.94 | 91.05 | 92.25 | 91.35 | 88.57 | 76.46 | 77.69 | 76.68 | 83.12 |
| Gemini-3-Flash | 82.03 | 87.04 | 83.81 | 87.50 | 77.19 | 80.89 | 78.30 | 77.97 | 89.97 | 94.10 | 91.48 | 90.18 | 82.46 | 84.16 | 82.95 | 80.62 |
| Seed-2.0-Pro | 79.04 | 69.48 | 72.16 | 79.69 | 73.47 | 67.27 | 68.98 | 67.66 | 89.67 | 85.81 | 86.97 | 86.79 | 73.72 | 62.87 | 66.16 | 70.62 |
| Seed-2.0-Lite | 56.31 | 54.27 | 53.90 | 77.44 | 50.52 | 48.74 | 48.54 | 65.16 | 67.95 | 69.30 | 67.74 | 89.11 | 63.61 | 67.53 | 64.42 | 80.31 |
| Models | Visual Haystack | MM-NIAH | Video-MME | LongVideoBench | ||||||||||||
| Proprietary Models | ||||||||||||||||
| Gemini-3-Pro | 81.08 | 76.74 | 78.08 | 76.75 | 79.88 | 68.34 | 72.05 | 97.81 | 81.17 | 73.86 | 75.95 | 90.00 | 74.78 | 64.92 | 67.56 | 80.62 |
| Gemini-3-Flash | 69.43 | 80.60 | 73.44 | 73.25 | 87.65 | 91.98 | 89.20 | 95.78 | 72.27 | 76.99 | 73.31 | 91.25 | 72.52 | 71.49 | 70.17 | 78.12 |
| Seed-2.0-Pro | - | - | - | - | - | - | - | - | 72.72 | 47.50 | 55.40 | 89.17 | 73.55 | 45.19 | 53.85 | 76.56 |
| Seed-2.0-Lite | - | - | - | - | - | - | - | - | 49.25 | 38.56 | 40.87 | 86.94 | 52.23 | 42.10 | 44.58 | 72.50 |
| Metric | Inter-Annotator Agreement (Krippendorff’s ) | Human-Model Agreement | ||
| Qwen2.5-VL-7B | Doubao-1.6 | Qwen2.5-VL-7B | Doubao-1.6 | |
| Citation Precision | 0.8578 | 0.8361 | 0.9095 (Spearman) | 0.9005 (Spearman) |
| Citation Recall | 0.9218 | 0.8977 | 0.9272 (Spearman) | 0.9196 (Spearman) |
| Correctness | 0.9132 | 0.9726 | 0.9074 (Kappa) | 0.9393 (Kappa) |
| Models | Single-Source Visual QA | Multi-Source Visual QA | Visual Grounding | Video Understanding | ||||||||||||
| Proprietary Models | ||||||||||||||||
| Gemini-3-Pro(think) | 87.06 | 86.14 | 85.86 | 85.31 | 83.49 | 85.01 | 83.90 | 91.07 | 80.20 | 71.88 | 74.63 | 81.34 | 78.71 | 69.56 | 72.19 | 86.72 |
| Gemini-3-Flash(think) | 82.95 | 87.20 | 84.34 | 86.80 | 87.29 | 90.44 | 88.44 | 91.14 | 82.65 | 87.81 | 84.54 | 78.44 | 69.90 | 73.44 | 69.77 | 86.25 |
| Seed-2.0-Pro | 81.47 | 73.06 | 75.15 | 81.02 | 84.97 | 76.60 | 79.29 | 84.90 | - | - | - | - | 70.68 | 45.34 | 52.92 | 84.59 |
| Seed-2.0-Lite | 59.90 | 56.82 | 57.02 | 80.08 | 68.70 | 70.37 | 68.63 | 89.62 | - | - | - | - | 44.71 | 36.89 | 38.17 | 81.14 |
| Judge | Evaluated model | CP | CR | Correctness |
| GPT-5.2 | Qwen2.5-VL-7B | 0.9095 | 0.9272 | 0.9074 |
| Qwen3.6-27B | Qwen2.5-VL-7B | 0.9322 | 0.9385 | 0.9124 |
| GPT-5.2 | Doubao-1.6 | 0.9005 | 0.9196 | 0.9393 |
| Qwen3.6-27B | Doubao-1.6 | 0.8161 | 0.8489 | 0.7948 |
| Metric | Spearman | Kendall | Qwen GPT mean shift |
| Citation F1 | 0.999 | 0.989 | |
| Correctness | 0.993 | 0.959 |
| Metric | Qwen-family | Non-Qwen |
| Citation F1 | ||
| Correctness |
| Model | Task | Citation F1 [95% CI] | Correctness [95% CI] |
| Gemini-3-Pro | Single-Source Visual QA | 83.36 [81.19, 85.41] | 80.86 [77.97, 83.67] |
| Multi-Source Visual QA | 84.02 [81.99, 85.90] | 85.85 [83.07, 88.46] | |
| Visual Grounding | 75.07 [72.74, 77.39] | 87.28 [85.08, 89.42] | |
| Video Understanding | 71.75 [68.64, 74.90] | 85.31 [81.25, 89.06] | |
| Qwen2.5-VL-7B | Single-Source Visual QA | 28.63 [25.69, 31.50] | 46.09 [42.42, 49.77] |
| Multi-Source Visual QA | 33.10 [30.21, 35.96] | 52.59 [48.79, 56.36] |
| Model | Task | Citation F1 [95% CI] | Correctness [95% CI] |
| Gemini-3-Pro | Single-Source Visual QA | ||
| Multi-Source Visual QA | |||
| Visual Grounding | |||
| Video Understanding | |||
| Qwen2.5-VL-7B | Single-Source Visual QA | ||
| Multi-Source Visual QA |
| Model | Metric | Middle edge difference [95% CI] |
| Gemini-3-Pro | Citation F1 | |
| Gemini-3-Pro | Correctness | |
| Qwen2.5-VL-7B | Citation F1 | |
| Qwen2.5-VL-7B | Correctness |
| Model | Max citations | CP | CR | F1 | Correctness |
| Qwen2.5-VL-7B | 1 | 33.69 | 28.86 | 30.37 | 54.31 |
| 3 | 30.75 | 32.24 | 29.64 | 54.59 | |
| 5 | 30.49 | 31.32 | 29.00 | 54.62 | |
| 10 | 28.54 | 30.41 | 27.30 | 55.25 | |
| Qwen3-VL-30B-A3B | 1 | 61.51 | 50.99 | 54.27 | 68.86 |
| 3 | 65.59 | 53.16 | 56.61 | 69.46 |
| Model | LongDocURL | MMLongBench- Doc | SlideVQA | 2WikiMultiHopQA |
| Qwen2.5-VL-7B-Instruct | 16.56 | 5.00 | 15.71 | 9.38 |
| Qwen3-VL-30B-A3B-Instruct | 19.84 | 8.44 | 18.21 | 20.00 |
| Model | Citation F1 | Correctness | ||
| Image | Text | Image | Text | |
| Qwen2.5-VL-7B | 21.18 | 23.08 | 41.25 | 39.06 |
| Qwen3-VL-30B-A3B | 59.36 | 53.75 | 61.72 | 62.97 |
| Setting | Above Qwen2.5-VL limit | Above Qwen3-VL limit |
| 4-image grids | 0.00% | 0.00% |
| 16-image grids | 4.11% | 0.48% |