TopoGraphRAG-Bench: Evaluating Multimodal GraphRAG on Layout-Grounded Evidence Reasoning
Organizations: Department of Computer Science, North Carolina State University, USA · Independent Researcher · Department of Information Science, University of North Texas, USA · The Anuradha and Vikas Sinha Department of Data Science, University of North Texas, USA · School of Computing, University of Wyoming, USA
Abstract
Real-world documents distribute evidence across text, tables, figures, and captions within complex page layouts. Answering complex questions over such documents therefore requires more than retrieving relevant passages: systems must recover the evidence topology that connects heterogeneous evidence units. Existing GraphRAG evaluations remain largely text-centered, while multimodal document RAG benchmarks assess cross-modal retrieval and generation without directly evaluating recovery of the intended evidence topology. We introduce TOPOGRAPHRAG-BENCH, a layout-grounded benchmark for multimodal evidence reasoning in GraphRAG, comprising 2,024 questions over 201 long, visually rich documents. Questions are constructed bottom-up from text, figure, and table evidence units under three controlled topologies: single-hop retrieval, bridge-chain reasoning, and multi-source synthesis. To ensure that questions preserve their intended structure, we apply counterfactual validation for shortcut resistance, modality necessity, and evidence necessity. We evaluate text-only GraphRAG, page-level visual retrieval, and multimodal GraphRAG systems using retrieval, generation, and topology-aware reasoning metrics. Multimodal GraphRAG systems achieve the strongest overall performance, but still fail when visual-textual evidence alignment or multi-unit composition is incomplete. Text-only GraphRAG struggles when key dependencies are grounded in figures or tables, while page-level visual retrieval lacks the fine-grained structure needed for topology recovery. These findings motivate GraphRAG systems that move beyond text-derived entity relation graphs to explicitly model document layouts, cross-modal evidence alignment, and the reasoning roles of evidence units. Code and data are available at https://richardlrc.github.io/TopoGraphRAG-Bench/.
Figures & tables
| Statistic | Number |
|---|---|
| Documents | 201 |
| - Pages/doc | 45.5 / 27 / 318 |
| - Evidence units/doc | 335.1 / 207 / 2,423 |
| Total Questions | 2,024 |
| - Single-hop | 302 (14.9%) |
| - Bridge-chain | 1,014 (50.1%) |
| System | Retrieval | Generation | ||||
|---|---|---|---|---|---|---|
| Context Precision | Context Recall | Answer Accuracy | Faithfulness | Response Relevancy | Step Coverage | |
| LightRAG | 0.6952 | 0.6353 | 0.4045 | 0.7317 | 0.7269 | 0.6031 |
| HippoRAG | 0.3834 | 0.4353 | 0.3388 | 0.6517 | 0.5924 | 0.3952 |
| Microsoft GraphRAG | 0.6018 | 0.5839 | 0.3111 | 0.6802 | 0.6754 | 0.5402 |
| VisRAG | – | – | 0.3459 | – | 0.7688 | 0.4337 |
| RAG-Anything | 0.8093 | 0.7215 | 0.5300 | 0.8346 | 0.7808 | 0.7010 |
| Composition | LightRAG | HippoRAG | Microsoft GraphRAG | VisRAG | RAG-Anything | MegaRAG |
|---|---|---|---|---|---|---|
| Text-only | 46.6 | 44.4 | 45.9 | 34.1 | 54.5 | 51.0 |
| Figure-only | 36.3 | 29.2 | 31.9 | 30.3 | 46.2 | 46.7 |
| Table-only | 30.1 | 22.8 | 27.5 | 22.1 | 34.4 | 38.4 |
| Cross-modal | 27.4 | 18.4 | 27.2 | 22.7 | 43.2 | 42.7 |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| System | Input unit | Indexing / graph construction | Retrieval configuration | Model / retriever |
|---|---|---|---|---|
| LightRAG | Text layout chunks | Default LightRAG indexing pipeline for constructing a text-derived graph over chunks, entities, and relations | Mixed retrieval mode; top- | Qwen3-30B-A3B-Instruct; Qwen3-Embedding-8B |
| HippoRAG | Text layout chunks | Default HippoRAG indexing pipeline over text chunks and its native retrieval graph | Default retrieval; top- | Qwen3-30B-A3B-Instruct; Qwen3-Embedding-8B |
| MS-GraphRAG | Concatenated text file | Default Microsoft GraphRAG indexing workflow for text-unit chunking, entity and relationship extraction, and community construction | Local search; top-3 entities and top-3 relationships; max context tokens 12000 | Qwen3-30B-A3B-Instruct; Qwen3-Embedding-8B |
| VisRAG | Page images | Default VisRAG page-image indexing using VisRAG-Ret embeddings; no entity-relation graph is constructed. | Cosine page retrieval; top- pages. | Qwen3-VL-30B-A3B-Instruct; VisRAG-Ret |
| RAG-Anything | Multimodal layout chunks | Default RAG-Anything multimodal indexing pipeline, with LightRAG as the internal graph backbone and native processing for text, figures, and tables | Mixed retrieval mode; top- | Qwen3-VL-30B-A3B-Instruct; Qwen3-VL-Embedding-8B |
| MegaRAG | Page-content JSON | Default MegaRAG multimodal knowledge-graph construction over page-level text, figure, table, and image representations | Mixed retrieval mode; chunk top- ; reranking disabled | Qwen3-VL-30B-A3B-Instruct; Qwen3-VL-Embedding-8B |
| Runtime (s/doc) | LLM Usage (/doc) | Graph / Index Size (/doc) | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| System | Indexing | Retrieval | Generation | Calls | Prompt Tok. | Completion Tok. | Entities | Relationships | Chunks | Communities |
| LightRAG | 2347.73 | 3.81 | 8.94 | 206.43 | 848415.46 | 100313.98 | 847.55 | 865.26 | 116.55 | – |
| HippoRAG | 117.95 | 2.47 | 3.85 | 552.51 | 255395.94 | 56394.44 | 1686.16 | 2108.58 | 266.32 | – |
| MS-GraphRAG | 1942.94 | 50.82 | 11.90 | 531.37 | 909167.87 | 412309.60 | 667.77 | 1004.43 | 94.24 | 115.56 |
| RAG-Anything | 5844.99 | 18.74 | 21.70 | 668.80 | 2470006.99 | 286043.06 | 1542.71 | 2786.12 | 287.04 | – |
| VisRAG | 2.94 | 0.15 | 22.87 | – | – | – | – | – | – | – |
| System | Qwen3.5 | GLM-5.1 | GPT-5.1 |
|---|---|---|---|
| LightRAG | 0.6031 | 0.6121 | 0.6743 |
| HippoRAG | 0.3952 | 0.4153 | 0.4858 |
| Microsoft GraphRAG | 0.5402 | 0.5721 | 0.6029 |
| VisRAG | 0.4337 | 0.5024 | 0.4783 |
| RAG-Anything | 0.7010 | 0.7349 | 0.7442 |
| MegaRAG | 0.7329 | 0.7723 | 0.7601 |
| Overall | Bridge-chain | Synthesis | ||||
|---|---|---|---|---|---|---|
| Judge pair | Agreement | Agreement | Agreement | |||
| Qwen3.5 / GLM-5.1 | 0.8592 | 0.7208 | 0.8663 | 0.7328 | 0.8529 | 0.7105 |
| Qwen3.5 / GPT-5.1 | 0.8512 | 0.7054 | 0.8479 | 0.7002 | 0.8541 | 0.7103 |
| GLM-5.1 / GPT-5.1 | 0.8958 | 0.7923 | 0.9028 | 0.8039 | 0.8896 | 0.7814 |
| Comparison | Agreement | Cohen’s |
|---|---|---|
| Human / Qwen3.5 | 0.8300 | 0.6561 |
| Human / GLM-5.1 | 0.8800 | 0.7580 |
| Human / GPT-5.1 | 0.9100 | 0.8190 |
| System | Context Precision | Context Recall | Answer Accuracy | Faithfulness | Response Relevancy | Step Coverage |
|---|---|---|---|---|---|---|
| LightRAG | 0.695 [0.675, 0.715] | 0.635 [0.619, 0.652] | 0.405 [0.389, 0.420] | 0.732 [0.720, 0.744] | 0.727 [0.715, 0.739] | 0.603 [0.585, 0.621] |
| HippoRAG | 0.383 [0.362, 0.405] | 0.435 [0.419, 0.452] | 0.339 [0.322, 0.356] | 0.652 [0.637, 0.666] | 0.592 [0.576, 0.609] | 0.395 [0.377, 0.413] |
| Microsoft GraphRAG | 0.602 [0.580, 0.623] | 0.584 [0.567, 0.601] | 0.311 [0.298, 0.324] | 0.680 [0.669, 0.691] | 0.675 [0.662, 0.689] | 0.540 [0.522, 0.558] |
| VisRAG | – | – | 0.346 [0.330, 0.362] | – | 0.769 [0.759, 0.779] | 0.434 [0.416, 0.451] |
| RAG-Anything | 0.809 [0.792, 0.826] | 0.722 [0.707, 0.736] | 0.530 [0.515, 0.545] | 0.835 [0.826, 0.843] | 0.781 [0.772, 0.790] | 0.701 [0.685, 0.717] |
| MegaRAG | 0.713 [0.693, 0.732] | 0.663 [0.647, 0.679] | 0.535 [0.521, 0.549] | 0.993 [0.991, 0.995] | 0.817 [0.811, 0.823] | 0.733 [0.718, 0.748] |
| System | Single-hop | Bridge-chain | Synthesis | |||
|---|---|---|---|---|---|---|
| Context Precision | Context Recall | Context Precision | Context Recall | Context Precision | Context Recall | |
| LightRAG | 0.765 | 0.724 | 0.749 | 0.700 | 0.589 | 0.505 |
| HippoRAG | 0.702 | 0.711 | 0.449 | 0.500 | 0.154 | 0.226 |
| Microsoft GraphRAG | 0.729 | 0.707 | 0.618 | 0.617 | 0.524 | 0.484 |
| RAG-Anything | 0.848 | 0.814 | 0.881 | 0.793 | 0.691 | 0.580 |
| MegaRAG | 0.911 | 0.906 | 0.740 | 0.768 | 0.589 | 0.410 |
| Bridge Path | Number | LightRAG | HippoRAG | Microsoft GraphRAG | VisRAG | RAG-Anything | MegaRAG |
|---|---|---|---|---|---|---|---|
| text text | 337 | 57.6 | 49.5 | 40.1 | 39.0 | 63.3 | 55.8 |
| text figure | 204 | 30.9 | 17.5 | 17.9 | 27.5 | 42.3 | 51.5 |
| text table | 119 | 37.2 | 21.9 | 25.4 | 16.0 | 27.1 | 39.5 |
| figure text | 62 | 49.6 | 39.9 | 30.2 | 38.3 | 71.0 | 60.1 |
| figure figure | 99 | 25.5 | 10.1 | 16.4 | 32.9 | 58.3 | 63.4 |
| figure table | 54 | 12.0 | 2.8 | 8.8 | 9.3 | 51.8 | 51.8 |
| Bridge Path | Number | LightRAG | HippoRAG | Microsoft GraphRAG | VisRAG | RAG-Anything | MegaRAG |
|---|---|---|---|---|---|---|---|
| text text | 337 | 84.6 | 65.4 | 75.2 | 59.8 | 87.8 | 85.2 |
| text figure | 204 | 62.5 | 40.0 | 53.4 | 50.5 | 68.4 | 80.6 |
| text table | 119 | 62.6 | 44.5 | 58.4 | 40.8 | 64.7 | 71.4 |
| figure text | 62 | 71.0 | 54.8 | 64.5 | 58.1 | 91.1 | 88.7 |
| figure figure | 99 | 51.0 | 25.2 | 40.4 | 60.1 | 76.8 | 87.9 |
| figure table | 54 | 35.2 | 10.2 | 35.2 | 29.6 | 68.5 | 80.6 |
| Pattern | Number | LightRAG | HippoRAG | Microsoft GraphRAG | VisRAG | RAG-Anything | MegaRAG |
|---|---|---|---|---|---|---|---|
| Panorama | 354 | 37.6 | 29.1 | 36.7 | 29.7 | 48.5 | 49.6 |
| Comparison | 130 | 28.7 | 19.8 | 27.5 | 18.3 | 39.2 | 34.8 |
| Contradiction | 88 | 34.7 | 34.7 | 29.8 | 28.1 | 45.2 | 40.3 |
| Trend | 73 | 44.5 | 38.7 | 41.4 | 32.2 | 50.7 | 48.6 |
| Causal | 63 | 37.7 | 43.7 | 41.3 | 34.9 | 54.0 | 54.0 |
| Topology | Total | Breakdown | Count |
| Single-hop | 72 | Text | 35 |
| Figure | 37 | ||
| Bridge-chain | 236 | Text text | 82 |
| Text figure | 68 | ||
| Other paths | 86 | ||
| Synthesis | 165 | 3 evidence points | 92 |
| System | Context Precision | Context Recall | Answer Accuracy | Faithfulness | Response Relevancy | Step Coverage |
|---|---|---|---|---|---|---|
| LightRAG | 0.7324 | 0.6518 | 0.5356 | 0.7871 | 0.6942 | 0.5716 |
| VisRAG | – | – | 0.3800 | – | 0.7133 | 0.5021 |
| RAG-Anything | 0.8473 | 0.7259 | 0.5842 | 0.8847 | 0.8159 | 0.7632 |