ARK: A Dual-Axis Multimodal Retrieval Benchmark along Reasoning and Knowledge
Organizations: College of Computer Science, Sichuan University, Chengdu, China · National Key Laboratory of Fundamental Algorithms and Models for Engineering Simulation, Sichuan University, China
Abstract
Existing multimodal retrieval benchmarks largely emphasize semantic matching on daily-life images and offer limited diagnostics of professional knowledge and complex reasoning. To address this gap, we introduce ARK, a benchmark designed to analyze multimodal retrieval from two complementary perspectives: (i) knowledge domains (five domains with 17 subtypes), which characterize the content and expertise retrieval relies on, and (ii) reasoning skills (six categories), which characterize the type of inference over multimodal evidence required to identify the correct candidate. Specifically, ARK evaluates retrieval with both unimodal and multimodal queries and candidates, covering 16 heterogeneous visual data types. To avoid shortcut matching during evaluation, most queries are paired with targeted hard negatives that require multi-step reasoning. We evaluate 25 representative text-based and multimodal retrievers and observe a pronounced gap between knowledge- and reasoning-intensive retrieval, with fine-grained visual and spatial reasoning as persistent bottlenecks. We further show that enhancements such as re-ranking, rewriting, and agentic retrieval yield consistent gains, but substantial headroom remains.
Figures & tables
| Knowledge Domain | Subtype | Query | Gallery |
| Natural Science | Biology | 83 | 1834 |
| Physics | 80 | 712 | |
| Chemistry | 105 | 1127 | |
| Formal Science | Mathematics | 152 | 889 |
| Code-Drawing | 100 | 551 | |
| Engineering & Technology | Computer-Science | 80 | 5231 |
| Model | Visual Cognition | Natural Science | Formal Science | Humanities & Social Science | Engineering & Technology | Avg. | ||||||||||||
| Daily. | Fine. | SCog. | S3D. | Bio. | Phy. | Che. | Mat. | Cod. | Eco. | Comic. | Art | Cal. | Met. | Compu. | Mecha. | Ele. | ||
| Text Embedding Models | ||||||||||||||||||
| BGE-M3-0.6B | 37.29 | 2.17 | 0.00 | 0.00 | 32.53 | 15.00 | 0.95 | 3.29 | 2.00 | 36.25 | 1.15 | 2.00 | 6.25 | 7.50 | 25.00 | 9.76 | 0.99 | 10.71 |
| Diver-Embed-1.7B | 35.59 | 2.17 | 0.00 | 0.00 | 26.51 | 25.00 | 2.86 | 11.84 | 2.00 | 23.75 | 1.15 | 4.00 | 8.75 | 6.25 | 38.75 | 4.88 | 0.00 | 11.38 |
| Diver-Embed-4B | 46.19 | 1.09 | 0.00 | 0.76 | 39.76 | 22.50 | 2.86 | 14.47 | 2.00 | 55.00 | 2.30 | 5.00 | 6.25 | 10.00 | 46.25 | 10.98 | 0.00 | 15.61 |
| Qwen3-Embed-4B | 42.37 | 2.17 | 3.70 | 0.76 | 27.71 | 21.25 | 1.90 | 10.53 | 3.00 | 36.25 | 1.15 | 4.00 | 5.00 | 3.75 | 33.75 | 10.98 | 2.97 | 12.43 |
| Model | Cognitive Demand | Reasoning Skill | |||||||||
| Perception- Only | Knowledge- Intensive | Reasoning- Intensive | Both | Knowledge Reasoning | Conceptual Abstraction | Fine-Grained Visual Reasoning | Logical Reasoning | Spatial Reasoning | Symbolic Reasoning | Average | |
| Text Embedding Models | |||||||||||
| BGE-M3-0.6B | 37.29 | 12.09 | 7.89 | 10.12 | 11.70 | 18.34 | 3.08 | 4.29 | 2.31 | 3.90 | 7.27 |
| Diver-Embed-1.7B | 36.02 | 14.86 | 9.88 | 13.20 | 14.84 | 21.83 | 2.31 | 5.52 | 2.69 | 7.79 | 9.16 |
| Diver-Embed-4B | 46.19 | 17.73 | 11.94 | 15.95 | 18.12 | 27.07 | 3.85 | 6.13 | 1.54 | 9.31 | 11.00 |
| Qwen3-Embed-4B | 41.53 | 12.09 | 7.73 | 9.90 | 11.98 | 17.90 | 1.54 | 4.29 | 2.69 | 5.19 | 7.26 |
Appendix figures & tables43 assets
Supplementary material from the paper’s appendix.
Appendix
| Benchmarks | # Queries | # Tasks | Knowledge- Intensive | Reasoning- Intensive | Multi- Domain | Diverse Reasoning Type | Heterogeneous Visual Type | Fine-grained Visual Reasoning |
| COCO [ 7 ] | 25K | 1 | ||||||
| Flickr30K [ 8 ] | 5, 000 | 1 | ||||||
| CIRR [ 14 ] | 4,148 | 1 | ||||||
| FashionIQ [ 13 ] | 60K | 1 | ||||||
| CIRCO [ 57 ] | 1,020 | 1 | ||||||
| WebQA [ 58 ] | 7,540 | 1 | ✓ |
| Meta-Task | Subtype | Visual Type | Reasoning Skill | Adapted From |
| Visual Cognition | Daily-Life | natural images | – | COCO [ 7 ] and Fashion200K [ 63 ] |
| Fine-Grained | natural images | fine-grained visual reasoning | DIV8K [ 64 ] | |
| Spatial-Cogmap | cognitive maps | spatial reasoning | MindCube [ 65 ] | |
| Spatial-3D | 3D images | spatial reasoning | Spatial-DISE-12K [ 66 ] | |
| Natural Science | Biology | natural images | – | INQUIRE [ 67 ] |
| Physics | physics phenomenon illustrations | knowledge reasoning | Manually collected |
| Modality | Model Name | Param. | Reasoning | Model Link |
| Text | BGE-M3 [ 78 ] | 600M | HF: BAAI/bge-m3 | |
| Qwen3-Embedding [ 53 ] | 4B | HF: Qwen/Qwen3-Embedding-4B | ||
| Qwen3-Embedding [ 53 ] | 8B | HF: Qwen/Qwen3-Embedding-8B | ||
| Diver-Embed [ 34 ] | 1.7B | ✓ | HF: AQ-MedAI/Diver-Retriever-1.7B | |
| Diver-Embed [ 34 ] | 4B | ✓ | HF: AQ-MedAI/Diver-Retriever-4B | |
| ReasonIR [ 35 ] | 8B | ✓ | HF: reasonir/ReasonIR-8B |
| Modality | Model Name | Param. | Model Link |
| Text | BGE-Reranker [ 78 ] | 600M | HF: BAAI/bge-reranker-v2-m3 |
| Multi-modal | Qwen3-VL-Reranker [ 32 ] | 8B | HF: Qwen/Qwen3-VL-Reranker-8B |
| Jina-Reranker [ 85 ] | 2B | HF: jinaai/jina-reranker-m0 |
| Modality | Model Name | Param. | Model Link |
| Text | A-RAG [ 86 ] | 0.6B | GPT-5-mini + Qwen3-Embedding-0.6B |
| A-RAG [ 86 ] | 8B | GPT-5-mini + Qwen3-Embedding-8B | |
| Agentic-R [ 87 ] | 7B | HF: liuwenhan/search-agent |
| Model | Knowledge perspective | Reasoning perspective | ||||||||||
| R@1 | R@5 | R@10 | N@5 | N@10 | N@20 | R@1 | R@5 | R@10 | N@5 | N@10 | N@20 | |
| Query Rewrite | ||||||||||||
| Seed1.6-embedding | 19.32 | 37.60 | 48.43 | 28.79 | 32.28 | 34.77 | 15.30 | 31.69 | 40.70 | 23.80 | 26.71 | 29.18 |
| + Rewrite | 23.42 | 43.85 | 54.49 | 34.04 | 37.52 | 39.87 | 17.49 | 35.09 | 44.12 | 26.59 | 29.53 | 31.96 |
| Reranker Models Based on Seed1.6-embedding | ||||||||||||
| BGE-Reranker | 12.23 | 23.95 | 32.47 | 18.29 | 21.05 | 24.11 | 8.19 | 18.35 | 26.01 | 13.40 | 15.88 | 18.51 |
| Model | Knowledge Reasoning | Conceptual Abstraction | Fine-Grained Visual Reasoning | Logical Reasoning | Spatial Reasoning | Symbolic Reasoning | Average |
| Query Rewrite | |||||||
| Seed1.6-embedding | 22.08 | 33.33 | 4.35 | 10.71 | 6.19 | 15.16 | 15.30 |
| + Rewrite | 24.96 | 31.00 | 10.00 | 18.40 | 5.00 | 15.58 | 17.49 |
| Reranker Models Based on Seed1.6-embedding | |||||||
| BGE-Reranker | 14.55 | 21.40 | 3.85 | 3.07 | 1.54 | 4.76 | 8.19 |
| + Rewrite | 16.98 | 21.40 | 5.38 | 17.18 | 3.08 | 5.84 | 11.64 |
| Model | Visual Cognition | Natural Science | Formal Science | Humanities & Social Science | Engineering & Technology | Avg. | ||||||||||||
| Daily. | Fine. | SCog. | S3D. | Bio. | Phy. | Che. | Mat. | Cod. | Eco. | Comic. | Art | Cal. | Met. | Compu. | Mecha. | Ele. | ||
| Text Embedding Models | ||||||||||||||||||
| BGE-M3-0.6B | 57.63 | 6.52 | 0.00 | 1.53 | 63.86 | 30.00 | 8.57 | 10.53 | 4.00 | 56.25 | 3.45 | 9.00 | 7.50 | 31.25 | 31.25 | 10.98 | 0.99 | 19.61 |
| Diver-Embed-1.7B | 48.73 | 7.61 | 0.00 | 0.76 | 57.83 | 46.25 | 9.52 | 26.97 | 7.00 | 50.00 | 2.30 | 7.00 | 10.00 | 30.00 | 52.50 | 14.63 | 1.98 | 21.95 |
| Diver-Embed-4B | 61.02 | 6.52 | 0.00 | 0.76 | 68.67 | 40.00 | 15.24 | 31.58 | 15.00 | 75.00 | 5.75 | 11.00 | 7.50 | 38.75 | 60.00 | 15.85 | 4.95 | 26.92 |
| Qwen3-Embed-4B | 57.20 | 7.61 | 4.94 | 2.29 | 71.08 | 41.25 | 14.29 | 28.95 | 13.00 | 63.75 | 3.45 | 9.00 | 7.50 | 23.75 | 51.25 | 20.73 | 6.93 | 25.12 |
| Model | Visual Cognition | Natural Science | Formal Science | Humanities & Social Science | Engineering & Technology | Avg. | ||||||||||||
| Daily. | Fine. | SCog. | S3D. | Bio. | Phy. | Che. | Mat. | Cod. | Eco. | Comic. | Art | Cal. | Met. | Compu. | Mecha. | Ele. | ||
| Text Embedding Models | ||||||||||||||||||
| BGE-M3-0.6B | 63.98 | 11.96 | 1.23 | 1.53 | 68.67 | 37.50 | 17.14 | 15.13 | 11.00 | 67.50 | 6.90 | 13.00 | 7.50 | 45.00 | 43.75 | 14.63 | 4.95 | 25.37 |
| Diver-Embed-1.7B | 56.78 | 9.78 | 0.00 | 1.53 | 72.29 | 50.00 | 19.05 | 31.58 | 14.00 | 57.50 | 2.30 | 12.00 | 12.50 | 50.00 | 58.75 | 19.51 | 7.92 | 27.97 |
| Diver-Embed-4B | 68.64 | 13.04 | 0.00 | 0.76 | 78.31 | 48.75 | 23.81 | 38.82 | 23.00 | 82.50 | 6.90 | 17.00 | 10.00 | 53.75 | 67.50 | 19.51 | 6.93 | 32.90 |
| Qwen3-Embed-4B | 65.68 | 15.22 | 16.05 | 3.82 | 79.52 | 48.75 | 19.05 | 37.50 | 19.00 | 76.25 | 5.75 | 13.00 | 8.75 | 50.00 | 58.75 | 23.17 | 14.85 | 32.65 |
| Model | Visual Cognition | Natural Science | Formal Science | Humanities & Social Science | Engineering & Technology | Avg. | ||||||||||||
| Daily. | Fine. | SCog. | S3D. | Bio. | Phy. | Che. | Mat. | Cod. | Eco. | Comic. | Art | Cal. | Met. | Compu. | Mecha. | Ele. | ||
| Text Embedding Models | ||||||||||||||||||
| BGE-M3-0.6B | 48.34 | 4.27 | 0.00 | 0.76 | 48.88 | 23.44 | 5.13 | 7.10 | 2.89 | 47.21 | 2.22 | 5.62 | 6.73 | 19.37 | 28.37 | 10.23 | 0.99 | 15.39 |
| Diver-Embed-1.7B | 42.55 | 4.62 | 0.00 | 0.30 | 41.85 | 36.70 | 6.37 | 19.97 | 4.27 | 37.19 | 1.64 | 5.43 | 9.54 | 17.94 | 45.96 | 10.33 | 0.88 | 16.80 |
| Diver-Embed-4B | 54.27 | 3.70 | 0.00 | 0.76 | 55.07 | 32.82 | 8.98 | 23.41 | 8.35 | 64.75 | 4.09 | 8.25 | 7.04 | 24.23 | 53.74 | 13.81 | 2.48 | 21.51 |
| Qwen3-Embed-4B | 50.07 | 5.20 | 4.32 | 1.63 | 50.08 | 32.21 | 8.31 | 20.23 | 8.10 | 51.32 | 2.60 | 6.31 | 6.16 | 13.77 | 42.66 | 16.43 | 4.97 | 19.08 |
| Model | Visual Cognition | Natural Science | Formal Science | Humanities & Social Science | Engineering & Technology | Avg. | ||||||||||||
| Daily. | Fine. | SCog. | S3D. | Bio. | Phy. | Che. | Mat. | Cod. | Eco. | Comic. | Art | Cal. | Met. | Compu. | Mecha. | Ele. | ||
| Text Embedding Models | ||||||||||||||||||
| BGE-M3-0.6B | 50.39 | 6.01 | 0.39 | 0.76 | 50.49 | 25.81 | 7.93 | 8.63 | 5.14 | 50.90 | 3.36 | 6.82 | 6.73 | 23.71 | 32.54 | 11.38 | 2.31 | 17.25 |
| Diver-Embed-1.7B | 45.15 | 5.25 | 0.00 | 0.55 | 46.45 | 37.91 | 9.47 | 21.46 | 6.48 | 39.67 | 1.64 | 7.10 | 10.35 | 24.25 | 48.01 | 11.89 | 2.85 | 18.73 |
| Diver-Embed-4B | 56.74 | 5.78 | 0.00 | 0.76 | 58.24 | 35.64 | 11.72 | 25.83 | 10.93 | 67.15 | 4.50 | 10.15 | 7.90 | 29.10 | 56.26 | 14.96 | 3.13 | 23.46 |
| Qwen3-Embed-4B | 52.81 | 7.54 | 7.82 | 2.12 | 52.82 | 34.58 | 9.87 | 22.94 | 10.01 | 55.59 | 3.26 | 7.55 | 6.52 | 22.21 | 45.21 | 17.18 | 7.51 | 21.50 |
| Model | Visual Cognition | Natural Science | Formal Science | Humanities & Social Science | Engineering & Technology | Avg. | ||||||||||||
| Daily. | Fine. | SCog. | S3D. | Bio. | Phy. | Che. | Mat. | Cod. | Eco. | Comic. | Art | Cal. | Met. | Compu. | Mecha. | Ele. | ||
| Text Embedding Models | ||||||||||||||||||
| BGE-M3-0.6B | 52.68 | 7.61 | 1.92 | 1.16 | 52.64 | 27.62 | 10.59 | 11.11 | 6.14 | 54.08 | 4.25 | 7.76 | 6.73 | 29.55 | 33.71 | 13.54 | 2.54 | 19.04 |
| Diver-Embed-1.7B | 47.83 | 6.64 | 0.60 | 1.13 | 48.88 | 39.80 | 11.89 | 23.79 | 8.49 | 42.31 | 1.95 | 8.58 | 10.94 | 29.00 | 49.27 | 12.49 | 4.07 | 20.45 |
| Diver-Embed-4B | 58.72 | 7.38 | 0.32 | 0.95 | 60.07 | 37.29 | 12.92 | 28.04 | 12.67 | 68.75 | 4.50 | 11.65 | 10.77 | 33.72 | 56.93 | 15.92 | 4.15 | 24.99 |
| Qwen3-Embed-4B | 54.43 | 7.84 | 9.41 | 2.89 | 54.90 | 35.86 | 11.78 | 23.98 | 11.27 | 57.47 | 3.57 | 8.84 | 7.10 | 28.58 | 46.74 | 18.12 | 8.99 | 23.05 |
| Model | Cognitive Demand | Reasoning Skill | |||||||||
| Perception- Only | Knowledge- Intensive | Reasoning- Intensive | Both | Knowledge Reasoning | Conceptual Abstraction | Fine-Grained Visual Reasoning | Logical Reasoning | Spatial Reasoning | Symbolic Reasoning | Average | |
| Text Embedding Models | |||||||||||
| BGE-M3-0.6B | 57.63 | 24.44 | 19.07 | 22.00 | 23.25 | 31.88 | 12.31 | 14.11 | 11.92 | 14.07 | 17.92 |
| Diver-Embed-1.7B | 49.15 | 28.74 | 21.82 | 26.84 | 28.67 | 37.55 | 13.85 | 15.34 | 8.46 | 21.21 | 20.85 |
| Diver-Embed-4B | 61.02 | 32.68 | 24.73 | 30.91 | 33.38 | 43.67 | 13.85 | 17.79 | 8.46 | 24.68 | 23.64 |
| Qwen3-Embed-4B | 57.63 | 25.96 | 18.38 | 22.66 | 26.68 | 35.81 | 7.69 | 7.36 | 8.46 | 16.23 | 17.04 |
| Model | Cognitive Demand | Reasoning Skill | |||||||||
| Perception- Only | Knowledge- Intensive | Reasoning- Intensive | Both | Knowledge Reasoning | Conceptual Abstraction | Fine-Grained Visual Reasoning | Logical Reasoning | Spatial Reasoning | Symbolic Reasoning | Average | |
| Text Embedding Models | |||||||||||
| BGE-M3-0.6B | 63.98 | 29.81 | 23.74 | 27.17 | 27.67 | 37.12 | 17.69 | 17.18 | 14.62 | 20.56 | 22.47 |
| Diver-Embed-1.7B | 56.78 | 34.65 | 26.80 | 31.90 | 33.38 | 40.61 | 21.54 | 19.02 | 11.54 | 27.06 | 25.53 |
| Diver-Embed-4B | 69.92 | 39.03 | 30.70 | 37.29 | 39.94 | 49.34 | 23.08 | 23.31 | 11.54 | 30.52 | 29.62 |
| Qwen3-Embed-4B | 65.25 | 32.59 | 26.34 | 29.81 | 33.95 | 43.23 | 20.77 | 12.88 | 17.31 | 23.81 | 25.32 |
| Model | Cognitive Demand | Reasoning Skill | |||||||||
| Perception- Only | Knowledge- Intensive | Reasoning- Intensive | Both | Knowledge Reasoning | Conceptual Abstraction | Fine-Grained Visual Reasoning | Logical Reasoning | Spatial Reasoning | Symbolic Reasoning | Average | |
| Text Embedding Models | |||||||||||
| BGE-M3-0.6B | 48.34 | 18.58 | 13.65 | 16.35 | 17.73 | 25.26 | 7.34 | 9.35 | 7.26 | 9.27 | 12.70 |
| Diver-Embed-1.7B | 42.77 | 22.21 | 16.18 | 20.44 | 22.13 | 30.47 | 8.15 | 10.41 | 5.75 | 15.01 | 15.32 |
| Diver-Embed-4B | 54.21 | 25.68 | 18.65 | 23.84 | 26.17 | 36.12 | 8.97 | 12.15 | 5.07 | 17.20 | 17.61 |
| Qwen3-Embed-4B | 49.92 | 19.48 | 13.40 | 16.73 | 19.90 | 27.62 | 4.98 | 5.84 | 5.60 | 11.08 | 12.50 |
| Model | Cognitive Demand | Reasoning Skill | |||||||||
| Perception- Only | Knowledge- Intensive | Reasoning- Intensive | Both | Knowledge Reasoning | Conceptual Abstraction | Fine-Grained Visual Reasoning | Logical Reasoning | Spatial Reasoning | Symbolic Reasoning | Average | |
| Text Embedding Models | |||||||||||
| BGE-M3-0.6B | 50.39 | 20.33 | 15.18 | 18.04 | 19.19 | 26.99 | 9.04 | 10.37 | 8.17 | 11.41 | 14.20 |
| Diver-Embed-1.7B | 45.23 | 24.12 | 17.78 | 22.07 | 23.64 | 31.43 | 10.53 | 11.62 | 6.77 | 16.90 | 16.81 |
| Diver-Embed-4B | 57.07 | 27.73 | 20.60 | 25.90 | 28.29 | 37.94 | 12.02 | 13.92 | 6.09 | 19.12 | 19.56 |
| Qwen3-Embed-4B | 52.47 | 21.63 | 15.98 | 19.06 | 22.28 | 30.10 | 9.11 | 7.59 | 8.45 | 13.52 | 15.17 |
| Model | Cognitive Demand | Reasoning Skill | |||||||||
| Perception- Only | Knowledge- Intensive | Reasoning- Intensive | Both | Knowledge Reasoning | Conceptual Abstraction | Fine-Grained Visual Reasoning | Logical Reasoning | Spatial Reasoning | Symbolic Reasoning | Average | |
| Text Embedding Models | |||||||||||
| BGE-M3-0.6B | 52.68 | 22.31 | 17.11 | 19.98 | 21.26 | 28.72 | 11.54 | 12.22 | 9.52 | 13.67 | 16.16 |
| Diver-Embed-1.7B | 47.88 | 25.87 | 19.48 | 23.89 | 25.48 | 32.76 | 12.32 | 13.47 | 7.85 | 19.05 | 18.49 |
| Diver-Embed-4B | 58.53 | 29.15 | 22.03 | 27.36 | 29.81 | 38.41 | 14.16 | 14.87 | 7.09 | 20.95 | 20.88 |
| Qwen3-Embed-4B | 54.32 | 22.99 | 17.51 | 20.39 | 23.61 | 31.19 | 11.39 | 8.79 | 10.36 | 15.02 | 16.73 |
| Model | Knowledge domain avg. | Reasoning skill avg. | ||
| R@1 | Hits@1 | R@1 | Hits@1 | |
| Text Embedding Models | ||||
| BGE-M3-0.6B | 10.71 | - | 7.27 | - |
| Diver-Embed-1.7B | 11.38 | - | 9.16 | - |
| Diver-Embed-4B | 15.61 | - | 11.00 | - |
| Qwen3-Embed-4B | 12.43 | - | 7.26 | - |
| Text embedding model | Qwen | Kimi | GLM | Kimi Qwen | GLM Qwen |
| BGE-M3-0.6B | 10.71 | 8.97 | 10.08 | ||
| Diver-Embed-1.7B | 11.38 | 9.70 | 10.40 | ||
| Diver-Embed-4B | 15.61 | 12.97 | 14.77 | ||
| Qwen3-Embed-4B | 12.43 | 11.05 | 13.30 | ||
| Qwen3-Embed-8B | 14.90 | 11.95 | 13.67 | ||
| ReasonIR-8B | 12.72 | 11.09 | 12.37 |
| Text embedding model | Qwen | Kimi | GLM |
| BGE-M3-0.6B | 7.27 | 5.42 | 6.12 |
| Diver-Embed-1.7B | 9.16 | 5.29 | 5.64 |
| Diver-Embed-4B | 11.00 | 9.29 | 9.54 |
| Qwen3-Embed-4B | 7.26 | 7.18 | 8.35 |
| Qwen3-Embed-8B | 8.92 | 7.63 | 8.58 |
| ReasonIR-8B | 10.59 | 7.93 | 8.44 |