Advancing spatial intelligence in Multimodal Large Language Models (MLLMs) is bottlenecked by the scarcity of complex, scalable 3D question-answer (QA) data. While manual annotation is labor-intensive, directly utilizing LLMs to synthesize these QA pairs often fails due to their inherent deficiencies in spatial and geometric computation. We introduce Exemplar2VQA, a scalable exemplar-driven visual question answering generation framework that rapidly synthesizes large-scale spatial QA pairs in simulated environments via multi-agent coding. By equipping collaborative agents with a meticulously designed library of geometric utilities, Exemplar2VQA bypasses LLMs' spatial reasoning flaws through deterministic code execution. Crucially, the framework exhibits remarkable versatility: taking diverse static object-centric spatial query templates as exemplars, it seamlessly and autonomously scales them into massive, high-fidelity synthetic datasets. Fine-tuning Qwen2.5-VL (3B/7B) exclusively on Exemplar2VQA-generated synthetic indoor data yields significant performance improvements across various diverse benchmarks. Furthermore, its effectiveness is not limited to in-domain indoor datasets but also robustly extends to outdoor and mixed-scene benchmarks. These results establish Exemplar2VQA as a scalable and powerful paradigm for bridging the sim-to-real gap in Embodied AI. Our code is at https://github.com/yingjiayu12/Exemplar2VQA
Figures & tables
Figure 1: The camera generation track of the Exemplar2VQA. The Camera Generation Track: A four-agent pipeline (Architect, Coder, Reviewer, Refiner) interprets natural language instructions to autonomously plan trajectories and capture specific multi-view observations within the simulator.
Figure 2: The QA generation track of the Exemplar2VQA. The QA Generation Track: Utilizing the captured scenes and scene metadata, an agent group adapts QA exemplars into scaled QA pairs.
Method
Overall
Cam.-Cam.
Cam.-Obj.
Cam.-Reg.
Obj.-Obj.
Obj.-Reg.
Reg.-Reg.
Meas.
Cam.
MSR.
Appr.*
Obj.*
Baseline
InternVL3-38B [ 42 ]
26.3
21.5
23.3
25.3
20.2
35.3
33.3
39.1
16.2
25.8
21.2
31.6
InternVL2.5-38B [ 43 ]
27.9
18.3
22.1
34.9
22.3
38.8
35.8
37.5
14.9
25.3
25.8
38.2
Qwen2.5-VL-32B [ 17 ]
27.7
24.7
22.1
31.3
26.6
32.9
29.6
31.2
18.9
27.8
24.2
35.5
InternVL3-8B [ 42 ]
25.7
25.8
25.6
28.9
31.9
35.3
37.0
23.4
16.2
14.6
24.2
32.9
DeepSeek-VL2 [ 44 ]
27.1
23.7
36.0
22.9
31.9
30.6
22.2
28.1
28.4
28.3
15.2
26.3
Table 1: Performance comparison on the MMSI-Bench dataset . The best results are shown in bold , and the second-best are underlined . Metrics marked with an asterisk (*) are evaluated in a strictly zero-shot setting. A superscript plus sign ( + ) indicates a performance improvement over the base Qwen2.5-VL-7B model.
Method
Overall
Entity
Scene
Size
Obj.-Obj.
Obj.-Scene
Entity
Function
Spatial
Quant.
Quant.
Assess.
Relation
Relation
Presence
Reasoning
Planning
Qwen2.5-VL-7B (Base) [ 17 ]
33.3
32.7
36.9
36.9
35.3
32.3
27.6
34.2
27.5
Qwen2.5-VL-7B (MMSI-Bench) [ 27 ]
39.1
39.4
25.6
48.7
39.4
31.6
44.5
43.4
35.0
Qwen2.5-VL-7B (ViewSpatial) [ 24 ]
40.6
36.1
14.3
58.4
47.7
35.6
41.1
51.5
42.5
Exemplar2VQA (Ours)
42.2 +
38.4 +
24.2
55.1 +
45.5 +
35.4 +
45.9 +
50.3 +
32.5 +
Table 3: Zero-shot performance comparison on the SpaCE-10 benchmark. We evaluate the OOD generalization of our Exemplar2VQA-finetuned model against its base model and baselines fine-tuned on other spatial datasets. The best results are shown in bold . A superscript plus sign ( + ) indicates a performance improvement over the base Qwen2.5-VL-7B model.
Table 5: Ablation Study on Multi-Agent Configurations and Generator Capacity. We evaluate the generation success rate (%) given 100 seed examples under various agent topologies and generator backbones. Merged roles (indicated by ’+’) share the same context window, whereas separated roles (separated by ’,’) operate as distinct agents.
Method
Rel. Dist
Rel. Dir
Route Plan
Appr. Order
Avg.
Qwen2.5-VL-3B (Base)
34.7
42.6
28.9
35.0
35.3
Qwen2.5-VL-3B (LLM annotation)
31.7
40.4
34.5
6.3
28.2
Exemplar2VQA (Ours)
43.8
52.1
35.1
40.5
42.9
Table 6: Comparison with Direct LLM Annotation. We use Gemini3.5-Flash to directly annotate 10K synthetic training examples for the multiple-choice subset of VSI-Bench and fine-tune the same base model. Direct LLM annotation underperforms our code-based generation and even degrades the base model on average.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A1: Visual examples of diversified spatial QA pairs and corresponding 3D scenes programmatically generated by the Exemplar2VQA framework. This visualization demonstrates the framework’s versatility in adapting and scaling diverse query templates sourced from the OSR-Bench for Object Counting , Relative Distance , and Relative Direction across different generated environments to create massive and high-fidelity spatial reasoning datasets.
Method
Avg.
Object Count
Relative Distance
Relative Direction
Baseline
Qwen2.5-VL-72B [ 17 ]
33.5
49.8
32.5
18.1
LLaVA-v1.5-13B [ 55 ]
31.6
55.3
22.6
16.9
DeepSeek-VL2 [ 44 ]
25.5
57.9
5.5
13.1
Qwen2.5-VL-7B (Base) [ 17 ]
28.3
46.3
31.1
7.4
Exemplar2VQA (Ours)
36.5 +
59.0 +
41.1 +
9.3 +
Appendix
Table A1: Performance comparison on the OSR-Bench dataset . The best results are shown in bold , and the second-best are underlined . A superscript plus sign ( + ) indicates a performance improvement over the base Qwen2.5-VL-7B model.
Method
1_obj
2_obj
Overall
Qwen2.5-VL-7B (Base) [ 17 ]
83.86
62.90
69.71
Exemplar2VQA (Ours)
84.46 +
63.85 +
70.50 +
Appendix
Table A2: Zero-shot performance comparison on the Spatial-Obj dataset . Quantitative results demonstrating the spatial reasoning capabilities of our Exemplar2VQA-finetuned model on one-object and two-object scenarios. The best results are shown in bold . A superscript plus sign ( + ) indicates a performance improvement over the base Qwen2.5-VL-7B model.
Method
Avg.
Rel. Dist.
Rel. Dir.
Route Plan*
Appr. Order
GPT-4o [ 46 ]
34.6
37.0
41.3
31.5
28.5
Gemini-1.5 Pro [ 47 ]
42.1
51.3
46.3
36.0
34.6
InternVL2-2B [ 50 ]
28.2
32.1
44.1
30.4
6.3
InternVL2-8B [ 50 ]
36.7
38.0
33.4
28.9
46.4
Exemplar2VQA (Ours, InternVL2-2B)
35.6 +
33.6 +
47.6 +
36.0 +
25.2 +
Exemplar2VQA (Ours, InternVL2-8B)
46.1 +
50.4 +
47.2 +
36.1 +
50.6 +
Appendix
Table A3: Generalization across backbone architectures. We fine-tune InternVL2-2B and InternVL2-8B on the same synthetic dataset generated by Exemplar2VQA and evaluate on VSI-Bench. The best results are shown in bold , and the second-best are underlined . A superscript plus sign ( + ) indicates a performance improvement over the corresponding base model.
Figure A2: Customizable QA category generation versus static benchmark distributions. (Top) Existing spatial reasoning datasets (e.g., ViewSpatial-Bench, MMSI-Bench) exhibit fixed and inherently imbalanced QA category distributions. (Bottom) In contrast, our proposed pipeline enables arbitrary, programmatic control over the generation ratios. Driven by a theoretically infinite generation capacity—bounded only by the diversity of available 3D scenes—our framework allows users to dynamically configure the distribution of specific spatial queries. This paradigm effectively overcomes the static distribution biases prevalent in traditional human-annotated benchmarks.
Method
Overall
Cam.-Cam.
Cam.-Obj.
Cam.-Reg.
Obj.-Obj.
Obj.-Reg.
Reg.-Reg.
Meas.
Cam.
MSR.
Appr.*
Obj.*
Exemplar2VQA
29.33 ± 0.99
31.54 ± 2.24
33.72 ± 1.16
37.35 ± 4.34
26.96 ± 2.66
30.98 ± 2.96
30.04 ± 1.42
30.73 ± 4.78
20.27 ± 2.34
28.96 ± 1.05
24.24 ± 0.00
25.00 ± 1.32
Appendix
Table A4: Stability of MMSI-Bench results across random seeds. We rerun the MMSI-Bench fine-tuning experiment with three different random seeds using the same training data each time, and report the mean with the standard deviation as a superscript. Metrics marked with an asterisk (*) are evaluated in a strictly zero-shot setting.
Method
Obj. Count
Abs. Dist
Obj. Size
Room Size
Avg.
GPT-4o [ 46 ]
46.2
5.3
43.8
38.2
33.4
Qwen2.5-VL-3B (Base) [ 17 ]
21.9
18.4
22.4
27.2
22.5
Exemplar2VQA (Ours)
33.6 +
29.7 +
58.8 +
27.4 +
37.4 +
Appendix
Table A5: Performance comparison on the four numerical tasks of VSI-Bench . The best results are shown in bold , and the second-best are underlined . A superscript plus sign ( + ) indicates a performance improvement over the base Qwen2.5-VL-3B model.
Figure A3: Distribution of error types in the Exemplar2VQA pipeline.
Figure A4: Additional qualitative examples of diverse spatial reasoning QA pairs autonomously synthesized across simulated environments.
Figure A5: Additional qualitative examples of diverse spatial reasoning QA pairs autonomously synthesized across simulated environments.
Figure A6: Additional qualitative examples of diverse spatial reasoning QA pairs autonomously synthesized across simulated environments.
Figure A7: Additional qualitative examples of diverse spatial reasoning QA pairs autonomously synthesized across simulated environments.
Figure A8: Additional qualitative examples of diverse spatial reasoning QA pairs autonomously synthesized across simulated environments.
Figure A9: Additional qualitative examples of diverse spatial reasoning QA pairs autonomously synthesized across simulated environments.
Spatial question answering is the dominant paradigm for evaluating spatial intelligence in Vision-Language Models (VLMs), but it leaves a complementary axis of spatial competence under-evaluated: holistic 3D layout inference, which predicts every visible object's pose and extent from a single image in a structured form. To this end, we introduce IDEAL-Bench, an evaluation suite that requires VLMs to predict structured 3D layouts on photorealistic indoor scenes across 10 room types, scored along five numerical dimensions and a perceptual render-and-compare protocol. By operating on semantically realistic scenes with full asset substitution under controlled lighting and viewpoint, IDEAL-Bench moves beyond CLEVR-style simple geometric primitives so that any image-space discrepancy reflects spatial reasoning alone. The benchmark is built on IDEAL-Scenes, a procedurally generated dataset of 1,000 re-renderable Blender environments with ground-truth layouts. Evaluating 15 prominent VLMs reveals three findings: the task remains substantially unsolved, with the strongest model reaching only 62.1/100 overall; all models exhibit a sharp asymmetry between object recognition and geometric regression, indicating that current VLMs are trained to describe scenes rather than to measure them; model rankings partially diverge from those on QA-based and primitive-reconstruction benchmarks: top-tier consensus holds, but mid-tier rankings shift substantially. Collectively, these findings establish IDEAL-Bench as a diagnostic suite, targeting the geometric and structural competencies that QA-based evaluation cannot surface, and paving the way towards more rigorous evaluation of spatial intelligence in next-generation VLMs. Together, these findings position IDEAL-Bench as a principled diagnostic for whether future VLMs achieve genuine spatial understanding rather than linguistic approximations of it.
We present M3-VQA, a novel knowledge-based Visual Question Answering (VQA) benchmark, to enhance the evaluation of multimodal large language models (MLLMs) in fine-grained multimodal entity understanding and complex multi-hop reasoning. Unlike existing VQA datasets that focus on coarse-grained categories and simple reasoning over single entities, M3-VQA introduces diverse multi-entity questions involving multiple distinct entities from both visual and textual sources. It requires models to perform both sequential and parallel multi-hop reasoning across multiple documents, supported by traceable, detailed evidence and a curated multimodal knowledge base. We evaluate 16 leading MLLMs under three settings: without external knowledge, with gold evidence, and with retrieval-augmented input. The poor results reveal significant challenges for MLLMs in knowledge acquisition and reasoning. Models perform poorly without external information but improve markedly when provided with precise evidence. Furthermore, reasoning-aware agentic retrieval surpasses heuristic methods, highlighting the importance of structured reasoning for complex multimodal understanding. M3-VQA presents a more challenging evaluation for advancing the multimodal reasoning capabilities of MLLMs. Our code and dataset are available at https://github.com/CASIA-IVA-Lab/M3VQA.
Jiatong Ma, Longteng Guo, Yuchen Liu +4
Institute of Automation, Chinese Academy of Sciences · School of Artificial Intelligence, University of Chinese Academy of Sciences
Recent advancements in Large Vision-Language Models (VLMs) have demonstrated exceptional semantic understanding, yet these models consistently struggle with spatial reasoning, often failing at fundamental geometric tasks such as depth ordering and precise coordinate grounding. Recent efforts introduce spatial supervision from scene-centric datasets (e.g., multi-view scans or indoor video), but are constrained by the limited number of underlying scenes. As a result, the scale and diversity of such data remain significantly smaller than those of web-scale 2D image collections. To address this limitation, we propose SpatialForge, a scalable data synthesis pipeline that transforms in-the-wild 2D images into spatial reasoning supervision. Our approach decomposes spatial reasoning into perception and relation, and constructs structured supervision signals covering depth, layout, and viewpoint-dependent reasoning, with automatic verification to ensure data quality. Based on this pipeline, we build SpatialForge-10M, a large-scale dataset containing 10 million spatial QA pairs. Extensive experiments across multiple spatial reasoning benchmarks demonstrate that training on SpatialForge-10M significantly improves the spatial reasoning ability of standard VLMs, highlighting the effectiveness of scaling 2D data for 3D-aware spatial reasoning.