Exemplar2VQA: A Scalable Exemplar-Driven Visual Question Answering Generation Framework via Multi-Agent Coding
Organizations: East China Normal University · Shanghai Jiao Tong University · Fudan University · Tencent
Abstract
Advancing spatial intelligence in Multimodal Large Language Models (MLLMs) is bottlenecked by the scarcity of complex, scalable 3D question-answer (QA) data. While manual annotation is labor-intensive, directly utilizing LLMs to synthesize these QA pairs often fails due to their inherent deficiencies in spatial and geometric computation. We introduce Exemplar2VQA, a scalable exemplar-driven visual question answering generation framework that rapidly synthesizes large-scale spatial QA pairs in simulated environments via multi-agent coding. By equipping collaborative agents with a meticulously designed library of geometric utilities, Exemplar2VQA bypasses LLMs' spatial reasoning flaws through deterministic code execution. Crucially, the framework exhibits remarkable versatility: taking diverse static object-centric spatial query templates as exemplars, it seamlessly and autonomously scales them into massive, high-fidelity synthetic datasets. Fine-tuning Qwen2.5-VL (3B/7B) exclusively on Exemplar2VQA-generated synthetic indoor data yields significant performance improvements across various diverse benchmarks. Furthermore, its effectiveness is not limited to in-domain indoor datasets but also robustly extends to outdoor and mixed-scene benchmarks. These results establish Exemplar2VQA as a scalable and powerful paradigm for bridging the sim-to-real gap in Embodied AI. Our code is at https://github.com/yingjiayu12/Exemplar2VQA
Figures & tables
| Method | Overall | Cam.-Cam. | Cam.-Obj. | Cam.-Reg. | Obj.-Obj. | Obj.-Reg. | Reg.-Reg. | Meas. | Cam. | MSR. | Appr.* | Obj.* |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Baseline | ||||||||||||
| InternVL3-38B [ 42 ] | 26.3 | 21.5 | 23.3 | 25.3 | 20.2 | 35.3 | 33.3 | 39.1 | 16.2 | 25.8 | 21.2 | 31.6 |
| InternVL2.5-38B [ 43 ] | 27.9 | 18.3 | 22.1 | 34.9 | 22.3 | 38.8 | 35.8 | 37.5 | 14.9 | 25.3 | 25.8 | 38.2 |
| Qwen2.5-VL-32B [ 17 ] | 27.7 | 24.7 | 22.1 | 31.3 | 26.6 | 32.9 | 29.6 | 31.2 | 18.9 | 27.8 | 24.2 | 35.5 |
| InternVL3-8B [ 42 ] | 25.7 | 25.8 | 25.6 | 28.9 | 31.9 | 35.3 | 37.0 | 23.4 | 16.2 | 14.6 | 24.2 | 32.9 |
| DeepSeek-VL2 [ 44 ] | 27.1 | 23.7 | 36.0 | 22.9 | 31.9 | 30.6 | 22.2 | 28.1 | 28.4 | 28.3 | 15.2 | 26.3 |
| Method | Overall | Entity | Scene | Size | Obj.-Obj. | Obj.-Scene | Entity | Function | Spatial |
|---|---|---|---|---|---|---|---|---|---|
| Quant. | Quant. | Assess. | Relation | Relation | Presence | Reasoning | Planning | ||
| Qwen2.5-VL-7B (Base) [ 17 ] | 33.3 | 32.7 | 36.9 | 36.9 | 35.3 | 32.3 | 27.6 | 34.2 | 27.5 |
| Qwen2.5-VL-7B (MMSI-Bench) [ 27 ] | 39.1 | 39.4 | 25.6 | 48.7 | 39.4 | 31.6 | 44.5 | 43.4 | 35.0 |
| Qwen2.5-VL-7B (ViewSpatial) [ 24 ] | 40.6 | 36.1 | 14.3 | 58.4 | 47.7 | 35.6 | 41.1 | 51.5 | 42.5 |
| Exemplar2VQA (Ours) | 42.2 + | 38.4 + | 24.2 | 55.1 + | 45.5 + | 35.4 + | 45.9 + | 50.3 + | 32.5 + |
| Agent Configuration (Roles) | QA Generation (%) | Camera Trajectory (%) |
|---|---|---|
| Agent Topology | ||
| 1 Agent (All roles merged) | 34.0 | 60.0 |
| 2 Agents (Architect, Coder+Reviewer+Refiner) | 72.0 | 74.0 |
| 3 Agents (Architect, Coder, Reviewer+Refiner) | 85.0 | 80.0 |
| 4 Agents (Exemplar2VQA Full: Arch., Coder, Rev., Ref.) | 92.0 | 84.0 |
| Generator Capacity (4 Agents) | ||
| Method | Rel. Dist | Rel. Dir | Route Plan | Appr. Order | Avg. |
|---|---|---|---|---|---|
| Qwen2.5-VL-3B (Base) | 34.7 | 42.6 | 28.9 | 35.0 | 35.3 |
| Qwen2.5-VL-3B (LLM annotation) | 31.7 | 40.4 | 34.5 | 6.3 | 28.2 |
| Exemplar2VQA (Ours) | 43.8 | 52.1 | 35.1 | 40.5 | 42.9 |
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | Avg. | Object Count | Relative Distance | Relative Direction |
|---|---|---|---|---|
| Baseline | ||||
| Qwen2.5-VL-72B [ 17 ] | 33.5 | 49.8 | 32.5 | 18.1 |
| LLaVA-v1.5-13B [ 55 ] | 31.6 | 55.3 | 22.6 | 16.9 |
| DeepSeek-VL2 [ 44 ] | 25.5 | 57.9 | 5.5 | 13.1 |
| Qwen2.5-VL-7B (Base) [ 17 ] | 28.3 | 46.3 | 31.1 | 7.4 |
| Exemplar2VQA (Ours) | 36.5 + | 59.0 + | 41.1 + | 9.3 + |
| Method | 1_obj | 2_obj | Overall |
|---|---|---|---|
| Qwen2.5-VL-7B (Base) [ 17 ] | 83.86 | 62.90 | 69.71 |
| Exemplar2VQA (Ours) | 84.46 + | 63.85 + | 70.50 + |
| Method | Avg. | Rel. Dist. | Rel. Dir. | Route Plan* | Appr. Order |
|---|---|---|---|---|---|
| GPT-4o [ 46 ] | 34.6 | 37.0 | 41.3 | 31.5 | 28.5 |
| Gemini-1.5 Pro [ 47 ] | 42.1 | 51.3 | 46.3 | 36.0 | 34.6 |
| InternVL2-2B [ 50 ] | 28.2 | 32.1 | 44.1 | 30.4 | 6.3 |
| InternVL2-8B [ 50 ] | 36.7 | 38.0 | 33.4 | 28.9 | 46.4 |
| Exemplar2VQA (Ours, InternVL2-2B) | 35.6 + | 33.6 + | 47.6 + | 36.0 + | 25.2 + |
| Exemplar2VQA (Ours, InternVL2-8B) | 46.1 + | 50.4 + | 47.2 + | 36.1 + | 50.6 + |
| Method | Overall | Cam.-Cam. | Cam.-Obj. | Cam.-Reg. | Obj.-Obj. | Obj.-Reg. | Reg.-Reg. | Meas. | Cam. | MSR. | Appr.* | Obj.* |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Exemplar2VQA | 29.33 0.99 | 31.54 2.24 | 33.72 1.16 | 37.35 4.34 | 26.96 2.66 | 30.98 2.96 | 30.04 1.42 | 30.73 4.78 | 20.27 2.34 | 28.96 1.05 | 24.24 0.00 | 25.00 1.32 |
| Method | Obj. Count | Abs. Dist | Obj. Size | Room Size | Avg. |
|---|---|---|---|---|---|
| GPT-4o [ 46 ] | 46.2 | 5.3 | 43.8 | 38.2 | 33.4 |
| Qwen2.5-VL-3B (Base) [ 17 ] | 21.9 | 18.4 | 22.4 | 27.2 | 22.5 |
| Exemplar2VQA (Ours) | 33.6 + | 29.7 + | 58.8 + | 27.4 + | 37.4 + |