cs.CVSep 29, 2026

Exemplar2VQA: A Scalable Exemplar-Driven Visual Question Answering Generation Framework via Multi-Agent Coding

Authors: Jiayu Ying, Qijian Tian, Ruijie Xu, Xinnan Zhu, Daoguo Dong, Jiachen Xu, Xin Tan

Organizations: East China Normal University · Shanghai Jiao Tong University · Fudan University · Tencent

Abstract

Advancing spatial intelligence in Multimodal Large Language Models (MLLMs) is bottlenecked by the scarcity of complex, scalable 3D question-answer (QA) data. While manual annotation is labor-intensive, directly utilizing LLMs to synthesize these QA pairs often fails due to their inherent deficiencies in spatial and geometric computation. We introduce Exemplar2VQA, a scalable exemplar-driven visual question answering generation framework that rapidly synthesizes large-scale spatial QA pairs in simulated environments via multi-agent coding. By equipping collaborative agents with a meticulously designed library of geometric utilities, Exemplar2VQA bypasses LLMs' spatial reasoning flaws through deterministic code execution. Crucially, the framework exhibits remarkable versatility: taking diverse static object-centric spatial query templates as exemplars, it seamlessly and autonomously scales them into massive, high-fidelity synthetic datasets. Fine-tuning Qwen2.5-VL (3B/7B) exclusively on Exemplar2VQA-generated synthetic indoor data yields significant performance improvements across various diverse benchmarks. Furthermore, its effectiveness is not limited to in-domain indoor datasets but also robustly extends to outdoor and mixed-scene benchmarks. These results establish Exemplar2VQA as a scalable and powerful paradigm for bridging the sim-to-real gap in Embodied AI. Our code is at https://github.com/yingjiayu12/Exemplar2VQA

Figures & tables

Appendix figures & tables14 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. IDEAL-Bench: Indoor Dataset and Evaluation suite for Analyzing 3D Layout reasoning

    Jul 3, 2026Yuening Cai, Junwei Zhou, Youran Qu +13D Layout GenerationVisual Question Answering Benchmarks

  2. M3^3-VQA: A Benchmark for Multimodal, Multi-Entity, Multi-Hop Visual Question Answering

    Apr 28, 2026Jiatong Ma, Longteng Guo, Yuchen Liu +4Knowledge-Based Visual Question AnsweringMultimodal Reasoning

  3. SpatialForge: Bootstrapping 3D-Aware Spatial Reasoning from Open-World 2D Images

    May 12, 2026Zishan Liu, Ruoxi Zang, Yanglin Zhang +53D Spatial ReasoningSpatial Reasoning