GroundSight at GroundLM 2026 Shared Tasks: GoldenViewVQA
Authors: Kun Wang, Yupeng Hu, Ruping Cao, Hao Liu, Zhiran Li, Qianlong Xiang, Harry Cheng
Organizations: School of Software, Shandong University, Jinan, China · School of Computing, National University of Singapore, Singapore · School of Computer Science and Technology, Harbin Institute of Technology (Shenzhen), Shenzhen, China
GoldenViewVQA requires models to jointly answer driving-scene questions and identify the camera view containing the supporting visual evidence, making precise evidence localization as important as answer correctness. We present \textbf{CoVeR-VQA}, a training-free multi-stage verification and correction framework for grounded multi-view VQA. Starting from GPT-5.6 zero-shot predictions, CoVeR-VQA progressively applies view-specific verification with Gemini-3.6-Flash, prior-guided joint verification with Claude-Opus-5, and cross-split group-level verification that exploits semantically filtered question groups from shared multi-view scenes and validation-derived prior knowledge. On the official GoldenViewVQA test set, the four-stage CoVeR-VQA pipeline achieves 84.75% Joint Accuracy, improving the GPT-5.6 zero-shot baseline by 13.56 percentage points, while reaching 94.92% Answer Accuracy and 86.44% View Accuracy. The final submitted run achieves 88.14% Joint Accuracy after two additional evaluator-informed post-hoc corrections. Our analysis shows that supporting-view localization remains the primary source of residual errors, highlighting the importance of explicit evidence verification for reliable multi-view multimodal reasoning.
Figures & tables
Figure 1: Overview of the proposed CoVeR-VQA framework and its four-stage prediction and verification pipeline.
Method
Joint
Ans.
View
V-Macro
Organizer Baseline
20.34
23.73
77.97
16.67
Qwen3-VL-8B Zero Shot
66.10
88.14
71.19
49.88
\raisebox{-.15ex}{\scriptsize1}⃝ GPT-5.6 Zero Shot
\raisebox{-.15ex}{\scriptsize3}⃝ Claude Joint Review
79.66
93.22
83.05
54.83
Table 1: Results on the GoldenViewVQA test set (%). Shaded rows marked with \raisebox{-.15ex}{\scriptsize1}⃝ – \raisebox{-.15ex}{\scriptsize4}⃝ denote the four stages of the CoVeR-VQA pipeline, while “+” denotes an additional ablation setting.
Figure 2: Distribution of answer and supporting-view correctness for the final submission on the GoldenViewVQA test set.
Figure 3: Example of a view-selection error caused by confusing object saliency with decisive spatial evidence.
Figure 4: Example of unsupported semantic inference under ambiguous visual evidence.
Establishing a clear link between model predictions and the visual evidence that supports them is critical for transparency and reliability in multimodal reasoning, yet current multimodal large language model (MLLM) evaluations do not explicitly enforce this alignment. Existing benchmarks assess either textual answer correctness or pixel-level localization in isolation, leaving the coupling of reasoning and grounding an open challenge. We introduce VISTAQA, a comprehensive benchmark for joint evaluation of free-form answer correctness and pixel-level evidence grounding in visual question answering. VISTAQA comprises 1,157 expert-curated samples spanning six task types and six visual domains, ranging from direct perception to compositional and relational reasoning. VISTAQA requires models to not only answer correctly, but to also provide precise segmentation masks that support their answers. It also includes hallucination-aware examples where no valid visual evidence exists. To support this enhanced evaluation, we introduce GROVE, a unified evaluation metric that enforces joint correctness by combining textual accuracy and grounding quality via a per-sample geometric mean, ensuring neither dimension can compensate for deficiencies in the other. Comprehensive experiments across grounding-aware models and hybrid pipelines with general-purpose MLLMs reveal that even the strongest systems achieve limited performance under GROVE, highlighting a substantial gap between answer accuracy and visual evidence alignment.
Mozhgan Nasr Azadani, Yimu Wang, Yongpeng Zhu +5
University of Waterloo · Stanford University · NVIDIA
Multimodal large language models (MLLMs) achieve strong results on visual reasoning benchmarks, but answer accuracy alone does not indicate whether a model relied on the correct visual evidence. This gap is particularly important in multi-view driving scenes used for autonomous driving, where a model can produce a plausible answer while grounding it in the wrong camera view. We introduce a multi-view visual question answering benchmark for evaluating evidence-source identification: given six synchronized NuScenes views and a question, the model must identify the supporting camera view and answer the question. The benchmark contains 122 conflict-centric question-answer pairs from 73 scenes, spanning causality, counterfactual reasoning, and intent prediction. View labels are proposed by an automatic conflict-mining pipeline and manually verified by annotators. We evaluate three settings: camera-view selection, oracle QA given the golden view, and joint prediction in which the model selects a view and answers in one pass. Answers are evaluated in both multiple-choice and free-form formats, using exact match for structured predictions and an LLM judge for free-form responses. By explicitly separating visual-source identification from answer correctness, the benchmark exposes grounding failures that answer-only evaluation misses.
Manual annotation of high-quality visual question answering with grounding (VQA-G) datasets, which pair visual questions with evidential grounding, is crucial for advancing vision-language models (VLMs), but remains unscalable. Existing automated methods are often hindered by two key issues: (1) inconsistent data fidelity due to model hallucinations; (2) brittle verification mechanisms based on simple heuristics. To address these limitations, we introduce AutoVQA-G, a self-improving agentic framework for automated VQA-G annotation. AutoVQA-G employs an iterative refinement loop where a Consistency Evaluation module uses Chain-of-Thought (CoT) reasoning for fine-grained visual verification. Based on this feedback, a memory-augmented Prompt Optimization agent analyzes critiques from failed samples to progressively refine generation prompts. Our experiments show that AutoVQA-G generates VQA-G datasets with superior visual grounding accuracy compared to leading multimodal LLMs, offering a promising approach for creating high-fidelity data to facilitate more robust VLM training and evaluation. Code: https://github.com/rohnson1999/AutoVQA-G
Rongsheng Hu, Runwei Guan, Yicheng Di +2
School of Artificial Intelligence and Computer Science, Jiangnan University, Wuxi, China · Thrust of Artificial Intelligence, HKUST(GZ), Guangzhou, China