cs.CVOct 8, 2026

GroundSight at GroundLM 2026 Shared Tasks: GoldenViewVQA

Authors: Kun Wang, Yupeng Hu, Ruping Cao, Hao Liu, Zhiran Li, Qianlong Xiang, Harry Cheng

Organizations: School of Software, Shandong University, Jinan, China · School of Computing, National University of Singapore, Singapore · School of Computer Science and Technology, Harbin Institute of Technology (Shenzhen), Shenzhen, China

Abstract

GoldenViewVQA requires models to jointly answer driving-scene questions and identify the camera view containing the supporting visual evidence, making precise evidence localization as important as answer correctness. We present \textbf{CoVeR-VQA}, a training-free multi-stage verification and correction framework for grounded multi-view VQA. Starting from GPT-5.6 zero-shot predictions, CoVeR-VQA progressively applies view-specific verification with Gemini-3.6-Flash, prior-guided joint verification with Claude-Opus-5, and cross-split group-level verification that exploits semantically filtered question groups from shared multi-view scenes and validation-derived prior knowledge. On the official GoldenViewVQA test set, the four-stage CoVeR-VQA pipeline achieves 84.75% Joint Accuracy, improving the GPT-5.6 zero-shot baseline by 13.56 percentage points, while reaching 94.92% Answer Accuracy and 86.44% View Accuracy. The final submitted run achieves 88.14% Joint Accuracy after two additional evaluator-informed post-hoc corrections. Our analysis shows that supporting-view localization remains the primary source of residual errors, highlighting the importance of explicit evidence verification for reliable multi-view multimodal reasoning.

Figures & tables

Explore similar work

CardsList
  1. VISTAQA: Benchmarking Joint Visual Question Answering and Pixel-Level Evidence

    May 20, 2026Mozhgan Nasr Azadani, Yimu Wang, Yongpeng Zhu +5VLM EvaluationVisual Question Answering

  2. AutoVQA-G: Self-Improving Agentic Framework for Automated Visual Question Answering and Grounding Annotation

    Apr 19, 2026Rongsheng Hu, Runwei Guan, Yicheng Di +2Visual Question AnsweringLLM-Assisted Annotation