cs.CVSep 27, 2026

Informative Viewpoint Selection for Episodic-Memory Embodied Question Answering using Omnidirectional Images

Authors: Kaname Kitamura, Asako Kanezaki

Organizations: Department of Computer Science, Institute of Science Tokyo, Tokyo, Japan · RIKEN, Japan · Tohoku University, Sendai, Japan

Abstract

Embodied Question Answering (EQA) requires agents to answer natural language questions about surrounding environments from visual observations. In this work, we focus on open-vocabulary episodic-memory EQA (EM-EQA), where an agent answers free-form questions using recorded observation histories. Omnidirectional images are promising for this task, as they provide wide field-of-view observations that can capture surrounding context without requiring explicit camera rotations. However, omnidirectional images introduce two challenges for EQA: (i) equirectangular projection causes severe geometric distortion that degrades vision-language model (VLM) recognition accuracy, and (ii) feeding equirectangular images directly into VLMs introduces excessive irrelevant background information, reducing answer accuracy and increasing the visual-token burden. To address these challenges, we propose a viewpoint selection method for EM-EQA using omnidirectional images. Our method converts equirectangular observations into perspective views via cubemap projection, estimates question-conditioned relevance with fine-tuned BLIP-2, and selects informative and diverse viewpoints through diversity-aware greedy selection. Experiments on the Habitat-Matterport 3D (HM3D) subset of OpenEQA show that our method achieves state-of-the-art model performance among the reported model results with equirectangular observations. Moreover, after removing rotation views, which reduces observation frames by 65.5%, our method largely maintains its answer accuracy.

Figures & tables

Explore similar work

CardsList
  1. ScoutVLA: UAV-Centric Active Perception via a Dual-Expert VLA Model for Open-World Embodied Question Answering

    Jun 9, 2026Wenhao Lu, Zhengqiu Zhu, Xiaofeng Wang +7Aerial Vision-Language NavigationVisual Question Answering Benchmarks

  2. Hierarchical Floorplan-Guided Vision-Language Exploration for Embodied Question Answering

    Sep 22, 2026Albert Gassol Puigjaner, Kostas AlexisVision-Language NavigationKnowledge-Based Visual Question Answering

  3. Beyond Episodic Evaluation: Memory Architectural Bottlenecks in Sequential Embodied Question Answering

    Jul 23, 2026Zikui Cai, Kaushal Janga, Tan Dat Dao +15Visual MemoryEmbodied