Embodied Question Answering (EQA) requires agents to answer natural language questions about surrounding environments from visual observations. In this work, we focus on open-vocabulary episodic-memory EQA (EM-EQA), where an agent answers free-form questions using recorded observation histories. Omnidirectional images are promising for this task, as they provide wide field-of-view observations that can capture surrounding context without requiring explicit camera rotations. However, omnidirectional images introduce two challenges for EQA: (i) equirectangular projection causes severe geometric distortion that degrades vision-language model (VLM) recognition accuracy, and (ii) feeding equirectangular images directly into VLMs introduces excessive irrelevant background information, reducing answer accuracy and increasing the visual-token burden. To address these challenges, we propose a viewpoint selection method for EM-EQA using omnidirectional images. Our method converts equirectangular observations into perspective views via cubemap projection, estimates question-conditioned relevance with fine-tuned BLIP-2, and selects informative and diverse viewpoints through diversity-aware greedy selection. Experiments on the Habitat-Matterport 3D (HM3D) subset of OpenEQA show that our method achieves state-of-the-art model performance among the reported model results with equirectangular observations. Moreover, after removing rotation views, which reduces observation frames by 65.5%, our method largely maintains its answer accuracy.
Figures & tables
Figure 1 : Overview of our viewpoint selection method for EM-EQA using omnidirectional images. Given an episode history of omnidirectional images and a natural language question, we convert images to perspective views via cubemap projection, compute relevance scores using fine-tuned BLIP-2, and select diverse informative viewpoints through our Diversity-Aware Greedy Selection algorithm for input to the VLM.
Figure 2 : Cubemap projection converts an equirectangular image into six perspective views, and Diversity-Aware Greedy Selection selects informative viewpoints.
Figure 3 : Architecture of fine-tuned BLIP-2 for relevance estimation.
Setting
Avg. frames
Reduction rate
Original
170.48
–
w/o rotation
58.76
65.5%
Table 1 : Average number of stored observation frames. Reduction rate is relative to Original .
Method
Setting
Original
w/o rotation
Human baseline [ 25 ]
perspective
85.1
–
GPT-4 (Blind LLMs) [ 25 ]
–
35.5
–
LLaMA-2 (Blind LLMs) [ 25 ]
–
29.0
–
GPT-4 w/ LLaVA-1.5 [ 25 ]
perspective
40.0
–
LLaMA-2 w/ LLaVA-1.5 [ 25 ]
perspective
31.1
–
GPT-4 w/ CG [ 25 ]
perspective
34.0
–
Table 2 : OpenEQA evaluation on EM-EQA in HM3D. Scores are maximum LLM-Match (%) over K∈{1,…,10} for our experiments. Previously reported entries are quoted from the cited papers. Human baseline denotes the reported human performance, Blind LLMs denotes text-only models without visual input, CG denotes ConceptGraphs, and SVM denotes Sparse Voxel Map.
Equirectangular
Standard perspective
Original
w/o rotation
Original
w/o rotation
K
Multi
RAG
RAG cubemap
Ours
Multi
RAG
RAG cubemap
Ours
Multi
RAG
Ours
Multi
RAG
Ours
1
35.0
49.7
41.7
50.5
34.7
49.4
37.4
51.3
30.1
44.9
52.7
30.6
37.5
41.8
2
45.6
49.2
43.7
53.4
45.4
49.1
40.2
54.4
35.8
47.4
54.9
38.2
37.1
45.0
3
51.4
48.2
43.8
57.0
49.6
51.1
40.4
54.9
44.5
49.2
57.2
42.0
38.0
44.9
4
50.8
48.7
42.1
55.7
50.9
50.7
40.8
55.3
43.9
50.1
59.4
41.1
39.1
44.7
Table 3 : Full LLM-Match scores (%) over the number of selected frames K . Multi denotes Multi-Frame VLMs, RAG denotes RAG with BLIP, and bold indicates the highest score over K for each method and observation setting.
Table 7
Table 6 : Ablation studies at K=5 on equirectangular observations in the Original setting. Scores are LLM-Match (%). FT denotes Fine-tuning, and DAGS denotes Diversity-Aware Greedy Selection. R-EQA-style ranks views by caption similarity without BLIP-2, and MMR uses λ=0.5 . The similarity sweep fixes γ=0.1 , and the score sweep fixes τ=0.97 .
Figure 4 : Qualitative comparison of frame selection.
Equirectangular
Standard perspective
Original
w/o rotation
Original
w/o rotation
Stage
Multi
RAG
RAG (cube)
Ours
Multi
RAG
RAG (cube)
Ours
Multi
RAG
Ours
Multi
RAG
Ours
Selection
0.00
9.54
16.56
25.23
0.00
3.47
5.59
9.40
0.00
8.64
11.50
0.00
3.14
4.11
Answer
8.39
8.21
1.89
2.53
8.32
8.29
1.89
2.70
8.34
8.48
8.47
8.47
8.52
8.46
Total
8.39
17.75
18.45
27.76
8.32
11.76
7.48
12.10
8.34
17.12
19.97
8.47
11.66
12.57
Table 7 : Inference time components at K=5 . Times are seconds per question. Selection denotes viewpoint selection time, Answer denotes answer generation time, and Total is their sum. Multi denotes Multi-Frame VLMs, RAG denotes RAG with BLIP, and RAG (cube) denotes cubemap-based RAG with BLIP.
λ
0.1
0.2
0.3
0.4
0.5
0.6
0.7
0.8
0.9
LLM-Match
50.6
51.7
54.4
55.3
55.8
54.5
53.8
54.0
54.2
Table S1 : MMR with different relevance weights λ .
Figure S1 : Views selected by our method for two qualitative examples.
Fallback
Questions
Triggered
Not triggered
Uniform sampling ( ∣Chigh∣=0 )
48 (8.6%)
53.1
57.9
Similarity-constraint relaxation
28 (5.0%)
39.3
58.5
Filling from Clow
22 (3.9%)
23.9
58.9
Table S2 : Frequency and LLM-Match of each fallback mechanism. Chigh and Clow denote the candidates whose ITM scores are at least and below γ , respectively.
Relevance model
LLM-Match
Fine-tuned BLIP-2 (ours)
55.2
Pretrained BLIP-2
41.7
Pretrained SigLIP 2 So400m
27.4
Table S3 : Top- K selection with different relevance models.
Component
Time (s)
Image decoding
8.03
Cubemap generation
3.95
BLIP-2 preprocessing
4.98
BLIP-2 relevance estimation
8.20
Candidate embedding extraction
0.90
Diversity-Aware Greedy Selection
0.001
Table S4 : Runtime of each stage (seconds per question), measured on four NVIDIA H100 GPUs with a batch size of 32 per GPU.
Aerial Embodied Question Answering (EQA) requires Unmanned Aerial Vehicles (UAVs) to actively perceive the environment and answer natural language questions. Existing outdoor EQA systems usually stop once the target enters the UAV's field of view, leaving the fine-grained viewpoint adjustment needed for evidence-seeking questions largely unresolved. To address this issue, we introduce FG-EQA, a fine-grained active perception EQA benchmark with more than 40K simulated trajectories and 1K real-world trajectories. Drawing inspiration from the ``waggle dance'' of scout bees, which iteratively adjust their flight paths to verify target information, we propose ScoutVLA, an evidence-driven Vision-Language-Action model for outdoor EQA. To emulate this active exploration behavior, ScoutVLA features a decoupled dual-expert architecture: a vision-language expert infers the semantic intent to identify missing evidence, while an independent action expert employs high-DoF flow matching to generate continuous viewpoint-refinement trajectories. To balance the competing demands of continuous control and semantic reasoning, we devise a decoupled training strategy with a knowledge insulation mechanism that prevents the action gradients from erasing the model's multimodal reasoning ability. Extensive simulated experiments and a qualitative real-world field study both verify the superiority of ScoutVLA over the state-of-the-art baselines, demonstrating a 10.48× higher average strict success rate and a 7.72× higher average QA correctness.
Wenhao Lu, Zhengqiu Zhu, Xiaofeng Wang +7
1National Key Laboratory of Digital Intelligent Modeling and Simulation, National University of Defense Technology · 2GigaAI
Embodied Question Answering (EQA) requires an agent to explore a previously unseen environment, gather relevant information, and answer questions about the scene. Recent approaches leverage Vision-Language Models (VLMs) together with semantic maps or scene graphs to guide exploration. However, exploration is typically driven only by local observations, while structural priors about the environment remain largely unused. We propose HFLEX-EQA, a hierarchical EQA framework that combines online scene graph construction, VLM- based planning, semantic frontier exploration, and floorplan priors. The system incrementally builds a hierarchical scene graph and an open-vocabulary occupancy map from RGB-D observations, enabling a VLM to jointly reason over the scene graph, task-relevant visual observations, exploration history, and an estimated topological floorplan. Furthermore, we introduce a room-discovery strategy that leverages the floorplan and open-vocabulary frontier semantics to guide exploration toward semantically relevant yet currently unobserved room types. We evaluate HFLEX-EQA on the OpenEQA and ExploreEQA benchmarks and demonstrate deployment on a quadruped robot in real indoor environments. Our results demonstrate the benefit of combining VLM-based hierarchical planning with structural floorplan priors for the EQA task.
Albert Gassol Puigjaner, Kostas Alexis
Autonomous Robots Lab, Norwegian University of Science and Technology (NTNU), Trondheim, Norway
Embodied question answering (EQA) is traditionally evaluated under an episodic formulation, where agents solve each task independently and reset internal state between episodes. However, real-world robots operate continuously and must accumulate, retain, and selectively reuse information acquired from prior interactions. Despite this practical requirement, the architectural mechanisms needed to support sequential memory in EQA remain underexplored. In this work, we investigate how different memory architectures behave when EQA agents are evaluated sequentially, with multiple questions answered in the same scene while memory is carried forward across queries. We find that simply preserving existing memory is often insufficient. Agents that retain only traversability information, such as 2D occupancy maps, remember where the robot has explored but not the visual-semantic evidence needed for later questions. Agents trained on short-horizon episodic data face a different challenge: when exposed to continuous, multi-query histories, their inherited context suffers from severe temporal mismatch, rather than forming a reusable scene representation. To overcome this architectural bottleneck, we highlight the necessity of structured, spatially grounded memory: architectures that map persistent visual observations onto metric 3D geometry preserve visual-semantic evidence in a coherent scene representation. Extensive experiments in simulated environments reveal that this form of memory breaks the accuracy-efficiency tradeoff in sequential settings, simultaneously achieving higher answer accuracy and lower navigation costs. We further validate these findings on a real-world mobile robot, demonstrating that spatially grounded visual memory is critical for enabling continuous, intelligent operation in physical environments.
Zikui Cai, Kaushal Janga, Tan Dat Dao +15
University of Maryland, College Park. · The University of Texas at Austin. · University of Illinois Urbana-Champaign.