Embodied Question Answering (EQA) requires agents to answer natural language questions about surrounding environments from visual observations. In this work, we focus on open-vocabulary episodic-memory EQA (EM-EQA), where an agent answers free-form questions using recorded observation histories. Omnidirectional images are promising for this task, as they provide wide field-of-view observations that can capture surrounding context without requiring explicit camera rotations. However, omnidirectional images introduce two challenges for EQA: (i) equirectangular projection causes severe geometric distortion that degrades vision-language model (VLM) recognition accuracy, and (ii) feeding equirectangular images directly into VLMs introduces excessive irrelevant background information, reducing answer accuracy and increasing the visual-token burden. To address these challenges, we propose a viewpoint selection method for EM-EQA using omnidirectional images. Our method converts equirectangular observations into perspective views via cubemap projection, estimates question-conditioned relevance with fine-tuned BLIP-2, and selects informative and diverse viewpoints through diversity-aware greedy selection. Experiments on the Habitat-Matterport 3D (HM3D) subset of OpenEQA show that our method achieves state-of-the-art model performance among the reported model results with equirectangular observations. Moreover, after removing rotation views, which reduces observation frames by 65.5%, our method largely maintains its answer accuracy.
Figures & tables
Figure 1 : Overview of our viewpoint selection method for EM-EQA using omnidirectional images. Given an episode history of omnidirectional images and a natural language question, we convert images to perspective views via cubemap projection, compute relevance scores using fine-tuned BLIP-2, and select diverse informative viewpoints through our Diversity-Aware Greedy Selection algorithm for input to the VLM.
Figure 2 : Cubemap projection converts an equirectangular image into six perspective views, and Diversity-Aware Greedy Selection selects informative viewpoints.
Figure 3 : Architecture of fine-tuned BLIP-2 for relevance estimation.
Setting
Avg. frames
Reduction rate
Original
170.48
–
w/o rotation
58.76
65.5%
Table 1 : Average number of stored observation frames. Reduction rate is relative to Original .
Method
Setting
Original
w/o rotation
Human baseline [ 25 ]
perspective
85.1
–
GPT-4 (Blind LLMs) [ 25 ]
–
35.5
–
LLaMA-2 (Blind LLMs) [ 25 ]
–
29.0
–
GPT-4 w/ LLaVA-1.5 [ 25 ]
perspective
40.0
–
LLaMA-2 w/ LLaVA-1.5 [ 25 ]
perspective
31.1
–
GPT-4 w/ CG [ 25 ]
perspective
34.0
–
Table 2 : OpenEQA evaluation on EM-EQA in HM3D. Scores are maximum LLM-Match (%) over K∈{1,…,10} for our experiments. Previously reported entries are quoted from the cited papers. Human baseline denotes the reported human performance, Blind LLMs denotes text-only models without visual input, CG denotes ConceptGraphs, and SVM denotes Sparse Voxel Map.
Equirectangular
Standard perspective
Original
w/o rotation
Original
w/o rotation
K
Multi
RAG
RAG cubemap
Ours
Multi
RAG
RAG cubemap
Ours
Multi
RAG
Ours
Multi
RAG
Ours
1
35.0
49.7
41.7
50.5
34.7
49.4
37.4
51.3
30.1
44.9
52.7
30.6
37.5
41.8
2
45.6
49.2
43.7
53.4
45.4
49.1
40.2
54.4
35.8
47.4
54.9
38.2
37.1
45.0
3
51.4
48.2
43.8
57.0
49.6
51.1
40.4
54.9
44.5
49.2
57.2
42.0
38.0
44.9
4
50.8
48.7
42.1
55.7
50.9
50.7
40.8
55.3
43.9
50.1
59.4
41.1
39.1
44.7
Table 3 : Full LLM-Match scores (%) over the number of selected frames K . Multi denotes Multi-Frame VLMs, RAG denotes RAG with BLIP, and bold indicates the highest score over K for each method and observation setting.
Table 7
Table 6 : Ablation studies at K=5 on equirectangular observations in the Original setting. Scores are LLM-Match (%). FT denotes Fine-tuning, and DAGS denotes Diversity-Aware Greedy Selection. R-EQA-style ranks views by caption similarity without BLIP-2, and MMR uses λ=0.5 . The similarity sweep fixes γ=0.1 , and the score sweep fixes τ=0.97 .
Figure 4 : Qualitative comparison of frame selection.
Equirectangular
Standard perspective
Original
w/o rotation
Original
w/o rotation
Stage
Multi
RAG
RAG (cube)
Ours
Multi
RAG
RAG (cube)
Ours
Multi
RAG
Ours
Multi
RAG
Ours
Selection
0.00
9.54
16.56
25.23
0.00
3.47
5.59
9.40
0.00
8.64
11.50
0.00
3.14
4.11
Answer
8.39
8.21
1.89
2.53
8.32
8.29
1.89
2.70
8.34
8.48
8.47
8.47
8.52
8.46
Total
8.39
17.75
18.45
27.76
8.32
11.76
7.48
12.10
8.34
17.12
19.97
8.47
11.66
12.57
Table 7 : Inference time components at K=5 . Times are seconds per question. Selection denotes viewpoint selection time, Answer denotes answer generation time, and Total is their sum. Multi denotes Multi-Frame VLMs, RAG denotes RAG with BLIP, and RAG (cube) denotes cubemap-based RAG with BLIP.
λ
0.1
0.2
0.3
0.4
0.5
0.6
0.7
0.8
0.9
LLM-Match
50.6
51.7
54.4
55.3
55.8
54.5
53.8
54.0
54.2
Table S1 : MMR with different relevance weights λ .
Figure S1 : Views selected by our method for two qualitative examples.
Fallback
Questions
Triggered
Not triggered
Uniform sampling ( ∣Chigh∣=0 )
48 (8.6%)
53.1
57.9
Similarity-constraint relaxation
28 (5.0%)
39.3
58.5
Filling from Clow
22 (3.9%)
23.9
58.9
Table S2 : Frequency and LLM-Match of each fallback mechanism. Chigh and Clow denote the candidates whose ITM scores are at least and below γ , respectively.
Relevance model
LLM-Match
Fine-tuned BLIP-2 (ours)
55.2
Pretrained BLIP-2
41.7
Pretrained SigLIP 2 So400m
27.4
Table S3 : Top- K selection with different relevance models.
Component
Time (s)
Image decoding
8.03
Cubemap generation
3.95
BLIP-2 preprocessing
4.98
BLIP-2 relevance estimation
8.20
Candidate embedding extraction
0.90
Diversity-Aware Greedy Selection
0.001
Table S4 : Runtime of each stage (seconds per question), measured on four NVIDIA H100 GPUs with a batch size of 32 per GPU.