Organizations: The University of Tokyo, Japan · National Institute of Informatics, Japan · Institute of Science Tokyo, Japan · NII LLMC, Japan · The University of Osaka, Japan
Active robot vision requires controlling the camera to reveal task-relevant information that is hidden from the current viewpoint. For example, determining what is inside a box may require raising the camera and looking down into it. For viewpoint-dependent question answering, the challenge is to select camera motions that expose the visual evidence needed to answer the question. Although vision-language models (VLMs) can interpret observed images, selecting such motions requires anticipating the usefulness of unseen views. We quantify this usefulness as answerability, a VLM's estimate that a view suffices to answer the question, and present Rendering-Free Lookahead (RFL), a viewpoint-selection policy that ranks candidate camera motions by predicted future answerability. RFL transfers visual lookahead from deployment to offline training. At training, a privileged teacher renders candidate future views in 3D Gaussian Splatting (3DGS) scenes and uses a frozen VLM to compute one- and two-step answerability targets. Through two-stage distillation, a student learns to predict these action values from the question, recent visual observations, and a candidate camera motion. At deployment, RFL uses these predicted values to select camera motions without rendering future views. On 377 E3VS-Bench test episodes in unseen environments, RFL improves the mean judge score by 43% over a direct-action baseline using the same VLM. These results support learning camera-control policies from privileged visual lookahead for viewpoint-dependent question answering.
Figures & tables
Fig. 1: Rendering-free lookahead. RFL selects camera motions by predicted answerability, which rises as the agent closes in on a view that reveals the answer (right), with no rendering or generation at deployment.
Fig. 2: Overview of RFL. A privileged teacher renders each candidate action’s one- and two-step future views in 3DGS and scores them with the frozen VLM’s answerability probe, yielding the lookahead targets Q1 and Q2 ; the student learns to predict these values from the question, past views, and the action, requiring no rendering at deployment.
Fig. 3: Answerability tracks viewpoint search. Mean probed answerability vs. normalized progress on Gemini 3.0 Flash rollouts over the 231 validation episodes (bands: ±1 SE; dotted: mean at annotated goal views; Sec. V-C ).
Method
OS
OST
OA
CGS
SR
CNT
Avg. ↑
Coll. ↓
Proprietary
GPT-5.1
3.34
2.73
2.96
2.87
2.38
2.28
2.81
0.38
Gemini 3.0 Flash
3.14
3.09
2.67
2.47
2.25
1.96
2.62
0.43
Open-source
Cosmos3 (Planning)
2.83
2.18
2.38
2.33
2.25
1.92
2.36
0.50
Qwen3.6-27B
2.79
1.82
2.09
2.47
1.50
1.64
2.14
0.55
TABLE I: Main results on the E3VS-Bench test set. Green is best and light green is second best per column. OS = object search, OST = object state, OA = object attribute, CGS = context-guided search, SR = spatial reasoning, CNT = counting.
Fig. 4: Answerability along test rollouts for the agents of Table I (bands: ±1 SE; dashed: mean at goal views).
Figure 6
Fig. 8: Qualitative comparison on E3VS-Bench (RFL: green; Gemini 3.0 Flash: magenta). Full answers (abbreviated in the figure): middle, RFL: “The lid on the travel mug is beige.”; right, Gemini: “The text on the front of the plastic sack is not legible.”, RFL: “The text is not clear enough to be read.”
Fig. 9: Real-world experiments. Mean judge score per question for RFL, the direct base VLM (Base), and the GPT-5.1 agent. Each bar averages ten trials of one method on one question, i.e., 3 questions × 3 methods × 10 trials =90 trials.
Output
FT
Avg. ↑
Steps ↓
Rep. ↓
Actions
2.14
20.1
0.78
✓
2.47
24.9
0.87
Answerability
2.04
18.8
0.68
✓
2.71
17.2
0.62
TABLE II: Output and fine-tuning of the same VLM (Qwen3.6-27B) with single-view conditioning.