JRDB-AVR: An Active Visual Reasoning Benchmark for Embodied Agents in Real-World Environments
Organizations: Monash University
Abstract
In complex embodied visual reasoning scenarios, an agent often has only a limited field of view, and the evidence needed to answer a question may be distributed across time, viewpoint, and interacting objects. A model may therefore give a plausible answer without ever observing the relevant object, time, or view that supports it. Current visual reasoning benchmarks largely evaluate passive observations and final answers, overlooking settings that require active reasoning and evidence acquisition. We introduce JRDB-AVR, a benchmark derived from existing real-world JRDB robotics data through a structured question-generation engine that turns this gap into an explicit evaluation: an embodied agentic system receives a visual reasoning question, requests bounded observations by timestamp and viewing angle, and is evaluated on both the final answer and the grounded visual evidence supporting it. The benchmark contains diverse questions over multiple real-world environments involving temporal search, viewpoint selection, and human-oriented compositional reasoning. We also introduce JRDB-AVR-Agent, a reference active reasoning agentic method that maintains an explicit observation-grounded graph-based world model and answers through solving. Experiments reveal a substantial gap between answer accuracy and evidence accuracy in current baselines, showing that current VLMs can produce unsupported correct answers and that active evidence-aware evaluation is necessary for embodied visual reasoning. Code and benchmark are available at https://github.com/ControlNet/JRDB-AVR.
Figures & tables
| Benchmark | Temporal | Spatial/View | Domain | Camera | Active Obs. | Evidence | Evaluation |
|---|---|---|---|---|---|---|---|
| GQA ( Hudson and Manning, 2019 ) | ✓ | Real | Static | Answer | |||
| CLEVRER ( Yi et al., 2019 ) | ✓ | Sim | Static | Answer | |||
| STAR ( Wu et al., 2021 ) | ✓ | ✓ | Real | Static | Answer | ||
| MindCube ( Wang et al., 2025 ) | ✓ | Real | Static | Answer | |||
| VIEW2SPACE ( Ke et al., 2026 ) | ✓ | Sim | Static | BBox | Answer | ||
| JRDB-Reasoning ( Jahangard et al., 2026 ) | ✓ | ✓ | Real | Moving | Answer |
| Name | Description | Type | Evaluation Protocol |
| Chain 2-hop | Ask about the target with 2-hop entity relations. | Multiple choice | Correct if the selected choice matches the ground truth. |
| Chain 3-hop | Extend the relational path to three hops before querying the target. | ||
| Chain fork-join | Merge two relational branches from a shared entity before querying the target. | ||
| Unique anchor | Locate a target from conditions and ask about a later state or action in another temporal location. | ||
| Long range | Track temporally distant consequences or long-range dependencies. | ||
| Chain hybrid | Combine a spatial relation with temporal and viewpoint search to locate the target timestamp and view angle. | Timestamp and viewpoint | Correct if the predicted timestamp is within second of the ground truth and the wrapped viewpoint error is within . |
| Backbone VLM | Method | Answer | Evidence | Combined | Hallucination |
|---|---|---|---|---|---|
| Gemma-4 E2B (5B) | Monolithic VLM | 29.12 | 0.05 | 0.05 | 99.83 |
| Chain-of-Thought | 25.31 | 0.05 | 0.05 | 99.81 | |
| Search-Recognize-Pipeline | 26.69 | 4.00 | 0.81 | 96.96 | |
| ReAct | 29.50 | 5.67 | 1.91 | 93.54 | |
| Gemma-4 E4B (8B) | Monolithic VLM | 29.50 | 0.14 | 0.14 | 99.52 |
| Chain-of-Thought | 31.22 | 0.29 | 0.10 | 99.69 |
| Active | World Model | Answer | Evidence | Combined | Hallucination |
|---|---|---|---|---|---|
| 25.31 | 6.86 | 2.72 | 89.25 | ||
| ✓ | 26.83 | 8.06 | 3.62 | 86.51 | |
| ✓ | 37.66 | 18.22 | 9.20 | 75.57 | |
| ✓ | ✓ | 38.51 | 27.50 | 14.54 | 62.25 |
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
| Family | Count | Template(s) | Answer mode | Active reasoning requirement |
|---|---|---|---|---|
| Action boundary | 67 | start, span | Frame identification / temporal localization | Localize the onset of an action or recover its full temporal interval instead of answering from a single static frame. |
| Chain 2-hop | 76 | action | Multiple choice | Follow a unique two-hop relational path from the anchor person to the final target before reading out the answer. |
| Chain 3-hop | 222 | action | Multiple choice | Extend the path to three hops, so the answer depends on correctly grounding an additional intermediate person. |
| Chain fork-join | 32 | action | Multiple choice | Ground two branches from the same anchor and merge them at a shared target, requiring branch consistency rather than a single linear chain. |
| Chain hybrid | 51 | action search | Frame-viewpoint pair | Combine a spatial chain with temporal search and viewpoint selection, then return both the decisive frame and the best supporting viewpoint. |
| Long range | 30 | chain, consequence, gap | Multiple choice / numeric value | Recover a target over a larger temporal gap and either predict a later action/consequence or a temporally defined numeric gap. |
| Generator | Count | Answer | Evidence | Combined | Hallucination |
|---|---|---|---|---|---|
| Action boundary | 67 | 37.31 | 56.72 | 23.88 | 36.00 |
| Chain family | 330 | 41.82 | 8.48 | 5.45 | 86.96 |
| Chain hybrid | 51 | 11.76 | 5.88 | 0.00 | 100.00 |
| Long range | 30 | 50.00 | 3.33 | 0.00 | 100.00 |
| Unique anchor | 1,620 | 38.52 | 31.30 | 16.73 | 56.57 |