RoboQuest: Generalist Physical Agents that Search, Inspect and Test
Organizations: Nanyang Technological University
Abstract
Recent advances in multimodal foundation models have made them capable generalist physical agents for a range of manipulation tasks. However, successful operation in an unfamiliar environment may require an agent to seek task-relevant information through interaction when it is absent from the observations: it may need to determine where a relevant object is, inspect an unobserved property, or discover the effect of an unfamiliar tool. We thus introduce RoboQuest, a benchmark for goal-directed embodied exploration, where agents must actively acquire task-relevant information through physical interaction, use the resulting evidence to adapt subsequent actions, and autonomously decide when to commit to task completion. RoboQuest comprises ten mobile manipulation tasks centered on three forms of uncertainty: search, manipulation-based inspection, and interactive testing. We evaluate five frontier multimodal agents through a common visuomotor interface, as well as a policy fine-tuned on the full-episode demonstrations we release. The best agent succeeds in only 23% of the episodes, and the fine-tuned policy almost never succeeds. Isolated tests of the execution skills the tasks are built from, with the hidden information supplied, show that the agents can carry out most of the required actions, and our failure analysis attributes only a minority of the failures to execution. Our failure analysis further finds that the agents often stop exploring too early as they make decisions before observing the required evidence for task completion. We also find that agents rarely prevent or repair the disturbances caused by their exploration. Moreover, learning by trial and error remains difficult for most models.
Figures & tables
| Benchmark features | Evidence seeking | |||||||
| Benchmark | Focus | Long horizon | Hidden | Memory | Mobile | Search | Inspect | Test |
| RLBench ( James et al., 2020 ) | Visuomotor skills | – | – | – | – | – | ||
| ManiSkill3 ( Tao et al., 2024 ) | Scalable manipulation | – | – | – | – | – | – | |
| RoboTwin 2.0 ( Chen et al., 2025 ) | Bimanual manipulation | – | – | – | – | – | ||
| LIBERO ( Liu et al., 2023 ) | Lifelong skill transfer | – | – | – | – | – | – | |
| CALVIN ( Mees et al., 2022 ) | Long-horizon manipulation | – | – | – | – | – | ||
| Family | Task | Hidden information | How it is revealed |
| Search | Locked Storage | target behind locks | find the tokens that unlock each compartment |
| Search Room | where the targets are | open storage places until the set is complete | |
| Blackout Search | targets in the dark | carry a light and search what it illuminates | |
| Inspect | Painted Cubes | marks on unseen faces | turn each cube to see every face |
| Marked Mugs | labels under vessels | lift or tilt each vessel to read its underside | |
| Unfamiliar Containers | how containers open | try each opening mechanism |
| GPT-6 Astra | Claude Opus 5.5 | GPT-6.1 Sol | Claude Fable 5.1 | Gemini 3.8 Flash | |||||||
| Family | Task | SR | Prog | SR | Prog | SR | Prog | SR | Prog | SR | Prog |
| Search | Locked Storage | 28.0 | 52.9 | 6.0 | 35.2 | 12.0 | 38.5 | 8.0 | 21.7 | 0.0 | 9.2 |
| Search Room | 0.0 | 31.0 | 0.0 | 26.7 | 0.0 | 19.4 | 0.0 | 19.2 | 2.0 | 17.5 | |
| Blackout Search | 0.0 | 23.7 | 0.0 | 17.4 | 0.0 | 18.5 | 0.0 | 17.1 | 0.0 | 4.4 | |
| Inspect | Painted Cubes | 30.0 | 74.8 | 14.0 | 59.2 | 18.0 | 69.2 | 14.0 | 61.5 | 2.0 | 11.3 |
| Marked Mugs | 38.0 | 69.7 | 16.0 | 51.4 | 6.0 | 37.2 | 12.0 | 31.8 | 0.0 | 0.8 | |
| Skill | Used in | GPT-6 Astra | Claude Opus 5.5 | GPT-6.1 Sol | |
| General | Pick and place (on counter) | All tasks | 95 (19) | 90 (21) | 95 (18) |
| Drive and pick | Most tasks | 100 (23) | 100 (23) | 90 (23) | |
| Task- specific | Open drawer/cabinet | Search tasks | 50 (22) | 60 (31) | 30 (36) |
| Pick from drawer/cabinet | Search tasks | 70 (15) | 60 (17) | 60 (20) | |
| Open named container | Containers | 100 (20) | 90 (16) | 95 (21) | |
| Pick from open container | Containers | 90 (20) | 90 (24) | 70 (19) |
| GPT-6 Astra | Opus 5.5 | GPT-6.1 Sol | |
| Open drawer/cabinet | |||
| Isolated | 50 | 60 | 30 |
| Full task | 90 | 78 | 75 |
| 40 | 18 | 45 | |
| Take target out | |||
| Isolated | 70 | 60 | 60 |
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Tool | Arguments | Default ticks |
| ARM | gripper-site position (required), orientation as a quaternion, and gripper command in , where opens and closes | 120 |
| BASE | body-frame forward, left, and yaw velocity, each in | 20 |
| WAIT | optional gripper command, holds the arm and base | 20 |
| STOP | none, abandons the episode, which fails | – |
| Effort per episode | Tokens per episode | Cost (USD) | ||||||
| Model | Decisions | Sim min. | Wall min. | In (M) | Out (M) | Cached (%) | $/episode | $/success |
| GPT-6 Astra | 143.7 | 10.4 | 53.8 | 5.82 | 0.02 | 91.8 | 11.29 | 48.7 |
| Claude Opus 5.5 | 143.0 | 5.4 | 32.7 | 10.70 | 0.07 | 93.5 | 6.83 | 49.5 |
| GPT-6.1 Sol | 157.4 | 11.2 | 62.7 | 6.94 | 0.02 | 92.4 | 1.99 | 16.3 |
| Claude Fable 5.1 | 148.8 | 6.1 | 54.7 | 14.55 | 0.11 | 94.2 | 19.22 | 168.6 |
| Gemini 3.8 Flash | 160.8 | 5.1 | 41.0 | 19.66 | 0.13 | 89.3 | 3.37 | 168.3 |
| Category | Hyperparameter | Value |
| Training | Optimizer | AdamW ( , ) |
| Weight Decay | ||
| Peak Learning Rate | ||
| Learning Rate Schedule | Cosine decay with 1000 warmup steps | |
| Global Batch Size | 32 | |
| Gradient Clipping Norm |
| Evidence acquisition | Evidence use | Interactive inference | Action organization | ||||||||
| Family | Task | Hidden information | DS | AP | EI | M | AD | XI | CI | CA | LP |
| Search | Locked Storage | target behind locks | – | – | – | ||||||
| Search Room | where the targets are | – | – | – | – | ||||||
| Blackout Search | targets in the dark | – | – | – | – | ||||||
| Inspect | Painted Cubes | marks on unseen faces | – | – | – | – | |||||
| Marked Mugs | labels under vessels | – | – | – | – | ||||||
| Model | Correct submission | Wrong submission | Stopped | Reply with no action | Out of decisions | Precision |
| GPT-6 Astra | 23.2 | 61.8 | 5.4 | 0.0 | 9.6 | 27.3 |
| Claude Opus 5.5 | 13.8 | 31.6 | 48.6 | 4.8 | 1.2 | 30.4 |
| GPT-6.1 Sol | 12.2 | 62.8 | 3.4 | 0.0 | 21.6 | 16.3 |
| Claude Fable 5.1 | 11.4 | 53.6 | 28.6 | 5.4 | 1.0 | 17.5 |
| Gemini 3.8 Flash | 2.0 | 55.6 | 0.8 | 0.0 | 41.6 | 3.5 |
| Skill | Reached when | Reached | Completed |
| Open drawer/cabinet | the gripper touched the handle | 85 / 85 / 75 | 50 / 60 / 30 |
| Pick from drawer/cabinet | the gripper touched the object | 90 / 80 / 100 | 70 / 60 / 60 |
| Shim the short leg | the shim came within 3 cm of the leg | 100 / 90 / 95 | 35 / 60 / 40 |
| Drive and pick | the base came within 0.8 m of the object | 100 / 100 / 95 | 100 / 100 / 90 |
| Stage | GPT-6 Astra | Opus 5.5 | GPT-6.1 Sol | Isolated test |
| Search Room and Locked Storage: target in a compartment | ||||
| Opened tried | 52/58 (90%) | 28/36 (78%) | 24/32 (75%) | Open drawer/cabinet: 50 / 60 / 30 |
| Decisions per opening (mean) | 10.4 | 25.7 | 15.2 | Open drawer/cabinet: 22.1 / 30.8 / 35.8 |
| Taken out grasp tried | 30/38 (79%) | 17/18 (94%) | 13/16 (81%) | Pick from drawer/cabinet: 70 / 60 / 60 |
| Unfamiliar Containers: item in a box | ||||
| Opened tried | 57/81 (70%) | 35/77 (45%) | 49/63 (78%) | Open named container: 100 / 90 / 95 |
| Task | Scripted information gathering | Minutes | Hours | Subtasks |
| Locked Storage | follows the token chain | 3.7 | 31.1 | 16 |
| Search Room | opens compartments nearest-first until the targets are seen | 3.0 | 25.1 | 11 |
| Blackout Search | carries the lamp along a nearest-first search | 7.5 | 62.1 | 33 |
| Painted Cubes | removes covers and turns each cube face by face | 5.1 | 42.4 | 14 |
| Marked Mugs | lifts and tilts each vessel to read its label | 3.4 | 28.0 | 15 |
| Unfamiliar Containers | visits boxes nearest-first and tries each knob’s actions until one opens the box | 6.4 | 53.0 | 29 |
| SR | Prog | |||||||
| GPT-6 Astra | Opus 5.5 | GPT-6.1 Sol | Pooled | GPT-6 Astra | Opus 5.5 | GPT-6.1 Sol | Pooled | |
| Occlusion (4 tasks) | ||||||||
| Visible | 28.8 | 16.3 | 14.0 | 19.7 | 57.6 | 45.8 | 39.5 | 47.7 |
| Look | 16.3 | 7.8 | 5.0 | 9.7 | 53.5 | 38.6 | 32.6 | 41.6 |
| Uncover | 22.3 | 7.5 | 6.1 | 12.0 | 53.1 | 35.9 | 34.1 | 41.1 |
| Look Visible | 12.5 | 8.6 | 9.0 | 10.0 | 4.1 | 7.2 | 7.0 | 6.1 |
| Missing | Wrong | Execution + side effect | All | |
| Occlusion (4 tasks) | ||||
| Visible | 0.62 | 1.24 | 0.72 | 2.58 |
| Look | 0.89 | 1.32 | 0.53 | 2.74 |
| Uncover | 0.94 | 1.28 | 0.72 | 2.94 |
| Look Visible | 0.27 | 0.08 | 0.18 | 0.16 |
| Uncover Visible | 0.32 | 0.04 | 0.00 | 0.36 |
| Task | Policy | Failed episodes | Units | Missing | Wrong | Side effect | Execution | Ran out |
| Locked Storage | GPT-6 Astra | 36 | 87 | 48 | 23 | 0 | 16 | 7 |
| Opus 5.5 | 47 | 118 | 76 | 29 | 0 | 13 | 1 | |
| GPT-6.1 Sol | 44 | 108 | 69 | 30 | 0 | 9 | 42 | |
| Search Room | GPT-6 Astra | 50 | 121 | 67 | 25 | 0 | 29 | 32 |
| Opus 5.5 | 50 | 118 | 81 | 19 | 0 | 18 | 0 | |
| GPT-6.1 Sol | 50 | 132 | 87 | 18 | 0 | 27 | 59 |
| Family | Object | GPT-6 Astra | Opus 5.5 | GPT-6.1 Sol |
| Missing evidence | Never placed | 31.9 | 33.5 | 34.2 |
| Placed before the evidence decided | 11.2 | 9.4 | 11.4 | |
| Wrong decision | Never placed | 12.8 | 19.1 | 15.9 |
| Placed against the evidence | 10.2 | 11.9 | 9.3 | |
| Both | Never placed | 44.6 | 52.6 | 50.1 |
| Placed wrongly | 21.4 | 21.3 | 20.7 |