Skip to content
AI.info

Research

RoboQuest: Generalist Physical Agents that Search, Inspect and Test

Overview Research area: Robotics — embodied AI benchmarks and evaluation of generalist robot policies (mobile manipulation, active perception, vision-language-action models). Technical level: Advanced

RoboQuest: Generalist Physical Agents that Search, Inspect and Test
arXiv
2610.10388
Published
2026-10-07
Authors
Liu Renhang, Navonil Majumder, Tej Deep Pala, Soujanya Poria

AI summary

Overview

Research area: Robotics — embodied AI benchmarks and evaluation of generalist robot policies (mobile manipulation, active perception, vision-language-action models).

Technical level: Advanced. The paper assumes familiarity with manipulation benchmarks, vision-language-action (VLA) policies, and multimodal agent harnesses.

Scope: RoboQuest is a ten-task simulated benchmark for goal-directed embodied exploration, measuring whether agents can physically search, inspect, and test an unfamiliar kitchen to acquire information that is absent from their observations, then act on it.

What This Paper Is About

Most manipulation benchmarks make all task-relevant information available in the current observation, so success mainly tests execution. RoboQuest asks what happens when the information an agent needs is not observable and must be obtained through physical interaction — where an object is, what an unseen face or underside shows, or how an unfamiliar mechanism behaves. The goal is to evaluate generalist agents on the full loop of acquiring evidence, using it to adapt later actions, and deciding autonomously when to commit by pressing a SUBMIT button.

Key Contributions

  1. RoboQuest, a benchmark for goal-directed embodied exploration. Ten mobile manipulation tasks organized into three families — Search, Inspect, and Test — covering three forms of uncertainty (location, hidden object properties, and interaction outcomes). Each task specifies only an externally verifiable physical goal, never the investigation strategy.

  2. A demonstration dataset. 5,000 successful episodes (500 per task, 366 hours of interaction at 20 Hz) generated by scripted oracles whose exploration is deliberately scripted as that of an agent without access to the private specification. Every frame carries a two-level language annotation (stage and subtask), written from the robot's viewpoint.

  3. A systematic evaluation of generalist and trained agents. Five frontier multimodal agents (GPT-6 Astra, Claude Opus 5.5, GPT-6.1 Sol, Claude Fable 5.1 with medium thinking, and Gemini 3.8 Flash with high thinking) run zero-shot through a common visuomotor interface, plus a supervised fine-tuned π_0.5 policy trained on the released demonstrations.

  4. Isolated skill tests and a detailed failure analysis. The three highest-scoring models are tested on the execution skills the tasks are built from, with hidden information supplied up front, and every failure is attributed to one of four families: missing evidence, wrong decision, side effect, or execution failure.

Main Findings

  • Frontier agents score low overall. GPT-6 Astra has the best overall success rate (SR) at 23.2% and the best mean progress (48.0%), and leads on 7 of the 10 tasks. Claude Opus 5.5 (13.8% SR, 37.3% progress), GPT-6.1 Sol (12.2%, 34.0%), and Claude Fable 5.1 (11.4%, 31.0%) cluster closer together, while Gemini 3.8 Flash trails at 2.0% SR and 9.2% progress.

  • Open-ended search is unsolved. No model solves a single Blackout Search episode. The only Search Room success across all five models is one episode of Gemini 3.8 Flash.

  • Puzzle Box is an outlier. It is the easiest task for every model, 44 to 58 points above the next-best task for the four stronger models (GPT-6 Astra 96.0% SR, 99.6% progress). Excluding it, GPT-6 Astra succeeds about 1.9 to 3.1 times as often as Opus 5.5, GPT-6.1 Sol, and Fable 5.1.

  • One task favors a different model. Wobbly Stand is the only task where another model clearly outperforms GPT-6 Astra: Opus 5.5 leads on both SR (24.0%) and progress (32.2%).

  • Execution skills are largely present in isolation. Averaged over all isolated skills (20 scenes each), the three top models score 80% (GPT-6 Astra), 81% (Opus 5.5), and 72% (GPT-6.1 Sol). General skills are strong: GPT-6 Astra and Opus 5.5 drive to an object and pick it up in every episode, GPT-6.1 Sol in 90%, and the three place an object on a counter target in 90–95% of trials. Task-specific skills are uneven: flipping a cube reaches 100% for all three, and opening a named container 90–100%, but opening a drawer or cabinet succeeds in only 30–60% of episodes, retrieving from one in 60–70%, and shimming a table leg in 35–60%.

  • Isolation is not the bottleneck. Opus 5.5 (81%) slightly exceeds GPT-6 Astra (80%) in isolated skills despite GPT-6 Astra succeeding 1.7 times as often on the full benchmark.

  • Skill performance is often better inside full tasks. GPT-6 Astra opens 90% of the compartments it tries in full tasks against 50% in isolation, using 10 decisions on average instead of 22. Across compartments holding a target in Search Room and Locked Storage, models open 75–90% of those they attempt (versus 30–60% isolated), using 10 / 26 / 15 decisions on average instead of 22 / 31 / 36.

  • Missing evidence dominates failures. Across 931 / 1143 / 1182 units from 384 / 431 / 439 failed episodes of GPT-6 Astra / Opus 5.5 / GPT-6.1 Sol, missing evidence is the largest family (43 / 43 / 46%), followed by wrong decisions (23 / 31 / 25%). Together they account for 66–74% of failures. Execution failures are 21 / 19 / 18% and side effects 13 / 7 / 12%.

  • Agents stop exploring too early. In about half of failures (45 / 53 / 50%), an object the task needed was never placed. In most of these cases (27–29% of all failures), the object stayed hidden inside a drawer, cabinet, box, or cover that was never opened. The hiding place itself was rarely missed — it typically appeared in the agent's observations five or more times — while hiding places never observed account for only 1–2% of failures.

  • Premature commitment. Roughly a tenth of all failures (9–11%) came from committing objects to locations before observing the evidence needed to decide. In Odd Parcel, 71–84% of wrongly selected parcels were boxed before weighings could reveal whether they were odd. In Stamp Composition, 20–23% of incorrectly placed stamps by GPT-6 Astra and GPT-6.1 Sol were printed without a prior test print, compared to 3% for Opus 5.5.

  • Occlusion hurts only through missed discovery. When task objects start out of view or under a cover, overall success drops by 8–10 points, and these extra failures fall entirely under missing evidence; wrong decisions, execution failures, and side effects do not increase.

  • Disturbances are rarely prevented or repaired. Roughly two-thirds of side-effect errors involve objects falling to the floor. An empty bowl was provided in each scene to safely hold balls, but it was used in only 2–8 of the 50 Marked Mugs episodes and never in Wobbly Stand. Only 4 of 38 subsequent recovery attempts for a loose ball succeeded.

  • Trial-and-error learning is hard. On Puzzle Box with chain lengths 4 and 5, hiding the bolt dependencies barely affects GPT-6 Astra (94.4% SR without the cover, 96.4% with it, at 84.7 versus 119.5 mean decisions), while Opus 5.5 drops from 94.4% to 36.8% and GPT-6.1 Sol falls from 72.2% to 60.1%. Roughly half of Opus 5.5's failed attempts pull a blocked bolt three or more times without updating the plan (12 of 26).

  • The fine-tuned VLA policy almost never succeeds. π_0.5 reaches 2.0% success with 43.3% average progress on Puzzle Box, and 0.0% success with near-zero progress on every other task, with most episodes failing by timeout. The paper attributes this to out-of-distribution generalization: evaluation holds out kitchen styles, materials, mechanisms, and target configurations, and the memoryless architecture compounds small errors.

  • Cost varies widely. Evaluating the five models on 500 episodes each cost $21,352 in total, from $997 for GPT-6.1 Sol to $9,610 for Fable 5.1. GPT-6.1 Sol is cheapest both per episode and per successful episode; GPT-6 Astra and Opus 5.5 cost about three times as much per success, and Fable 5.1 and Gemini 3.8 Flash are the least economical per success.

  • Giving up takes different forms. GPT-6 Astra and GPT-6.1 Sol almost always press SUBMIT, but only 27% and 16% of those submissions are correct. Opus 5.5 is the most precise when it does submit (30%) and often stops on its own, with a median of 24 decisions left. Gemini 3.8 Flash and GPT-6.1 Sol run out of the 200 decisions far more often. No model stops because it is done, and no stopped episode had its goal state met.

Methodology in Plain English

The researchers built ten tasks in RoboCasa365 kitchens simulated in MuJoCo, using a Franka

Authors’ abstract

Recent advances in multimodal foundation models have made them capable generalist physical agents for a range of manipulation tasks. However, successful operation in an unfamiliar environment may require an agent to seek task-relevant information through interaction when it is absent from the observations: it may need to determine where a relevant object is, inspect an unobserved property, or discover the effect of an unfamiliar tool. We thus introduce RoboQuest, a benchmark for goal-directed embodied exploration, where agents must actively acquire task-relevant information through physical interaction, use the resulting evidence to adapt subsequent actions, and autonomously decide when to commit to task completion. RoboQuest comprises ten mobile manipulation tasks centered on three forms of uncertainty: search, manipulation-based inspection, and interactive testing. We evaluate five frontier multimodal agents through a common visuomotor interface, as well as a $π_{0.5}$ policy fine-tuned on the full-episode demonstrations we release. The best agent succeeds in only 23\% of the episodes, and the fine-tuned policy almost never succeeds. Isolated tests of the execution skills the tasks are built from, with the hidden information supplied, show that the agents can carry out most of the required actions, and our failure analysis attributes only a minority of the failures to execution. Our failure analysis further finds that the agents often stop exploring too early as they make decisions before observing the required evidence for task completion. We also find that agents rarely prevent or repair the disturbances caused by their exploration. Moreover, learning by trial and error remains difficult for most models.

Read the original paper