Research
RobotWorld: Benchmarking Multimodal Agents for Robot Use Across Diverse Tasks and Embodiments
Overview Research area: Robotics and embodied AI — specifically, benchmarking general-purpose multimodal (vision-language) agents on physical robot tasks across many different robot bodies. Technical

- arXiv
- 2610.10409
- Published
- 2026-10-07
- Authors
- Zhiqin Yang, Chenxin Li, Xiaomeng Hu, Yibin Liu, Weidong Huang, Jiankai Sun, Haitao Li, Zijian Wu, Yuzhi Huang, Fanding Huang, Hanwen Sun, Jiashun Liu, Jingqi Tong, Mingxin Huang, Shaoli Hu, Shijue Huang, Tianyi Bai, Xinyuan Wang, Yunlong Lin, Zhengyang Tang, Zhexin Zhang, Zhuo Chen, Xierui Song, Juntao Dai, Boyuan Chen, Jiaming Ji, Fangneng Zhan, Mengkang Hu, Wei Xue, Yonggang Zhang, Han Hu, Tsung-Yi Ho, Yike Guo
AI summary
Overview
Research area: Robotics and embodied AI — specifically, benchmarking general-purpose multimodal (vision-language) agents on physical robot tasks across many different robot bodies.
Technical level: Intermediate. The paper is written in largely accessible prose, but it assumes familiarity with simulation benchmarks, robot control interfaces, and agent evaluation protocols.
Scope: RobotWorld is a simulation benchmark of 84 robot tasks drawn from 20 source projects that measures whether general-purpose multimodal agents can turn instructions and observations into successful physical action.
What This Paper Is About
General-purpose AI agents have become competent at digital work such as writing code, driving browsers, and completing multi-step tasks with tools. This paper asks how far those same capabilities carry into the physical world — whether an agent that can write a program and inspect an image can also make a robot reliably grasp, insert, balance, park, or fly. The authors build RobotWorld, a simulation testbed spanning manipulation, mobile manipulation, locomotion, driving, and aerial control, and use it to measure both what current agents can do and exactly where their execution breaks down.
Key Contributions
-
A challenging proving ground for "robot use." RobotWorld provides 84 task specifications from 20 source projects, spanning five domains (38 manipulation, 20 mobile manipulation, 11 locomotion, 11 driving, 4 aerial tasks). Each task has documented robot interfaces, explicit interaction budgets, and executable success checks rather than subjective scoring.
-
An interface-inclusive evaluation design. Unlike prior benchmarks that support only one control style, RobotWorld supports direct model-emitted actions, generated code control, and mixed calls within an episode, and covers navigation, balance, and driving/flight domains that the compared benchmarks (CaP-X, EmbodiedBench, VLABench, EmbodiedEval, Embodied Agent Interface, EmbodiedSWE-Bench) report only partially or not at all.
-
An empirical capability profile drawn from execution traces. By analysing trajectories rather than just aggregate scores, the authors identify which perception and control workflows agents spontaneously construct (colour segmentation, camera-geometry fitting, spatial estimation, numerical dynamics computation) and which failure mechanisms prevent those workflows from composing into task completion.
-
Cross-model comparison on shared tasks. The paper contrasts model behaviour on identical tasks, showing that different models succeed on different task types and that interaction cost and physical execution cost do not track each other.
Main Findings
-
Task completion is low across the board. Astra completes 16 of 84 tasks (19.0%), Opus 5.5 completes 13 (15.5%), Kimi K3 completes 2 (2.4%), and DeepSeek V4.1 Flash and Gemini 3.8 Flash complete 1 each (1.2%). Even the highest-scoring model leaves 68 tasks unfinished.
-
Models solve partly different tasks. Astra and Opus 5.5 share eight successes; Astra solves eight additional tasks and Opus 5.5 solves five more. Their combined coverage is 21/84 tasks (25.0%), leaving 63 tasks unsolved by any evaluated model. The successes of Kimi, DeepSeek, and Gemini all fall inside this shared set.
-
Splitting mobile manipulation widens the leading gap. Astra solves 9/38 manipulation tasks (23.7%) and 4/20 mobile manipulation tasks (20.0%); Opus 5.5 solves 7/38 (18.4%) but none of the mobile manipulation tasks. Kimi K3 completes one mobile manipulation task (1/20, 5.0%). Pooling the two manipulation domains gives 13/58 for Astra and 7/58 for Opus 5.5.
-
Dynamic control favours Opus 5.5. Opus 5.5 completes 3/4 aerial tasks (75.0%) versus 1/4 each for Astra and DeepSeek, and records the only locomotion success (quadruped push recovery, 1/11, 9.1%). The authors note the aerial set is small: one task changes the rate by 25 percentage points.
-
Driving and locomotion remain broadly unsolved. Astra and Opus 5.5 each complete courtyard and hairpin driving (2/11, 18.2%) and Kimi completes courtyard driving (1/11, 9.1%). No evaluated model completes the remaining nine driving tasks or ten of the eleven locomotion tasks.
-
Different goal types favour different models. Astra succeeds more often on spatial and constrained-contact goals, while Opus 5.5 succeeds more often on continuous-balance and timed-interaction goals.
-
Tool calls and control steps diverge. Among the eight tasks solved by both models, Astra uses fewer tool calls in seven but fewer control steps in only six — showing interaction cost and physical execution cost need not agree.
-
Agents build sophisticated but non-composing workflows. Traces show agents performing image segmentation, camera calibration, spatial estimation, and dynamics-based computation, and using feedback to revise motions and recover from errors. These capabilities do not consistently produce success: agents lose task-relevant object states despite reaching commanded poses, fail to correct ineffective actions, recover too late, or mistake an unfinished task for a completed one.
-
Success criteria are physical, not behavioural. Checkers test position and orientation, object state, contact and motion, persistence (e.g., payload hovering requires all stability bounds to hold together for at least the final two seconds, with a broken hold resetting its duration), events and order, and path/safety constraints. A favourable final image, a finished tool call, or the agent's own success claim are each insufficient.
Methodology in Plain English
The authors collected robot tasks from 20 existing simulation projects (including BEHAVIOR-1K, RoboCasa, RoboLab, RoboDojo, Bench2Dex, WheeledLab, OmniDrones, and VolleyBots) and wrapped each one in a uniform contract: a plain-language goal, a defined initial state, a documented set of permitted observations and commands, a bounded number of interaction steps, and an automatic checker that decides success from simulated states and events.
Agents then interact through a shared loop. The agent sees camera images, robot state, and execution feedback; it may pause physics to run auxiliary computation in an isolated workspace (image inspection, geometric estimation, numerical calculation); it then issues a bounded robot command. The environment validates the request, advances physics, and returns completed-step counts and errors so the agent can distinguish a rejected command from one that ran but failed. Across 30 September to 6 October 2026, the authors ran GPT-6 Astra, Claude Opus 5.5, Kimi K3, DeepSeek V4.1 Flash, and Gemini 3.8 Flash on all 84 tasks, one episode per model–task pair, for 420 scored episodes. Code control was disabled; the observation history contained the current observation plus up to four earlier observations sampled every two observation rounds; and per-task event budgets (15 consecutive and 30, 60, or 120 total events on 83 tasks, with volleyball 1v1 at 20 and 150) were calibrated from Astra interaction records. The authors also validated checkers with positive, negative, and boundary fixtures, while acknowledging that checker validation does not prove a task is solvable under every initial condition.
Why This Matters
Impact on research. RobotWorld reframes robot benchmarking away from "did the final image look right" toward auditable physical evidence — contact, persistence, event ordering, and path history. By publishing execution traces alongside scores, it turns a leaderboard into a diagnostic tool that points at specific missing capabilities rather than a single number.
Real-world applications (domains the benchmark covers):
- Warehouse and household manipulation, where insertion, articulated-object operation, and multi-stage assembly must succeed reliably rather than approximately.
- Mobile manipulation in home environments, drawing on BEHAVIOR-1K and RoboCasa scenes such as kitchen navigation and dishwasher loading.
- Ground vehicle autonomy, including parking in restricted spaces and passage through a moving gate.
- Aerial and legged operation, including payload stabilisation, interception of moving targets, and quadruped balance recovery under pushes.
Industry relevance. The finding that the best model finishes under a fifth of the suite sets a concrete headroom figure for teams building robot-use products. The specific failure modes — losing object state after reaching a commanded pose, failing to correct ineffective actions, recovering too late, and mistaking unfinished tasks for complete ones — map directly onto engineering requirements for state estimation, closed-loop correction, temporal planning, and completion verification. The observation that tool-call efficiency and control-step efficiency diverge is also directly relevant to cost planning for deployed systems.
Future Directions
-
Ground perception and computation in action. The paper explicitly names this as a training priority: agents can segment, calibrate, and compute, but these outputs must be tied to whether the physical goal was achieved.
-
Improve error recovery and temporal coordination. Agents need to detect that an action had no effect, correct it before the budget is spent, and coordinate over the duration a task actually requires rather than satisfying a condition momentarily.
-
Build reliable completion assessment. Agents must stop treating a finished tool call or their own success claim as evidence of success, and instead verify against the physical conditions the task defines.
-
Strengthen feasibility and reference coverage. The paper notes that reference coverage is incomplete and that checker validation does not establish a solution for every task or robustness across initial conditions — an open gap for the community. A further open question is raised by the paper's own limits: results come from one retained episode per model–task pair, and the authors state that these are task-level observations rather than estimates of repeated-run reliability.
Target Audience
Researchers and engineers working on embodied agents, multimodal foundation models, and robot learning who need a rigorous way to measure whether general-purpose models can act in the physical world. It is also useful for benchmark designers interested in executable success checkers and interaction budgets, and for practitioners planning robot-use systems who want a documented account of current failure modes and their causes. Readers looking for a beginner-level introduction to robotics will find the task descriptions accessible, but the evaluation protocol and interface details assume some background in simulation and robot control.
Authors’ abstract
General-purpose agents increasingly write code, use tools, and complete complex digital tasks, raising the question of how far these capabilities carry into the physical world. To investigate this, we introduce RobotWorld, a challenging simulation testbed for robot use: turning instructions and observations into physical task execution through robot interfaces. Its 84 tasks span manipulation, mobile manipulation, locomotion, driving, and aerial control, with explicit interaction budgets and executable success checks. By analysing task outcomes alongside execution traces, we identify both the capabilities that transfer and the gaps that prevent reliable completion. Furthermore, we find that current agents can construct sophisticated perception and control workflows, including image segmentation, camera calibration, spatial estimation, and dynamics-based computation. These capabilities, however, do not consistently compose into successful behaviour: agents lose task-relevant object states despite reaching commanded poses, fail to correct ineffective actions, recover too late, or mistake unfinished tasks for completion. This uneven transfer also differs across models: Astra succeeds more often on spatial and constrained-contact goals, whereas Opus 5.5 succeeds more often on continuous-balance and timed-interaction goals. By linking these outcomes to execution behaviour, RobotWorld provides both a rigorous proving ground and an empirical account of the remaining capability gaps, thereby establishing concrete targets for training and designing more reliable physical-world agents.