Research
RoboCasa365: A Large-Scale Simulation Framework for Training and Benchmarking Generalist Robots
Overview Research area: Robotics — simulation benchmarks and policy learning for generalist (multi-task, mobile-manipulation) robots, specifically household kitchen manipulation. Technical level: Inte
- arXiv
- 2603.04356
- Published
- 2026-03-04
- Authors
- Soroush Nasiriany, Sepehr Nasiriany, Abhiram Maddukuri, Yuke Zhu
AI summary
Overview
Research area: Robotics — simulation benchmarks and policy learning for generalist (multi-task, mobile-manipulation) robots, specifically household kitchen manipulation.
Technical level: Intermediate. The benchmark design and the six experiment tables are readable by a general machine-learning audience, but interpreting the results assumes familiarity with imitation learning, vision-language-action models, and the train-then-fine-tune paradigm in robotics.
Scope: The paper introduces RoboCasa365, a simulation framework that packages 365 kitchen tasks, 2,500 kitchen scenes, and over 2,000 hours of robot interaction data together with a suite of benchmarks for multi-task training, foundation model training, and lifelong learning, then uses it to study what drives generalization in generalist robot policies.
What This Paper Is About
Progress toward generalist robots is hard to measure because the field lacks a reproducible, large-scale benchmark: real-world evaluation is expensive, noisy, and slow, while existing simulators cover only narrow sets of tasks and rooms. The authors scale up the RoboCasa simulation platform into RoboCasa365 — many more tasks, scenes, assets, and demonstrations — and then run systematic experiments on it to isolate how task diversity, dataset scale, and environment variation affect policy performance. The goal is both to provide a shared testbed and to extract concrete findings about what makes generalist robot learning work.
Key Contributions
-
A much larger task suite and environment set. RoboCasa365 defines 365 everyday tasks (65 atomic and 300 composite) spanning 60 kitchen activities, evaluated across 2,500 pretraining kitchen scenes plus 10 target kitchen scenes. The pretraining scenes are built as 50 layouts × 50 styles, with layouts modeled after 50 real homes listed on Zillow.com across U.S. locations.
-
Expanded simulation assets. The object library adds 57 new object categories on top of the existing 2,509 objects across 153 categories, and the inventory of interactable fixtures and appliances grows from 20 instances across 4 categories in RoboCasa to 456 instances across 12 categories (including new categories such as toasters, toaster ovens, stand mixers, blenders, and electric kettles, with fridges, ovens, and dishwashers now articulated).
-
Large-scale datasets with defined pretraining and target splits. The pretraining data covers 300 tasks (30k human demonstrations at 100 per task, plus MimicGen synthetic data for 60 atomic tasks at 10k demonstrations each, a 100× scale-up). The target data covers 50 tasks split into Atomic (18), Composite-Seen (16), and Composite-Unseen (16), with 500 human demonstrations per task for 25k demonstrations total.
-
A systematic benchmark with three learning settings. The paper defines and runs evaluations for massively multi-task training, foundation model (pretrain-then-fine-tune) training, and lifelong learning, plus a pretraining data composition study and a real-world sim-and-real transfer experiment.
Main Findings
-
Multi-task scale is hard, and horizon length dominates difficulty. Training on 300 pretraining human datasets (30k demonstrations), GR00T N1.5 achieved the best average success rate at 20.0%, followed by π0.5 (16.9%), π0 (15.0%), and Diffusion Policy (6.1%). Every method scored best on Atomic tasks (e.g., GR00T N1.5 at 43.0%), far worse on Composite-Seen (9.6%), and worst on the zero-shot Composite-Unseen tasks (4.4%).
-
High-capacity vision-language-action models outperform Diffusion Policy. Diffusion Policy reached an average of 6.1% and scored 0.2% on Composite-Seen and 1.25% on Composite-Unseen, while π0, π0.5, and GR00T N1.5 all showed non-zero success on unseen tasks — which the authors treat as a sign of stronger generalization. The authors explicitly state they do not claim GR00T N1.5 is conclusively superior, since compute, data composition, and whether backbones are fine-tuned all influence results.
-
Pretraining yields roughly a 3× data-efficiency gain. Averaged across task types, target-only training reached 21.0% / 34.3% / 43.7% at 10% / 30% / 100% of target data, while pretraining plus post-training reached 35.9% / 42.2% / 51.1%. Pretraining alone (no target fine-tuning) averaged 15.1%, with 41.9% on atomic tasks but 0.0% on Composite-Seen and 0.2% on Composite-Unseen.
-
Pretraining helps most on unseen tasks. With pretraining plus post-training at 100% target data, Composite-Unseen reached 42.1% versus 33.3% for target-only training, and Composite-Seen reached 40.6% versus 35.0%.
-
More pretraining task diversity helps most in low-data regimes. At 10% target data, averages rose from 21.0% (no pretraining) to 34.7% (Human50) to 40.0% (Human300). At 100% target data the same progression was 43.7% → 50.0% → 52.5%. The largest gain from adding task diversity appeared on Composite-Unseen tasks (32.3% for Human300 versus 23.8% for Human50 at 10% target data).
-
Adding MimicGen synthetic data did not improve downstream results. The full mixture (Human300 + MG60) scored lower on average than human data alone (Human300) in both the 10% regime (35.9% vs 40.0%) and the 100% regime (51.1% vs 52.5%). The authors attribute this to varying quality in the synthetic demonstrations and flag better use of mixed-quality data as future work.
-
Lifelong learning shows catastrophic forgetting. Training across four phases of progressively longer-horizon tasks, atomic performance fell from 41.5% after Phase 1 to 10.6% by Phase 4. The 2–3 stage tasks fell from 24.5% (Phase 2) to 1.7% (Phase 4), 4–5 stage tasks from 11.3% (Phase 3) to 2.7% (Phase 4), and the newly learned 6+ stage tasks reached only 4.3% in Phase 4.
-
Simulation data improves real-world policies. On four real kitchen tasks with a DROID Panda arm and three cameras, a sim-and-real recipe reached 79.8% average success versus 61.8% for real-data-only training — an 18.1% average improvement — with the largest single gain on PickPlaceCounterToCabinet (84% vs 52%).
-
Datasets skew long-tailed in horizon. Most of the 365 tasks require one or two subtasks, but a few require 15 or more; episode lengths across the 55k human pretraining and target episodes mostly fall between 10 and 60 seconds, with a long tail past 3 minutes.
Methodology in Plain English
The authors start from the existing RoboCasa simulator and scale it up on every axis. They expand the 3D object library, add many more articulated appliances and fixtures, and build thousands of kitchen scenes by independently combining floor-plan "layouts" (50 of them, digitized from real homes) with decorative/asset "styles" (50 of them), while keeping the styles used at pretraining time separate from those used at evaluation time so that styles do not overlap.
Tasks are defined in two tiers. Atomic tasks exercise a single skill out of eight foundational skills the platform supports (pick-and-place, opening and closing doors, opening and closing drawers, turning levers, turning knobs, pressing buttons, insertion, navigation). Composite tasks chain multiple skills and were generated with an LLM pipeline: prompt for a list of kitchen activities (top 60), then for each activity prompt for task blueprints (name, description, objects and fixtures, skill sequence), then implement each blueprint in code.
Data is collected by teleoperating a Franka Panda Emika robot on an Omron mobile base, plus MimicGen for synthetic trajectories seeded from human demonstrations. Evaluation is organized as three benchmark suites. In multi-task training, four published methods are trained on the 300-task human pretraining mixture and evaluated on 50 held-out target tasks. In foundation model training, a model is pretrained on all pretraining data and then fine-tuned separately on each target split at 10%, 30%, and 100% of target data, with ablation variants that swap the pretraining mixture (no pretraining, Human50, Human300, Human300 + MG60). In lifelong learning, the model is fine-tuned sequentially through four phases and re-evaluated on all previously seen tasks. A final experiment transfers to a real robot by mid-training on the 150 highest-performing simulation tasks and then co-fine-tuning on real demonstrations plus matching re-rendered simulation data.
Why This Matters
Impact on research. The paper provides a shared, reproducible testbed that lets different policy-learning methods be compared on identical task distributions, and its results frame a concrete agenda: long-horizon composite tasks remain largely unsolved (single-digit to low-double-digit success rates), catastrophic forgetting is severe, and synthetic data quality is a bottleneck rather than free scale. It also offers open-source models as comparison points for the community.
Real-world applications:
- Household service robots that operate across many kitchen activities — cooking, cleaning, packing, storing — rather than mastering one task.
- Data-efficient adaptation of robot foundation models to a new home, where pretraining cuts the number of target demonstrations needed by roughly 3×.
- Sim-to-real pipelines for appliance interaction, such as operating kettles, toaster ovens, dishwashers, and cabinets, where simulation data measurably improves real-robot success.
- Long-lived deployed robots that must acquire new skills on site without regressing on old ones — the exact failure mode the lifelong learning benchmark measures.
Industry relevance. The gains from pretraining and from sim-and-real co-training directly affect the cost of collecting robot data, a major expense for robotics companies. The finding that a large synthetic dataset (MimicGen) can reduce downstream performance is an important operational signal for teams deciding where to spend data-generation budgets. The comparison of Diffusion Policy against π0, π0.5, and GR00T N1.5 also gives practitioners a reference point for choosing policy architectures at scale.
Future Directions
- Escaping the kitchen. The authors note the benchmark is currently limited to kitchen environments and question how well the findings transfer to other household settings or broader domains.
- Closing the simulation-to-reality gap. The paper states that its dataset does not capture the full sensory and physical complexity of the real world and treats bridging that gap as a significant open challenge.
- Better use of large, mixed-quality synthetic data. Since MimicGen data added scale without helping downstream performance, developing methods that exploit such datasets effectively is called out as important future work.
- Solving long-horizon and continual learning. Both Composite-Seen/Unseen success rates and the phase-by-phase forgetting curves leave substantial headroom, and the paper positions the benchmark as a testbed for improving on these results.
Target Audience
Researchers and engineers working on robot learning, imitation learning, and robot foundation models who need a large-scale simulation benchmark for evaluating generalist policies; practitioners building sim-to-real pipelines for mobile manipulation; and teams deciding how to allocate data collection and synthetic data generation budgets. Readers looking for a closed-form algorithmic contribution will not find one here — the value lies in the benchmark, the datasets, the asset inventory, and the comparative study.
Note on reported figures: the paper states "over 600 hours" of human demonstration data in the abstract and lists 612 hours total human data (404 hours pretraining plus 208 hours target) in its dataset statistics, while the foundation-model experiment section cites 411 hours for the human pretraining data. Both values are reproduced here as reported in their respective sections.
Authors’ abstract
Recent advances in robot learning have accelerated progress toward generalist robots that can perform everyday tasks in human environments. Yet it remains difficult to gauge how close we are to this vision. The field lacks a reproducible, large-scale benchmark for systematic evaluation. To fill this gap, we present RoboCasa365, a comprehensive simulation benchmark for household mobile manipulation. Built on the RoboCasa platform, RoboCasa365 introduces 365 everyday tasks across 2,500 diverse kitchen environments, with over 600 hours of human demonstration data and over 1600 hours of synthetically generated demonstration data -- making it one of the most diverse and large-scale resources for studying generalist policies. RoboCasa365 is designed to support systematic evaluations for different problem settings, including multi-task learning, robot foundation model training, and lifelong learning. We conduct extensive experiments on this benchmark with state-of-the-art methods and analyze the impacts of task diversity, dataset scale, and environment variation on generalization. Our results provide new insights into what factors most strongly affect the performance of generalist robots and inform strategies for future progress in the field.