Future Horizons
Why Robots Can't Learn Like Chatbots Did: The Embodied AI Data Bottleneck
Humanoid hype hides the real constraint: no internet-scale, action-paired corpus exists for robots. Simulation, VR teleoperation and egocentric video are 2026's three fixes, and the visual sim-to-real gap still breaks policies.

Gabriele Masetti ·
Why Aren't Robots Smarter Yet? The Humanoid Hype Meets a Data Wall
Humanoid robots dominate keynote stages in 2026. Figure's third-generation machines sort parts inside a BMW assembly hall in Spartanburg, 1X sold out a year of $20,000 Neo home robots in five days, Tesla has torn out the Fremont line that built the Model S so it can build Optimus on it instead, and venture money keeps arriving at rates that would have seemed absurd two years ago. Ask any engineer inside these companies why a robot still fumbles a wine glass or drops a fabric napkin, and the answer rarely mentions chips or actuators. It comes down to data.
Large language models learned to write, reason, and code by ingesting a text corpus assembled over three decades of the open internet — blogs, books, forums, code repositories. Robots have no equivalent corpus. Ken Goldberg, a UC Berkeley robotics researcher, put the gap in blunt terms: the text used to train today's large language models would take a human roughly 100,000 years to read.
"We don't have anywhere near that amount of data to train robots, and 100,000 years is just the amount of text that we have to train language models," he said. Nobody knows exactly how much physical interaction data a general-purpose robot needs, but every serious estimate places it far above anything collected so far.
The 100,000-Year Data Gap Behind Embodied AI Training Data
Text scales because copying and transmitting it is nearly free. A blog post written once gets read by millions of people. Robot data does not work that way. Every demonstration requires a physical robot, a physical environment, and — in almost every deployed system today — a human either piloting the arm directly or supervising it closely enough to intervene. Goldberg's second, more painful observation explains the ceiling directly: "every eight hours of work gives you just eight more hours of data."
A model like GPT was not trained by having one person type for eight hours; it drew on a corpus built by billions of people over decades. A teleoperated robot arm produces training examples at exactly the rate a human can move it.
Bessemer Venture Partners quantified the resulting scarcity in an April 2026 report: total global robot manipulation data adds up to roughly 300,000 hours, against on the order of a billion hours of internet video available for language and video pretraining. Open X-Embodiment, the most widely used cross-embodiment dataset, pooled by Google DeepMind and more than 30 partner institutions, holds just over a million real robot trajectories spanning 22 platforms — a meaningful research resource, but a rounding error next to the corpus that trained a large language model.
DROID, a multi-lab academic effort spanning 13 institutions on three continents, added roughly 76,000 trajectories and about 350 hours of teleoperated interaction across 564 scenes and 86 tasks. Useful, diverse, and still tiny next to what text-based AI had to work with.
| Dataset | Trajectories | Scope |
|---|---|---|
| Open X-Embodiment | 1M+ | 22 platforms, 30+ institutions |
| DROID | ~76,000 | 13 institutions, 564 scenes, 86 tasks |
| NVIDIA GR00T-Mimic (synthetic) | 780,000 | Generated in 11 hours from real demonstrations |
| Data source | Scale |
|---|---|
| Text used to train today's LLMs (Goldberg estimate) | ~100,000 years of human reading |
| Global robot manipulation data (Bessemer, 2026) | ~300,000 hours |
| Internet video available for pretraining | ~1 billion hours |
World Models for Robots: Simulation Becomes the First Fix
The first production answer is to stop collecting exclusively in the real world and generate synthetic experience a robot can learn from before it ever touches hardware. NVIDIA's Isaac GR00T platform pairs a vision-language-action foundation model, GR00T N1, with a synthetic-data pipeline called GR00T-Mimic that multiplies a small number of captured human demonstrations into much larger simulated datasets.
NVIDIA has reported generating 780,000 synthetic trajectories — the equivalent of roughly 6,500 hours of human demonstration, or nine months of continuous recording — in just 11 hours of compute. Training the company's Cosmos world-foundation models required a different kind of resource: roughly 10,000 H100 GPUs running for about three months.
NVIDIA is not alone in betting on video-trained world models rather than hand-coded physics. Google DeepMind's Genie 3 and Dreamer 4 generate interactive environments a policy can practice inside before deployment, and autonomous-driving specialist Wayve has applied the same idea to its own domain with an 8.4-billion-parameter model called GAIA-2. Serving a frontier world model is not cheap either — running Google's Genie 3 has been estimated at roughly $100 an hour of generated environment — which is part of why the money chasing this approach has scaled so fast.
Money has followed the bet that world models are the way to escape the data ceiling. Yann LeCun's new venture, AMI Labs, raised $1.03 billion in seed funding in March 2026 at a $3.5 billion pre-money valuation, explicitly to build world models rather than another large language model.
Fei-Fei Li's World Labs closed a $1 billion round on February 18, 2026, bringing total funding to about $1.23 billion, backed by Autodesk, AMD, NVIDIA, Fidelity, and Sea, to build spatial-intelligence models under its Marble product line. Bessemer estimated that roughly $6 billion flowed into six or seven world-model companies in the first quarter of 2026 alone, and separately projected that aggregate robotic data costs across the industry will top $3 billion over the following two years.
VR Teleoperation and the Race to Capture Dexterity
Simulation handles locomotion — walking, balancing, recovering from a stumble — reasonably well, because physics engines model rigid-body dynamics accurately. It handles dexterous manipulation far less reliably, so companies still lean on humans piloting robots directly, increasingly through consumer VR hardware rather than industrial rigs. Apple Vision Pro and Meta Quest 3 headsets now show up in teleoperation stacks because they track six-degree-of-freedom hand motion out of the box, though every system still needs a retargeting layer to map a human hand's 27 joints onto a gripper or a 12-to-16-degree-of-freedom dexterous hand.
Apple's own research group built ARMADA, a system that overlays a virtual robot onto the real world through Vision Pro so a person can demonstrate a task barehanded and still produce data compatible with a physical robot's limits, which matters because it removes the requirement to own a robot at all.
Older, cheaper rigs remain in wide use alongside the headsets: ALOHA 2, a leader-follower bimanual arm setup that costs roughly $15,000 to $20,000 to build, produces somewhere between 10 and 30 usable demonstration episodes per hour of human effort, and UMI, a wrist-mounted capture device, extends collection into ordinary kitchens and workshops without a robot present at all.
Figure AI opened a dedicated Helix Lab to run its VR and teleoperation pipeline at scale for its Helix vision-language-action model, and its job postings now describe "Helix data creators" and "humanoid robot pilots" who wear teleoperation equipment to guide robots through tasks and upload the results. What that pipeline feeds is now on a factory floor: Helix 02, Figure's pixels-to-actions model, drives the Figure 03 robots that arrived in Hall 52 of BMW's Spartanburg plant on June 30, 2026 to sequence unsorted parts into delivery trolleys, a harder job than the sheet-metal loading its predecessor did there across more than 30,000 X3s in 2025.
1X took the most direct version of the bet to consumers. Buyers of its Neo home robot — $20,000 for priority delivery in 2026, or $499 a month for a later slot — are told upfront that human teleoperators will be viewing inside their homes while piloting the robot, and 1X opened a Hayward, California factory on April 30, 2026 rated at 10,000 units a year to build them.
Founder Bernt Børnich was explicit about why: "If we don't have your data, we can't make the product better." Tesla moved in the opposite direction in 2025, dropping motion-capture suits and teleoperation rigs in favor of workers wearing helmet-and-backpack camera arrays that record ordinary tasks like folding a shirt, a change made specifically to collect data faster once Ashok Elluswamy took over the Optimus program from Milan Kovac.
Egocentric Video Pretraining: Borrowing Priors from Human Hands
The third fix skips robots and teleoperation rigs altogether and mines video of humans simply living their lives. Meta's Ego4D dataset, built with 13 university partners, contains nearly 3,700 hours of first-person video from 931 camera wearers across 74 locations in nine countries — recordings of cooking, repairing, gardening, and other everyday manipulation with no robot involved.
NVIDIA's EgoScale project, published in February 2026 with UC Berkeley and University of Maryland collaborators, pretrained a vision-language-action policy on 20,854 hours of egocentric human video captured with motion-tracked gloves, more than 20 times the scale of prior human-video pretraining efforts.
The paper found a clean log-linear scaling law between pretraining data volume and downstream robot performance — early evidence that robot policies improve predictably with more data the way language models do — and the resulting policy beat a no-pretraining baseline by 54% on a 22-degree-of-freedom dexterous hand.

Meta's V-JEPA 2 world model took a related but distinct approach: pretrain on roughly a million hours of general internet video to learn physical intuition, then fine-tune on under 62 hours of robot-specific footage drawn from DROID. The resulting system, V-JEPA 2-AC, deployed zero-shot on Franka robot arms in two separate labs it had never operated in before, achieved 65 to 80% success on pick-and-place tasks by imagining roughly 16 seconds of consequences before acting, without collecting a single frame of data in either lab beforehand.
Sim-to-Real Gap Explained: Why Specular and Transparent Objects Still Break Policies
None of these three fixes has closed the visual sim-to-real gap, which researchers increasingly point to as a leading cause of manipulation-policy failure once a robot leaves the lab. Domain randomization — varying textures, lighting, and object properties inside a simulator during training — remains the default mitigation, but it does not reliably cover specular and transparent surfaces: a stainless-steel bowl, a drinking glass, a plastic food container.
Rendering engines built for speed rather than physical accuracy often skip ray tracing, so they cannot reproduce how light actually bounces off a mirror-finish countertop or refracts through water, and a policy trained against that flawed rendering learns visual cues that do not exist in the real room.
Academic sim-to-real benchmarks have measured manipulation success rates dropping 24 to 30% when a policy trained purely in simulation meets real cameras and real lighting, and researchers building physically based rendering pipelines have found that accurate materials and plausible light are necessary, not optional, for closing that gap. Soft-body physics compounds the problem: fabric dynamics, sponge deformation, and drawer friction still are not modeled with enough fidelity for a simulator-trained grasp policy to transfer reliably.
The GPT-2.5 Moment: What 2026's Funding Surge Actually Buys
Bessemer Venture Partners framed where the field stands in April 2026 as "the GPT-2.5 moment for robotics" — capabilities are real and moving fast, and scaling laws are starting to show up in papers like EgoScale, but the distance between a lab demo and a 99.9% field-reliability bar remains wide. Physical Intelligence has shipped four generalist policies in eighteen months to make the point: π0 folded laundry and bussed tables after training on 68 tasks across seven robot platforms, and π0.7, released on April 16, 2026, reports zero-shot cross-embodiment transfer — a robot folding laundry it was never shown the task on — and operating an espresso machine out of the box.
Bessemer's own math shows why that bar matters: a policy succeeding 95% of the time at each step of a ten-step task completes the full task only around 60% of the time, a failure rate no warehouse or factory floor can absorb.
None of the three current fixes fully replaces the internet-scale corpus that made chatbots possible, and combining them is now the explicit strategy at every serious lab: simulation for locomotion and scale, teleoperation for the dexterity data only humans can generate reliably, and egocentric video for the broad physical priors neither of the other two sources supplies cheaply.
The fourth hope is that deployed fleets become data engines no dataset release could match, and Tesla's own year shows how slowly that arrives. On the January 28, 2026 earnings call Musk said the several hundred Optimus units inside Tesla's factories were there primarily for learning and data collection rather than productive work — data engines in the literal sense, and nothing else yet. Tesla had still not shown Optimus Gen 3 publicly by August 2026, when it told investors the design was finalised and commercial sales could start "as early as" the second half of 2027. Musk's guidance was blunt.
No, Optimus production will be extremely slow at first, as everything is new. This is not like making a car. — Elon Musk, CEO, Tesla
XPeng switched on an automated humanoid line in Guangzhou in September 2026 and is aiming at volume production by year end, which would be the first fleet large enough to test the data-engine theory. Until one of these lines runs, the corpus is still being assembled one arm-hour at a time, exactly as Goldberg described it.